Visualization method and system for test management in feature engineering
By using random forest algorithms and feature weight evaluation methods in the autonomous driving logistics vehicle ring test, the problem that traditional test management methods are difficult to deal with multi-dimensional data is solved, efficient test management and accurate feature recognition are achieved, and testing efficiency and result accuracy are significantly improved.
Patent Information
- Application Number
- CN202510340837.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the ring testing of autonomous logistics vehicles, traditional test management methods are difficult to efficiently handle multi-dimensional sensor input and control system status, resulting in inefficient testing and difficulty in quickly identifying key features.
Random forest algorithm is used to combine feature weights (PR value) evaluation to efficiently manage the test process through visual means. Specific steps include data preprocessing, feature selection, feature weight calculation, test case optimization and visual presentation.
It significantly improves the efficiency and coverage of autonomous driving tests, accurately identify key features and prioritizes the execution of relevant test cases, reduces resource waste and redundant operations, and improves the accuracy and reliability of test results.
Smart Images

Figure CN120144989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to a method for optimizing feature engineering in the in-loop testing of autonomous driving of logistics vehicles and a visualization system for test management based on machine learning. Background Art
[0002] In the test link of autonomous driving logistics vehicles, especially in the in-loop testing, the test data involves various complex sensor inputs and control system states. Traditional test management methods lack efficient processing of these multi-dimensional data, resulting in low test efficiency and difficulty in quickly identifying key features. The technical solution provided by the present invention uses a random forest algorithm combined with feature weight (PR value) evaluation to efficiently manage the test process through visualization means, improving the quality and accuracy of autonomous driving tests.
[0003] Therefore, the present solution particularly proposes a visualization method and system for test management in feature engineering to solve the above problems. Summary of the Invention
[0004] To overcome the defects of the prior art, the purpose of the present invention is to provide a visualization method and system for test management in feature engineering.
[0005] To achieve the above object, the technical solution of the present invention is realized as follows: A visualization method for test management in feature engineering, which is applied to the in-loop testing link of autonomous driving of logistics vehicles, and is carried out through the following steps: a. Data preprocessing step: Preprocess the original sensor data, control data and other environmental information during the in-loop testing of the logistics vehicle to generate a standardized feature vector; data preprocessing includes denoising, missing value filling, normalization and standardization processing, etc.; b. Feature selection step: Use the random forest algorithm to analyze the preprocessed feature data, and select the key features that have a significant impact on the performance of autonomous driving tests through the feature importance scoring mechanism in the tree model; the random forest algorithm trains multiple decision trees and comprehensively evaluates the features based on their results, which can effectively handle the non-linear relationships and high-dimensional data in the dataset; c. Feature weight calculation step: Based on the feature importance score calculated by the random forest model, further calculate the PR value (feature weight) of each feature; this PR value reflects the influence degree of each feature on the test result and is used in subsequent test management decisions; the calculation of the PR value can be carried out through feature importance evaluation or other feature selection algorithms in machine learning algorithms; d. Test case optimization step: Optimize the test cases according to the calculated PR values, and preferentially select those features that have a greater impact on system performance for testing, so as to improve the test coverage rate; the test cases with high feature weights will be executed preferentially to improve the test effectiveness; e. Visual display steps: Visualize the results of feature selection and PR value calculation to form intuitive test result charts and data distribution diagrams; through a dynamic test data visualization interface, display the contribution degree, importance of each feature and its correlation with system performance; this process includes but is not limited to: feature importance bar charts, PR value heat maps and test case priority diagrams; f. Test management decision-making steps: Combine the PR values of the features to automatically adjust the test strategy; based on the test priorities displayed on the visualization interface, automatically adjust the test order and test scenarios, so as to ensure the test priorities and comprehensiveness of key features.
[0006] Preferably, in the feature selection step, the random forest algorithm obtains the feature importance score by calculating the contribution degree of each feature to the decision tree model. The importance score of the feature reflects its influence on system performance, and then selects the features that have the most influence on the behavior of the autonomous driving system.
[0007] Preferably, the PR value is calculated by the following formula in the feature weight calculation step: where I i represents the importance score of the i-th feature, and n is the total number of features. The PR value reflects the relative importance of the feature to the test result. Features with higher PR values are considered to have a greater impact on system performance and should be tested preferentially.
[0008] Preferably, the visualization display step includes the following contents: a. Feature importance display: Display the weight (PR value) of each feature through a bar chart or a pie chart, so that testers can clearly understand the contribution degree of each feature to the test result; b. PR value heat map: Generate a heat map of the PR value to show the change of the PR value of each feature under different test scenarios, so as to help testers identify in which specific scenarios the feature importance is more prominent; c. Test case priority display: Generate a test case priority list according to the PR value calculation result to ensure that the test cases for high-weight features are executed preferentially and ensure the efficient allocation of test resources.
[0009] Preferably, the system includes the following modules: a. Data acquisition module: It is used to collect sensor data, control data, environmental data, etc. from the logistics vehicle autonomous driving system for subsequent feature engineering processing; b. Feature processing module: It uses the random forest algorithm to perform feature selection and processing on the collected raw data, generates a feature set and evaluates the feature importance; c. PR value calculation module: According to the feature importance evaluation in the random forest algorithm, it calculates the PR values of each feature and outputs weight data for the test management module to use; d. Visualization display module: It generates an intuitive graphical interface to display information such as the PR value of each feature, the feature importance distribution, and the test case priority, to help test engineers perform data analysis and decision-making.
[0010] e. Test management module: According to the PR value and the visualization result, it automatically adjusts the test strategy, preferentially selects features with high weights for testing, and optimizes the execution order and test coverage of test cases.
[0011] Preferably, the test management module can automatically adjust the priority of test cases according to the PR value calculation result, generate corresponding test reports and adjust the test strategy in real time to ensure the comprehensiveness and efficiency of testing.
[0012] Preferably, the visualization display module adopts a dynamic interactive interface, and test engineers can view the feature importance and test progress in real time based on the graphical interface for multi-dimensional analysis and strategy adjustment.
[0013] Preferably, the system continuously optimizes feature selection, weight calculation and test strategy through machine learning algorithms and feedback mechanisms to adapt to the requirements of different test scenarios and improve test accuracy.
[0014] The beneficial effects of the present invention are reflected in: Through optimizing feature selection, PR value calculation and test case optimization, the present invention significantly improves the efficiency and coverage of autonomous driving testing. By accurately identifying key features and preferentially executing relevant test cases, it reduces resource waste and redundant operations, and at the same time improves the accuracy and reliability of test results. The system has the ability to dynamically adjust strategies, can optimize the test process in real time according to the test progress and feedback, reduces the test cost, and provides an intuitive visualization management to enhance decision-making support. Generally speaking, the present invention provides an efficient, intelligent and sustainable management platform for autonomous driving testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In the drawings: Figure 1 : Schematic diagram of the test management system architecture; Figure 2 : Flow chart of feature selection and random forest algorithm; Figure 3 : Flowchart for PR value calculation and test case optimization. Specific implementation manners
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the invention without creative efforts shall fall within the protection scope of the invention.
[0017] In addition, "a plurality of" means two or more. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions conflicts with each other or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the invention.
[0018] Please refer to the accompanying specification Figures 1-3 , the present invention provides a visualization method and system for test management in feature engineering, which is mainly applied to the feature engineering optimization and test management of the in-loop test of logistics vehicle autonomous driving. The implementation manners of the present invention will be described in detail below with reference to specific embodiments.
[0019] Feature engineering optimization and test management based on autonomous driving test 1. Data collection and preprocessing During the process of logistics vehicle autonomous driving test, the system first needs to collect data from different types of sensors (such as lidar, camera, ultrasonic sensor, GPS, IMU, etc.) and in-vehicle control systems. This process includes the following aspects: Sensor data: Obtain road information, obstacle detection data, etc. from environmental sensors.
[0020] In-vehicle control data: Collect control data including vehicle speed, steering angle, acceleration, etc.
[0021] Environmental data: Obtain information such as weather and road conditions from in-vehicle systems or external information sources.
[0022] These data are transmitted to the data preprocessing module for the following processing: Denoising processing: Remove the noise caused by sensor or measurement errors, for example, process radar data through a Kalman filter.
[0023] Missing value filling: Use interpolation methods (such as linear interpolation) to fill the missing values in sensor data to ensure the integrity of the data.
[0024] Data normalization and standardization: Normalize features of different scales to standardize all data to the same scale, eliminating the impact of dimensional differences on subsequent algorithms.
[0025] The preprocessed data will be used as features and input into the subsequent feature selection module.
[0026] 2. Feature selection and importance evaluation Use the random forest algorithm to perform feature selection and importance evaluation on the data. This algorithm evaluates the influence of each feature on the prediction result by constructing multiple decision trees, and then selects the most influential features. The specific process is as follows: Train decision trees: The random forest algorithm randomly selects samples from the training data through the Bootstrap method and uses these samples to construct multiple decision trees. During the node splitting process of each tree, a random subset of features is selected to determine the best splitting feature.
[0027] Feature importance evaluation: Each decision tree evaluates the "importance" of each feature by calculating the Gini index or information gain when splitting nodes. The average value of the evaluation results of all decision trees is the importance score of each feature.
[0028] Feature selection: According to the importance score of each feature, select the features that have a greater impact on the system performance for subsequent analysis. The importance score of the feature can help us clarify which features need to be given priority in the test.
[0029] 3. Feature weight (PR value) calculation After feature selection, we need to calculate the PR value according to the contribution degree (i.e., importance score) of each feature. The calculation formula of the PR value is as follows: where I i represents the importance score of the i-th feature, and n is the total number of features. The PR value reflects the relative importance of the feature's influence on the test result. Features with higher PR values are considered to have a greater impact on the system performance and should be tested preferentially.
[0030] The calculation result of the PR value will provide a basis for subsequent test case optimization.
[0031] 4. Test case optimization and test priority determination According to the PR values of the features, the system will automatically generate optimized test cases and adjust the priorities of the test cases according to the weights of the features. The main steps are as follows: Test case generation: Test cases are generated by selecting different combinations of feature values, and these combinations cover the key scenarios that may affect the autonomous driving system.
[0032] Priority Sorting: Based on the feature-based PR value, the system sorts the test cases by priority to ensure that test resources are concentrated on testing high-PR-value features. Test cases corresponding to high-PR-value features will be executed first.
[0033] Improved Test Coverage: By reasonably allocating feature weights, the comprehensiveness of testing is ensured. Each test case corresponds to one or more high-weight features, thus improving the overall test coverage and efficiency.
[0034] 5. Visualization of the Testing Process The present invention provides a way of visual display, enabling test engineers to intuitively view and manage key data in the testing process. The visualization interface includes the following: Feature Importance Chart: Displays the PR value of each feature through a bar chart or pie chart, helping testers intuitively understand the impact degree of each feature on the test result.
[0035] PR Value Heat Map: Displays the distribution of PR values of different features in different test scenarios. Testers can identify which features have higher importance in specific scenarios based on the heat map, thereby adjusting the test strategy.
[0036] Test Case Priority Chart: The system generates a test case priority chart, showing the execution order and priority of different test cases, ensuring that test cases for high-PR-value features can be executed first.
[0037] Dynamic Progress Bar and Feedback: The visualization interface also shows the real-time test progress. Testers can see the execution status of the current test case and the coverage of feature testing.
[0038] 6. Test Management and Decision Support Based on the test strategy management of PR values, the system can dynamically adjust the test strategy according to the test progress and test case priority to ensure the efficient execution of test tasks. The system provides the following decision support functions: Automatically Adjust the Test Order: According to the PR value, the system can dynamically adjust the execution order of test cases to ensure that testing of high-weight features is prioritized.
[0039] Real-Time Feedback and Report: During the testing process, the system provides real-time feedback based on the test results and automatically generates a test report. The report includes feature importance, PR value calculation, test results, and test efficiency analysis, helping testers adjust the test strategy in a timely manner.
[0040] Optimization of Test Resources: The system can reasonably allocate test resources according to the PR value and test case priority to ensure that high-priority test cases receive sufficient execution resources.
[0041] 7. System Iteration and Adaptive Optimization The system of the present invention can not only improve efficiency in the first test, but also continuously optimize the test strategy through a feedback mechanism. Based on the continuously accumulated test data, the system will optimize the feature selection and test case generation algorithms to adapt to different test requirements: Model Adaptive Update: The system will automatically adjust the feature selection strategy and the PR value calculation model according to new test data and feedback. As the amount of test data increases, the system will gradually improve the accuracy of feature selection and optimize the test results.
[0042] Online Learning and Model Update: As new environmental data continuously pours in, the system can achieve online learning, dynamically update the parameters of the model, and further improve the test quality.
[0043] Implementation Effects Through the above specific implementation manners, the technical solution of the present invention effectively improves the feature engineering management and test efficiency of the in-loop test of the autonomous driving logistics vehicle. Through the automatic adjustment of the test case priority and the visualization management based on the PR value, the efficiency, comprehensiveness, and accuracy of the test process are ensured. The effective utilization of complex sensor and environmental data in the autonomous driving test in feature engineering optimization and test management significantly reduces the test cycle and cost, while improving the reliability of the test results.
[0044] This implementation manner combines the feature selection technology of machine learning, the test case optimization strategy, and the visualization management function, which not only improves the efficiency of the test process, but also provides a sustainable optimization solution for subsequent autonomous driving tests.
[0045] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0046] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
[0047] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A visualization method for test management in feature engineering, applied to the in-loop test of autonomous driving of logistics vehicles, characterized by: Proceed through the following steps: a. Data preprocessing steps: Preprocess the raw sensor data, control data and other environmental information during the logistics vehicle in-loop test to generate standardized feature vectors; data preprocessing includes denoising, missing value filling, normalization and standardization processing; b. Feature selection steps: The random forest algorithm is used to analyze the preprocessed feature data, and the key features that have a significant impact on the autonomous driving test performance are selected through the feature importance scoring mechanism in the tree model. The random forest algorithm trains multiple decision trees and integrates their results to evaluate features, which can effectively deal with nonlinear relationships and high-dimensional data in the data set. c. Feature weight calculation steps: Based on the feature importance score calculated by the random forest model, the PR value (feature weight) of each feature is further calculated; the PR value reflects the degree of influence of each feature on the test results and is used in subsequent test management decisions; the PR value can be calculated through feature importance evaluation or feature selection algorithms in other machine learning algorithms; d. Test case optimization steps: According to the calculated PR value, the test cases are optimized, and those features that have a greater impact on system performance are prioritized for testing, thereby improving the test coverage; test cases with high feature weights will be executed first to improve the effectiveness of the test; e. Visualization steps: Visualize the results of feature selection and PR value calculation to form intuitive test result charts and data distribution diagrams; display the contribution, importance and correlation of each feature with system performance through a dynamic test data visualization interface; This process includes but is not limited to: feature importance bar chart, PR value heat map and test case priority map; f. Test management decision steps: The test strategy is automatically adjusted based on the PR value of the feature. Based on the test priority displayed in the visual interface, the test sequence and test scenarios are automatically adjusted to ensure the test priority and comprehensiveness of key features.
2. The visualization method for test management in feature engineering according to claim 1, characterized in that: In the feature selection step, the random forest algorithm calculates the contribution of each feature to the decision tree model to obtain a feature importance score. The feature importance score reflects its impact on system performance, and then selects the features that have the greatest impact on the behavior of the autonomous driving system.
3. The visualization method for test management in feature engineering according to claim 1, characterized in that: The feature weight calculation step calculates the PR value by the following formula: Among them, I i It represents the importance score of the i-th feature, and n is the total number of features. The PR value reflects the relative importance of the feature on the test results. Features with higher PR values are considered to have a greater impact on system performance and should be tested first.
4. The visualization method for test management in feature engineering according to claim 1, characterized in that: The visual display step includes the following contents: a. Feature importance display: The weight (PR value) of each feature is displayed through a bar chart or pie chart, so that testers can clearly understand the contribution of each feature to the test results; b.PR value heat map: Generate a PR value heat map to show the PR value changes of each feature under different test scenarios, so as to help testers identify in which specific scenarios the importance of features is more prominent; c. Test case priority display: Generate a test case priority list based on the PR value calculation results to ensure that test cases with high-weight features are executed first and to ensure efficient allocation of test resources.
5. A test management system for implementing the method of claim 1, characterized in that: The system includes the following modules: a. Data acquisition module: used to collect sensor data, control data, and environmental data from the logistics vehicle's autonomous driving system for subsequent feature engineering processing; b. Feature processing module: Use the random forest algorithm to select and process the collected raw data, generate feature sets and evaluate feature importance; c.PR value calculation module: Calculate the PR value of each feature based on the feature importance evaluation in the random forest algorithm, and output weight data for use by the test management module; d. Visual display module: Generates an intuitive graphical interface to display information such as the PR value of each feature, feature importance distribution, test case priority, etc., to help test engineers perform data analysis and decision-making. e. Test management module: Automatically adjust the test strategy based on the PR value and visualization results, give priority to high-weight features for testing, and optimize the execution order and test coverage of test cases.
6. The system according to claim 5, characterized in that The test management module can automatically adjust the priority of test cases according to the PR value calculation results, generate corresponding test reports and adjust the test strategy in real time to ensure the comprehensiveness and efficiency of the test.
7. The system according to claim 5, characterized in that The visualization module uses a dynamic interactive interface, and test engineers can view the importance of features and test progress in real time based on the graphical interface, and perform multi-dimensional analysis and strategy adjustment.
8. The system according to claim 5, characterized in that The system continuously optimizes feature selection, weight calculation and testing strategy through machine learning algorithms and feedback mechanisms to adapt to the needs of different testing scenarios and improve test accuracy.