A Method for Evaluating the Intelligence of Autonomous Vehicles Based on Subjective-Objective Mapping of a Large Language Model
By constructing a subjective-objective mapping evaluation framework based on a large language model, the problems of model simplicity and subjective dependence in the evaluation of autonomous driving systems are solved, and a more accurate and efficient intelligence evaluation is achieved, which is applicable to the performance evaluation of autonomous vehicles and their subsystems at different levels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-26
AI Technical Summary
Existing quantitative evaluation methods for autonomous driving systems are too simplistic, making it difficult to fit nonlinear relationships, and they rely on subjective expert judgment, resulting in inaccurate and inconsistent evaluation results.
A subjective-objective mapping evaluation framework based on a large language model is constructed. By combining an external knowledge base and a multi-layer perceptron model with multi-dimensional quantitative indicators, a quantitative evaluation of the intelligence of autonomous driving is achieved, reducing subjective dependence and improving the objectivity and accuracy of the evaluation.
It improves the objectivity and efficiency of autonomous driving system evaluation, can fit nonlinear relationships, provides more accurate intelligence evaluation results, and reduces human resource consumption and data labeling costs.
Smart Images

Figure CN121597823B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of testing and evaluation of autonomous driving systems, and in particular to a method for evaluating the intelligence of autonomous vehicles based on the subjective-objective mapping of a large language model. Background Technology
[0002] Currently, autonomous vehicles are transitioning from Level 1 (driver assistance) and Level 2 (partial automation) to Level 3 (high automation) and higher. However, existing quantitative evaluations of autonomous driving systems still rely on traditional methods, primarily involving weighted and aggregated indicators. As the correlation between multidimensional indicators increases and the complexity of system behavior rises, these methods are gradually revealing their limitations.
[0003] First, the models are too simplistic and struggle to fit nonlinear relationships. Most current quantitative evaluation models weight and sum various indicators to produce a single comprehensive score that assesses the intelligence level of autonomous driving. However, in reality, these indicators may not be independent, and the aggregation of indicators may be nonlinear. All of this makes it difficult for simple linear models to accurately evaluate the intelligence level of autonomous driving.
[0004] Secondly, there is an over-reliance on subjective judgment. Determining indicator weights generally employs a subjective approach, incorporating expert opinions. However, this evaluation method heavily depends on expert subjective judgment, and the results may be influenced by the expert group's knowledge and experience. While objective analytical methods exist that derive weights from large-scale sample data, these methods rely on mathematical models and have high data quality requirements, thus limiting their applicability.
[0005] To address the aforementioned challenges, this invention proposes an evaluation framework based on the subjective-objective mapping of a large language model for assessing the intelligence level of advanced autonomous driving systems, which can provide more accurate and reasonable evaluation results. Summary of the Invention
[0006] The main object of this invention is an autonomous driving intelligence evaluation method based on the subjective-objective mapping of a large language model. First, an external knowledge base is constructed, and knowledge enhancement is used to improve the understanding of autonomous driving professional knowledge by the large language model. Then, autonomous driving interaction data converted into natural language text is input into the model, allowing the large language model to perform quantitative evaluation of the interaction data. Finally, a multilayer perceptron is used to fit the evaluation results, thereby constructing an evaluation model.
[0007] The specific technical solution of the present invention is as follows:
[0008] A method for evaluating the intelligence of autonomous vehicles based on the subjective-objective mapping of a large language model includes the following steps:
[0009] S1. Prepare the dataset;
[0010] First, 1,000 data points from real driving scenarios were selected from a public dataset as the training set for the model. Then, the data was filtered based on the interaction intensity, and data with high interaction intensity was extracted as the data source for knowledge enhancement of the large language model in S3.
[0011] S2 Determine the evaluation dimensions;
[0012] A five-dimensional evaluation system for autonomous driving intelligence is constructed, including safety, comfort, efficiency, social interactivity, and impact on the transportation system; specific evaluation indicators and their quantification methods are determined for each dimension, and the quantified evaluation indicators will serve as input layer variables for the neural network in S4.
[0013] S3 uses a large language model to quantitatively evaluate autonomous driving data;
[0014] First, a data-natural language mapping module is constructed to map the interactive fragment data selected in S1 into natural language prompts. Then, knowledge from the field of autonomous driving, such as traffic rules and driving regulations, is compiled as an external knowledge base for the large language model. Subsequently, these interactive information described by natural language are given to human experts for fuzzy evaluation. Finally, the expert fuzzy evaluation results are incorporated into the knowledge base, and the prompts are input into the knowledge-enhanced large model to obtain the intelligence evaluation results.
[0015] S4 utilizes neural networks to construct a quantitative evaluation model for intelligent autonomous driving;
[0016] First, the input data is standardized, including normalizing the intelligent scores output by the large language model and standardizing the evaluation indicators of each dimension. Then, a multilayer perceptron is used for regression, allowing the neural network to capture the non-linear relationship between the interactive data and the intelligent scores.
[0017] Beneficial effects
[0018] Compared with the prior art, the present invention has the following advantages:
[0019] (1) Improve the objectivity of the assessment. Traditional assessment methods rely heavily on human experts, and due to differences in the knowledge and experience of experts, the assessment results are highly subjective. In contrast, the large language model in this invention learns a large number of expert opinions through knowledge reinforcement, and can provide assessment results with higher objectivity and consistency.
[0020] (2) Improving evaluation efficiency while controlling costs. This invention constructs a fully trained autonomous driving evaluation perceptron network that can directly provide evaluation results based on autonomous driving interaction data. The entire process does not require the participation of human experts, resulting in higher evaluation efficiency. In addition, this invention uses a knowledge-enhanced large language model to complete data annotation work at low cost and high efficiency, solving the problem of high data annotation costs during supervised learning model training.
[0021] (3) It can fit nonlinear relationships. Traditional multi-index evaluation methods treat each dimension as an independent index and perform weighted summation. They also rely on fixed weights and cannot reflect nonlinear relationships and changes in the importance of dimensions under different scenarios. However, this invention uses a multilayer perceptron for regression, which can fully capture the nonlinear relationships between the indicators of each dimension. At the same time, it implicitly learns the differences in the weights of indicators under different scenarios, so that the model can more flexibly and accurately reflect the true comprehensive performance of the system. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the construction process of the intelligence evaluation model of this invention.
[0023] Figure 2 The Turing test results are from an embodiment of the present invention.
[0024] Figure 3 This is a comparison of the model validity verification results in the embodiments of the present invention. Detailed Implementation
[0025] This invention relates to an autonomous driving intelligence evaluation and testing system. Its core lies in constructing an intelligent evaluation framework based on the subjective-objective mapping of a large language model. This framework possesses universality and scalability, and can be applied to the performance evaluation of autonomous vehicles and their subsystems (such as decision-making and planning, behavioral interaction modules) at different levels. During implementation, it is required that each test scenario and each tested vehicle remain independent of each other. The following will further elaborate on the implementation methods of this invention through the actual configuration of model selection, data processing flow, and evaluation mechanism, combined with specific implementation steps.
[0026] Figure 1 This is a flowchart illustrating the construction process of the intelligence evaluation model of this invention.
[0027] Example
[0028] A method for evaluating the intelligence of autonomous vehicles based on the subjective-objective mapping of a large language model includes the following steps:
[0029] S1. Prepare the dataset;
[0030] This embodiment uses the Waymo public dataset, specifically designed for motion prediction tasks, which contains motion state data and map information from 1000 real-world driving scenarios. This dataset was then used as the basis for training and evaluating the model. The InterHub system was used to quantify the interaction intensity between the vehicle and surrounding vehicles using the minimum sum of absolute accelerations, and interaction events were identified based on trajectory conflict points to extract interaction segments. This embodiment extracted a total of 5200 valid interaction segments, which were categorized into three types based on the main interaction behaviors of vehicles on urban roads: following, crossing, and merging.
[0031] S11 State initialization and trajectory prediction;
[0032] First, the vehicle state vector is extracted, which includes position, velocity, and acceleration. Then, moving vehicles are identified as the objects for trajectory prediction. Finally, the trajectory within the next 5 seconds is predicted based on the current vehicle state.
[0033] S12 identifies trajectory intersections;
[0034] Two-dimensional geometric modeling of vehicle trajectories is performed using the Shapely geometry library, representing vehicle trajectories as linear geometric objects. Based on its provided geometric relationship judgment and intersection operation functions, the geometric intersection points between different vehicle trajectories are calculated, thereby identifying potential trajectory intersection locations. For the identified intersection points, the distance calculation function of the Shapely library is further used to calculate the path distance from each vehicle along its own trajectory to the intersection point. Combined with the vehicle's travel speed information, the arrival time of the vehicles at the intersection point is estimated, thus obtaining the time difference between different vehicles arriving at the intersection point.
[0035] S13 Determines valid interaction;
[0036] The system uses a buffer verification method to exclude cases where the vehicle has already passed safely: The system creates a 1.5-meter horizontal parallel buffer zone around the trajectory of the preceding vehicle and verifies whether the trajectory of the following vehicle intersects with the buffer zone; if they intersect and the trajectory of the following vehicle is contained within the buffer polygon, it means that the preceding vehicle has passed safely and is determined to be an invalid interaction; otherwise, it is determined to be a valid interaction.
[0037] S14 quantifies the interaction intensity;
[0038] The interaction intensity is quantified using the minimum absolute acceleration comprehensive quantization method: the system establishes a constrained optimization model, and the objective function is to minimize the sum of the absolute values of the accelerations of all participating vehicles.
[0039] By solving this optimization problem, the minimum combined acceleration requirement that enables both vehicles to avoid a collision is obtained, which serves as a quantitative indicator of the interaction intensity.
[0040] The constraints include: (1) the time difference between the arrival of the two vehicles at the collision point does not exceed 1.5 seconds; (2) the actual displacement of the vehicles deviates from the expected displacement by no more than 0.1 meters; and (3) the vehicle displacement and final speed are positive.
[0041] S15 Extract interactive segments;
[0042] The system sets an acceleration threshold (0.01 m / s²) and identifies time periods where the acceleration continuously exceeds the threshold, retaining only segments lasting longer than 0.3 seconds. For each segment, the maximum interaction intensity and minimum PET (post-intrusion time) value are calculated as key features of the interaction event.
[0043] S2 Determine the evaluation dimensions;
[0044] This invention constructs a five-dimensional evaluation system, which includes safety, comfort, efficiency, social interactivity, and impact on the transportation system, and quantifies these into quantitative evaluation indicators.
[0045] The evaluation indicators and quantification methods for each dimension are as follows:
[0046] S21 security;
[0047] This invention uses Time-of-Collision (TTC) to reflect a vehicle's control over safety. This indicator comprehensively considers both following distance and relative speed, and is therefore selected as one of the safety indicators. Its quantification method is as follows:
[0048]
[0049] in This indicates the distance between the rear of the car in front and the front of the car behind, with the direction from the rear car to the front car. This indicates the relative speed between the two vehicles. The TTC safety threshold is set to 1.5 seconds.
[0050] In addition, Post-Intrusion Time (PET) is another safety assessment metric, which focuses more on quantifying vehicle driving safety in intersection scenarios. Assuming that vehicles A and B are about to collide at an intersection, and vehicle A passes the collision point before vehicle B, the PET is calculated as follows:
[0051]
[0052] in It is the displacement of car A from the current moment until its body has just completely left the point of conflict. It is the displacement of car B from the current moment until any point on its body just enters the point of conflict. and The instantaneous speeds of the two vehicles are given. The PET safety threshold is set to 1.0 second.
[0053] S22 Comfort;
[0054] Autonomous vehicles should perform reasonable steering and acceleration / deceleration during operation to provide a comfortable riding experience for passengers. Through field tests, a systematic survey of passenger riding experience was conducted, and statistical methods were used to determine the threshold values of vehicle state parameters under different comfort levels, as shown in the table below:
[0055]
[0056] S23 Efficiency;
[0057] Metrics include: TaskTime and Average Delay (AvgDelay);
[0058] For the same distance traveled, efficiency is negatively correlated with task time (TaskTime). TaskTime is defined as the difference between the theoretical minimum task time and the actual task time. The ratio, while the theoretical shortest task time is determined by the minimum task path. Speed limits on road sections Calculated from:
[0059]
[0060] In addition, the average delay (AvgDelay) metric is defined to measure efficiency. This metric highlights the potential for acceleration in autonomous vehicles and is defined as follows:
[0061]
[0062] in The average speed of the main vehicle during this mission; and These represent the minimum and maximum average speeds of surrounding vehicles, respectively. If AvgDelay is small, it indicates that the speeds of surrounding vehicles are relatively high, meaning the vehicle's relative position in the traffic flow is gradually lagging behind, and there is still considerable room for acceleration.
[0063] S24 Social interactivity;
[0064] Autonomous vehicles should make reasonable interaction decisions based on right-of-way and the status of other vehicles. This invention extends the I / O metric to general scenarios to evaluate the rationality of autonomous vehicle decisions. The I / O calculation method is as follows:
[0065]
[0066]
[0067]
[0068] ITSI stands for Interactive Transient State Index. This represents the normalized time difference of a vehicle passing through the conflict zone. It is the normalized cooperative acceleration of the vehicle. The softmax function can compare... The magnitude of the value maps larger values to near 1 and smaller values to near 0, making the ITSI close to the larger of the two, thus highlighting the factors with higher risk coefficients. This represents the normalized vehicle motion behavior trajectory characteristics. This represents the longitudinal displacement of the main vehicle at the current moment. and These represent the minimum and maximum longitudinal displacements achievable by the vehicle under collision boundary acceleration. If IO is close to 0, it indicates that ITSI is... At the same time, it is close to 1, which means that the main vehicle still adopted a relatively aggressive interaction strategy while the interaction risk was high, so the interaction was considered unreasonable.
[0069] The impact of S25 on the transportation system;
[0070] Imperfections in autonomous driving systems can affect the normal operation of other vehicles, leading to a decrease in the traffic capacity of the transportation system. This impact is defined as the degree to which the autonomous vehicle, after entering the transportation system, causes a decrease in the driving efficiency of other vehicles that directly or indirectly interact with it, as shown in the following formula:
[0071]
[0072] in, The number of vehicles in the background. and These are the first and last times before and after an autonomous vehicle enters the transportation system. The speed of the affected background vehicle. A relatively large value indicates that the average impact of autonomous vehicles on the speed of other vehicles is relatively large, meaning that the overall impact on the traffic system is relatively small.
[0073] In summary, 10 quantitative evaluation indicators were obtained: time to collision (TTC), time to intrusion (PET), and longitudinal acceleration. lateral acceleration lateral speed Yaw angle, angular velocity yr, mission time Average delay (AvgDelay), interoperability index (IO), and impact on the traffic system. .
[0074] S3 uses a large language model to quantitatively evaluate autonomous driving data;
[0075] Using large language models to evaluate the intelligence performance of autonomous driving instead of human experts can effectively avoid problems such as inconsistent evaluations caused by subjectivity, while also reducing human resource consumption and improving evaluation efficiency.
[0076] S31 Data-Natural Language Mapping Module;
[0077] This invention constructs a data-natural language mapping module (DLMM) between the original interactive data and the large language model, which transforms complex autonomous driving data into a form that is easy for the large language model to understand, and finally adds it to the prompt word.
[0078] The state of an autonomous driving system during driving is represented by a state-space model:
[0079]
[0080] in, For the vehicles that interact, and For vehicle coordinates, and The longitudinal and lateral velocities and accelerations of the vehicle. This is the vehicle's heading angle. To identify the occurrence of interactive events, define... Define the effective interaction range. Based on different interaction types, define the effective interaction ranges for following, merging, and crossing. , and as follows:
[0081]
[0082]
[0083]
[0084] in, This is the lane vector of the interacting vehicle, with each element being the lane ID. The angle difference threshold, To maintain following distance, and The following distance threshold and conflict distance threshold. This represents the lane vector of the vehicle that changed lanes. The conjunction symbol, The number of points of conflict in the future trajectories of the two vehicles. For vehicles Distance to the point of conflict.
[0085] This formula describes the criteria for distinguishing the three types of interactions, and each condition must be satisfied simultaneously:
[0086] 1) Following: The two vehicles are in the same lane; the difference in heading angle between the two vehicles is small enough; the distance between the two vehicles is close enough.
[0087] 2) Merging: The heading angles of the two vehicles differ sufficiently; at least one of the vehicles is crossing a lane line.
[0088] 3) Crossing: There is a potential collision between the two vehicles; at least one of the vehicles is close enough to the point of collision.
[0089] function Used to detect abnormal behavior of the main vehicle:
[0090]
[0091] in, These are vehicle state quantities, including speed, acceleration, and heading angle. This is the abnormal threshold. This is an indicator function. It is triggered when the rate of change of any state variable exceeds a threshold. Identify and label abnormal moments.
[0092] definition The vehicle's behavior is described in natural language, which is determined by the vehicle's state. Interact with other vehicle statuses and abnormal behavior of the main vehicle Jointly decided. Definition. For the mapping module, the mapping from raw data to natural language description is as follows:
[0093]
[0094] The output of this module is the core mapping in the prompt word, which directly describes the entire interaction process.
[0095] Taking a merging interaction as an example, when the main vehicle exhibits a certain difference in steering angle and simultaneously crosses two lanes, the current state is determined to be a merging interaction. The final output of this module is a prompt that combines natural language statements with key data elements.
[0096] "...The lead vehicle maintained a speed of approximately 28 km / h in the right lane; after 3 seconds, the lead vehicle began preparing to merge into the left lane, gradually adjusting its heading angle from the initial 180° to 172°, with lateral acceleration stabilizing at approximately 0.35 m / s², while longitudinal acceleration increased to 0.9 m / s², and speed gradually increased to 32 km / h. During this period, the relative speed with the vehicle in front on the left was less than 1 km / h, and the PET (Penetration, Acceleration, and Threat) was maintained at 1.8 seconds without triggering a safety warning; after successfully crossing the lane line and merging into the left lane, the lead vehicle maintained a safe following distance of approximately 5.6 meters from the vehicle in front, cruised at 32 km / h for 5 seconds, the heading angle returned to 180° and remained stable, the lateral speed was controlled at 0.8 m / s without any sudden changes, and the speed of the vehicle in front on the left remained at approximately 33 km / h without any sudden acceleration or deceleration; subsequently, the lead vehicle's speed slowly decreased to 30 km / h, and the longitudinal acceleration was maintained between -0.25 and -0.10." The speed was between m / s², and the TTC with the car in front on the left was stable at 2.8 seconds. The interaction between the two vehicles was smooth throughout the merging process, and it did not have a significant impact on the surrounding traffic flow...
[0097] S32 Large Language Model Knowledge Enhancement;
[0098] In this embodiment, GPT-4o is selected as the large language model to complete this step.
[0099] First, an external knowledge base for autonomous driving is constructed, mainly including traffic rules, vehicle behavior norms, and autonomous driving-related knowledge, and then converted into a JSON format that is easy to search.
[0100] Secondly, high-quality interaction data from the dataset is selected, and DLMM is used to map this data into natural language prompts. These prompts are then given to human experts for evaluation. This invention does not require experts to give a final score, but rather summarizes and evaluates the interaction process, providing a fuzzy evaluation. This avoids the influence of subjective differences on the evaluation results.
[0101] Finally, the summaries and evaluations of human experts are incorporated into the external knowledge base for autonomous driving, serving as an information source for subsequent retrieval by the large language model.
[0102] S4 utilizes neural networks to construct a quantitative evaluation model for intelligent autonomous driving;
[0103] First, the natural language interaction segment processed by S31 is input into the large language model, and the large language model is required to search the external knowledge base and evaluate the interaction segment according to the quantitative evaluation index in S2, and give the intelligence evaluation result.
[0104] Secondly, the intelligence evaluation result is normalized, and its value range is scaled to between 0 and 1.
[0105] Subsequently, the quantitative evaluation indicators in step S2 are standardized and mapped to the range of 0 to 1 using a linear transformation matrix, and they are positively correlated with the system performance.
[0106] The processed data is used for neural network training, including the above intelligence evaluation results (1 output value as a label) and evaluation metrics (10 input features).
[0107] The 10 input features are the 10 specific quantitative evaluation indicators in the five-dimensional evaluation system of S2, including: collision time TTC, post-intrusion time PET, and longitudinal acceleration. lateral acceleration lateral speed Yaw angle, angular velocity yr, mission time Average delay (AvgDelay), interoperability index (IO), and impact on the traffic system. Standardize to the [0,1] interval.
[0108] The output value (label) is the intelligence evaluation result, standardized to the [0,1] interval.
[0109] The neural network uses a Multilayer Perceptron (MLP) as the evaluation model. The MLP network structure has an input layer with 10 neurons, each corresponding to a metric; and an output layer with 1 neuron. The number of hidden layers is set to 4, and the number of neurons in each layer is determined using an empirical formula.
[0110]
[0111] in, The number of neurons per layer; This represents the number of training samples; and These represent the number of neurons in the input and output layers, respectively. It is a constant between 2 and 10, and in this embodiment, it is taken as... .
[0112] The total number of samples was 5200. 20% of the data (1040 samples) was taken as the validation set, and the remaining 4160 samples were used as the training set. The number of neurons in the first two layers was calculated to be approximately 38, which was ultimately set to 40, a multiple of the number of input neurons. The number of neurons in the last two layers was set to 20 and 10, respectively.
[0113] The activation functions for the hidden layer and the output layer are the ReLU function and the Sigmoid function, respectively, and the other hyperparameters are set to batch_size=128 and epoch=125.
[0114] The output of this step is a fully trained MLP network, i.e., an autonomous driving intelligence evaluation model. For subsequent autonomous driving experimental data, it is first converted into the 10 evaluation indicators of the five-dimensional evaluation system in S2. Then, these evaluation indicators are provided as input features to the trained MLP network to obtain the intelligence output. The entire prediction process is integrated into a prediction module; inputting experimental data automatically completes the conversion and prediction, yielding the intelligence result, effectively improving the efficiency of autonomous driving intelligence evaluation.
[0115] Model performance evaluation
[0116] Reliability verification
[0117] The Turing test verifies whether the evaluation results of a model can approximate the judgment of human experts. In this invention, it is used to verify the reliability of the results obtained by using a Large Language Model (LLM) for evaluation. This invention selects 20 additional test segments and divides them into two groups. One group is evaluated by the LLM, and the other group is evaluated by human experts (10 doctoral students specializing in this field). The requirement is to produce both the evaluation logic and the results.
[0118] Subsequently, an expert in the field of autonomous driving was invited as a third-party evaluator to determine the source of the above assessment output. Options included human input, LLM (Leadership Management Model), and no determination possible. The Turing test results are as follows: Figure 2 As shown, where and The percentages of human expert assessments that were judged as "unresolved" and "LLM assessment" are respectively.
[0119] The overall pass rate for the Turing test is generally higher than 0.8. , The average pass rate reached 86.75%. This indicates that the model can, to a certain extent, replace human experts in evaluating the intelligence of autonomous driving, improving efficiency while ensuring the accuracy of the evaluation.
[0120] Validity verification
[0121] To verify the effectiveness of this invention, combinations of five typical autonomous driving decision planning and control algorithms (integrated planning and control, second-order programming + MPC, reinforcement learning + PID, IDM model + PID, and grid planner + MPC) were selected and tested in six typical traffic scenarios (conventional urban roads, urban roundabouts, urban intersections, highway ramps, expressways, and parking lots). The test data were evaluated using the method of this invention, and expert evaluation and an evaluation method based on AHP (Analytic Hierarchy Process) were used as comparisons. In the expert evaluation method, the testing process was presented to four experts in video format, supplemented by test data and visualization results. Experts were required to give a score of 0-10 and provide their reasons. In the AHP evaluation method, experts first calibrated the weights of the indicators, and then calculated the evaluation score based on the weights and indicators. The experts involved in the evaluation and calibration in this verification process were three experts from the National Intelligent Vehicle Integrated System Test Zone and one associate professor in the field of autonomous driving.
[0122] Validation results are as follows Figure 3 As shown, the model of this invention scores most algorithms around 0.8, but there are still cases where the scores are too high or too low. For example, Algorithm combination 1 scores 0.82 in scenario 1 and 0.88 in scenario 3, with an overall average score of about 0.85; while Algorithm combination 2 scores 0.81 in scenario 1 and 0.56 in scenario 3, with an average score of about 0.73. This difference in scores reflects the different characteristics of the algorithms: the ensemble planning and control of Algorithm combination 1 is more stable in multiple scenarios, while the second-order programming + MPC of Algorithm combination 2 has weaker adaptability in some complex scenarios.
[0123] Compared with expert evaluation methods, the overall ranking trend of the evaluation results of this invention is basically consistent with expert judgment, especially the rankings of algorithms 1-3, which are identical to expert judgments. This indicates that the model can effectively capture the scoring logic of experts. Compared with the AHP evaluation method, there are significant differences in the ranking of algorithm scores, especially between algorithms 1 and 3, and between 2 and 5, where rankings are interchanged. This is because AHP, by using linear weighting, excessively amplifies the influence of single indicators and is insufficient to accurately describe the complex relationships between indicators. The ability to capture non-linear dependencies between indicators is precisely the advantage of this invention.
[0124] The model of this invention can fit the nonlinear relationship between various indicators, accurately reflect the real performance of the algorithm in actual traffic scenarios, and conform to the subjective decision-making logic of experts based on multi-dimensional interaction. Moreover, it is more efficient than expert evaluation and more in line with the complex needs of actual scenarios than AHP evaluation.
[0125] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.
Claims
1. A method for evaluating the intelligence of autonomous vehicles based on the subjective-objective mapping of a large language model, characterized in that, Includes the following steps: S1. Prepare the dataset; First, 1,000 data points from real-world driving scenarios were selected from a public dataset as the training set for the autonomous driving intelligent metric evaluation model. Then, the data was filtered based on interaction intensity, and data with high interaction intensity was extracted as the data source for knowledge enhancement of the large language model in S3. S2 Determine the evaluation dimensions; A five-dimensional evaluation system for autonomous driving intelligence is constructed, including safety, comfort, efficiency, social interactivity, and impact on the transportation system; specific quantitative evaluation indicators and their quantification methods are determined for each dimension, and these quantitative evaluation indicators will serve as input layer variables for the neural network in S4. S3 uses a large language model to quantitatively evaluate autonomous driving data; First, a data-natural language mapping module is constructed to map the effective interaction fragments corresponding to the high interaction intensity data extracted from S1 into natural language prompts. Then, knowledge from traffic rules, driving regulations, and autonomous driving fields is summarized as an external knowledge base for the large language model. Subsequently, these interaction information described by natural language are handed over to human experts for fuzzy evaluation. Finally, the expert fuzzy evaluation results are incorporated into the external knowledge base, and the prompt word is input into the knowledge-enhanced large model to obtain the intelligence evaluation results. S4 utilizes neural networks to construct a quantitative evaluation model for intelligent autonomous driving; First, the intelligent score output by the large language model and the quantitative evaluation indicators of each dimension determined by S2 are standardized, including the normalization of the intelligent score output by the large language model and the standardization of the quantitative evaluation indicators of each dimension in S2. Then, a multilayer perceptron is used for regression to allow the neural network to capture the nonlinear relationship between the interactive data and the intelligent score.
2. The method for evaluating the intelligence of autonomous vehicles based on subjective-objective mapping of a large language model according to claim 1, characterized in that, Step S1 is as follows: S11 State initialization and trajectory prediction; First, the vehicle state vector is extracted, which includes position, velocity, and acceleration. Then, moving vehicles are identified as the objects for trajectory prediction. Finally, the trajectory within the next 5 seconds is predicted based on the current vehicle state. S12 identifies trajectory intersections; The Shapely geometry library is used to perform two-dimensional geometric modeling of vehicle trajectories, representing vehicle trajectories as linear geometric objects. Based on its provided geometric relationship judgment and intersection operation functions, the geometric intersection points between different vehicle trajectories are calculated, thereby identifying potential trajectory intersection locations. For the identified intersection points, the distance calculation function of the Shapely library is used to calculate the path distance from each vehicle along its own trajectory to the intersection point. Combined with the vehicle's driving speed information, the time for the vehicle to reach the intersection point is estimated, thus obtaining the time difference between different vehicles reaching the intersection point. S13 Determines valid interaction; The system uses a buffer verification method to exclude cases where the vehicle has already passed safely: The system creates a 1.5-meter horizontal buffer zone around the trajectory of the preceding vehicle and verifies whether the trajectory of the following vehicle intersects with the buffer zone; if they intersect and the trajectory of the following vehicle is contained within the buffer polygon, it means that the preceding vehicle has passed safely and is considered an invalid interaction; otherwise, it is considered a valid interaction. S14 quantifies the interaction intensity; The interaction intensity is quantified using the minimum absolute acceleration comprehensive quantization method: the system establishes a constrained optimization model, and the objective function is to minimize the sum of the absolute values of the accelerations of all participating vehicles; By solving the optimization problem, the minimum combined acceleration requirement that enables both vehicles to avoid a collision is obtained, which serves as a quantitative indicator of the interaction intensity. The design constraints include: (1) the time difference between the arrival of the two vehicles at the collision point does not exceed 1.5 seconds; (2) the actual displacement of the vehicles deviates from the expected displacement by no more than 0.1 meters; and (3) the vehicle displacement and final velocity are positive. S15 Extract interactive segments; The system sets an acceleration threshold, identifies time periods where the acceleration continuously exceeds the threshold, and retains only segments with a duration exceeding 0.3 seconds; for each segment, it calculates the maximum interaction intensity and the minimum PET value as key features of the interaction event.
3. The method for evaluating the intelligence of autonomous vehicles based on subjective-objective mapping of a large language model according to claim 1, characterized in that, Step S2 specifically involves constructing a five-dimensional evaluation system, with dimensions including safety, comfort, efficiency, social interactivity, and impact on the transportation system, and quantifying these dimensions to obtain quantitative evaluation indicators. The evaluation indicators and quantification methods for each dimension are as follows: S21 security; The Time-of-Collision (TTC) crash test is used to reflect a vehicle's control over safety, and its quantification method is as follows: in This indicates the distance between the rear of the car in front and the front of the car behind, with the direction from the rear car to the front car. This indicates the relative speed between the two vehicles; the TTC safety threshold is set to 1.5 seconds. Post-Intrusion Time (PET) is another safety assessment metric. Assuming vehicles A and B are about to collide at an intersection, and vehicle A passes the collision point before vehicle B, the PET is calculated as follows: in It is the displacement of car A from the current moment until its body has just completely left the point of conflict. It is the displacement of car B from the current moment until any point on its body just enters the point of conflict. and The instantaneous speeds of the two vehicles; S22 Comfort; The indicators include: longitudinal acceleration lateral acceleration lateral speed Yaw angle and angular velocity yr; and determine the threshold values of vehicle state parameters under different comfort conditions; S23 Efficiency; Metrics include: TaskTime and Average Delay (AvgDelay); TaskTime is defined as the theoretical shortest task time minus the actual task time. The ratio, while the theoretical shortest task time is determined by the minimum task path. Speed limits on road sections Calculated from: Define the average delay (AvgDelay) metric to measure efficiency as follows: in The average speed of the main vehicle during this mission; and These are the minimum and maximum average vehicle speeds among surrounding traffic vehicles, respectively. S24 Social interactivity; Autonomous vehicles should make reasonable interaction decisions based on right-of-way and the status of other vehicles. The I / O metric is extended to general scenarios to evaluate the rationality of autonomous vehicle decisions. The I / O calculation method is as follows: ITSI stands for Interactive Transient State Index. This represents the normalized time difference of a vehicle passing through the conflict zone. It is the normalized cooperative acceleration of the vehicle; the softmax function can compare... The size of the value maps larger values to near 1 and smaller values to near 0, making the ITSI close to the larger of the two, thus highlighting the factors with higher risk coefficients. This represents the normalized vehicle motion behavior trajectory characteristics. This represents the longitudinal displacement of the main vehicle at the current moment. and The minimum and maximum longitudinal displacements that the vehicle can achieve under collision boundary acceleration; The impact of S25 on the transportation system; The impact of an autonomous driving system on a traffic system is defined as the degree to which the driving efficiency of other vehicles that directly or indirectly interact with the autonomous vehicle decreases after the autonomous vehicle enters the traffic system, as shown in the following formula: in, The number of vehicles in the background. and These are the first and last times before and after an autonomous vehicle enters the transportation system. The speed of the affected background vehicle; In summary, 10 quantitative evaluation indicators were obtained: time to collision (TTC), time to intrusion (PET), and longitudinal acceleration. lateral acceleration lateral speed Yaw angle, angular velocity yr, mission time Average delay (AvgDelay), interoperability index (IO), and impact on the traffic system. .
4. The method for evaluating the intelligence of autonomous vehicles based on subjective-objective mapping of a large language model according to claim 1, characterized in that, Step S3 is as follows: S31 Data-Natural Language Mapping Module; A Data-Natural Language Mapping (DLMM) module was built between the raw interaction data and the large language model to convert the autonomous driving data into a form that is easy for the large language model to understand, and finally add it to the prompt word. The state of an autonomous driving system during driving is represented by a state-space model: in, For the vehicles that interact, and For vehicle coordinates, and The longitudinal and lateral velocities and accelerations of the vehicle. For vehicle heading angle; to identify the occurrence of interactive events, define Define the effective interaction range; based on different interaction types, define the effective interaction ranges for following, merging, and crossing. , and as follows: in, This is the lane vector of the interacting vehicle, with each element being the lane ID. The angle difference threshold, To maintain following distance, and The following distance threshold and conflict distance threshold. This represents the lane vector of the vehicle that changed lanes. The conjunction symbol, The number of points of conflict in the future trajectories of the two vehicles. For vehicles Distance to the point of conflict; function Used to detect abnormal behavior of the main vehicle: in, These are vehicle state quantities, including speed, acceleration, and heading angle. This is the abnormal threshold. This is an indicator function; when the rate of change of a certain state variable exceeds a threshold, it will... Identify and label abnormal moments; definition The vehicle's behavior is described in natural language, which is determined by the vehicle's state. Interact with other vehicle statuses and abnormal behavior of the main vehicle Jointly decided; defined For the mapping module, the mapping from raw data to natural language description is as follows: The output of this module is the core mapping in the prompt word, which directly describes the entire interaction process; S32 Large Language Model Knowledge Enhancement; First, build an external knowledge base for autonomous driving, including traffic rules, vehicle behavior norms, and autonomous driving-related knowledge, and convert it into a JSON format that is easy to search. Secondly, high-quality interaction data from the dataset were selected, and DLMM was used to perform natural language mapping on these data and convert them into prompts. The prompts were then given to human experts for evaluation. Instead of requiring experts to give a final score, the experts were asked to summarize and evaluate the interaction process and give a fuzzy evaluation. Finally, the summaries and evaluations of human experts are incorporated into the external knowledge base for autonomous driving, serving as an information source for subsequent retrieval by the large language model.
5. The method for evaluating the intelligence of autonomous vehicles based on subjective-objective mapping of a large language model according to claim 1, characterized in that, Step S4 is as follows: First, the natural language interaction segment processed by S3 is input into the large language model, and the large language model is required to search the external knowledge base. The interaction segment is evaluated according to the quantitative evaluation index in S2, and the intelligence evaluation result is given. Secondly, the intelligence evaluation result is normalized, and its value range is scaled to between 0 and 1. Subsequently, the quantitative evaluation indicators in step S2 are standardized and mapped to the interval of 0 to 1 using a linear transformation matrix. The processed data was used to train the neural network. The input to the neural network consisted of 10 specific quantitative evaluation indicators from the five-dimensional evaluation system in S2 after standardization, including: collision time TTC, post-intrusion time PET, and longitudinal acceleration. lateral acceleration lateral speed Yaw angle, angular velocity yr, mission time Average delay (AvgDelay), interoperability index (IO), and impact on the traffic system. ; The output value of the neural network is the intelligence evaluation result, which is standardized to the [0,1] interval; The neural network selects a multilayer perceptron (MLP) as the evaluation model; the input layer of the MLP network structure contains 10 neurons, each neuron corresponding to an index; the output layer contains 1 neuron. The trained MLP network is the autonomous driving intelligence evaluation model; For subsequent autonomous driving experiment data, it is first converted into 10 evaluation indicators of the five-dimensional evaluation system in S2, and then the evaluation indicators are used as input features to provide to the trained MLP network to obtain the output of intelligence.
6. The method for evaluating the intelligence of autonomous vehicles based on subjective-objective mapping of a large language model according to claim 5, characterized in that, The number of hidden layers in the neural network is set to 4, and the number of neurons in each layer is set using an empirical formula: in, The number of neurons per layer; This represents the number of training samples; and These represent the number of neurons in the input and output layers, respectively. A constant between 2 and 10; The activation functions for the hidden layer and the output layer are the ReLU function and the Sigmoid function, respectively, and the other hyperparameters are set to batch_size=128 and epoch=125.
Citation Information
Patent Citations
Assessment method and device of large language model and computer equipment
CN118535443A
Automatic driving evaluation method and device based on large language model
CN120297546A