Large-model-driven intelligent automobile test scene human-like evaluation method
Through the large-model-driven intelligent car test scenario human-like evaluation method, combined with multi-condition data acquisition and physiological signal prediction, multi-dimensional human-like evaluation is realized, solving the problem that existing methods are difficult to fully reflect multi-dimensional information, and reducing the acquisition cost.
Patent Information
- Application Number
- CN202510630192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing intelligent vehicle test scenario evaluation methods are difficult to fully reflect multi-dimensional nonlinear information, and the physiological data acquisition cost is high and the efficiency is low.
A large-model-driven method is adopted to realize multi-dimensional human-like evaluation of test scenarios through multi-condition data acquisition, physiological signal prediction, evaluation word-element matching and multi-word element induction.
It improves evaluation accuracy, reduces data acquisition and experiment costs, and realizes low-cost and high-precision quantitative evaluation of test scenarios.
Smart Images

Figure CN120145204A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous vehicle testing, and specifically relates to a human-like evaluation method for intelligent vehicle test scenarios driven by a large model. Background Art
[0002] Testing and evaluation technologies are the foundation and prerequisite for the industrialization of intelligent vehicles. Compared with traditional mileage-based testing methods, scenario-based testing methods abstract the real world into fragmented test scenarios and achieve high-efficiency world description and test execution through parametric design, which has become the core means for evaluating the performance of intelligent vehicles. Among them, key scenarios can expose potential defects of the system under test and have extremely high testing and application value. How to evaluate key scenarios is a hot issue in existing research and the research basis for subsequent scenario generation and automated testing.
[0003] The research on the criticality evaluation of intelligent vehicle test scenarios is currently mainly divided into two aspects: objective data evaluation and human evaluation. In terms of objective data evaluation, the objective threshold evaluation method based on vehicle dynamics was first proposed and widely applied. To further enhance the interpretability of the evaluation process, relevant scholars introduced the concept of risk potential field to describe the visual potential risk distribution of vehicles in complex traffic environments. The ultimate goal of the development of intelligent vehicles is to become an intelligent tool that comprehensively serves humans, and its commercial attributes determine that market acceptance highly depends on the actual user experience. Therefore, it is particularly important to establish a comprehensive evaluation system that can reflect human-like characteristics. In terms of human evaluation, the basic scheme based on questionnaires or scales has been widely applied, providing important references for the performance improvement of intelligent vehicles by collecting direct feedback from users on driving experiences. In addition, in the field of human-machine co-driving, relevant scholars design decision-making strategies by considering the physiological signals of occupants to better meet human driving needs. However, there are still some deficiencies in existing research: existing subjective evaluation methods only evaluate from the dimension of deterministic numerical values, and the obtained subjective scale values are difficult to comprehensively reflect multi-dimensional non-linear information in the testing process; physiological data can reflect human psychological characteristics, but high-precision physiological data acquisition devices, professional evaluation personnel, and comprehensive test sites all bring significant cost and efficiency problems to the evaluation process.
[0004] Therefore, there is an urgent need for an evaluation system that comprehensively considers multi-dimensional information, finds an effective alternative method for physiological data, and realizes low-cost and high-precision quantitative evaluation of key test scenarios. Summary of the Invention
[0005] To solve the above problems, the present invention provides a human-like evaluation method for intelligent vehicle test scenarios driven by a large model. Through multi-condition data collection, physiological signal prediction, evaluation token matching, and multi-token induction and summary, it finally realizes the multi-dimensional human-like evaluation output of the key points of the test scenario, which can effectively improve the evaluation accuracy and reduce the data collection and experimental costs in practical applications.
[0006] The technical solution of the present invention is described in conjunction with the accompanying drawings as follows: The present invention provides a human-like evaluation method for intelligent vehicle test scenarios driven by a large model, including the following steps: S1. Data collection and processing; Design the collection conditions, build the collection equipment, collect data at the test site, and perform preliminary data processing; S2. Physiological signal prediction; Conduct correlation analysis on the data, perform hierarchical processing in combination with data characteristics, design the architecture of the physiological signal prediction model, and perform model training; S3. Evaluation token matching; Process the evaluation text data, construct a token map, perform feature processing on the time-series data, design the architecture of the token matching model, and perform model training; S4. Evaluation summary generation; Select a large language model, construct a text knowledge base, develop a multi-reflection mechanism, and construct a token summary evaluation model.
[0007] Furthermore, the specific method of S1 is as follows: S11. Collection condition design; S12. Construction of collection equipment and on-site collection; S13. Data processing.
[0008] Furthermore, the specific method of S11 is as follows: S111. Based on the test standard procedures of SAE J2944, ISO 7401, and SAE J266, design collection conditions with prominent single-dimensional indicators and coupled multi-dimensional indicators, that is, the single and coupled aspects of horizontal, vertical, and longitudinal directions; S112. The designed conditions cover the longitudinal, horizontal, and vertical movements in vehicle dynamics; the test parameters follow the ISO8855 and ISO 10844 standard documents; The specific method of S12 is as follows: S121. Use the BioNomadix wireless physiological recorder to synchronously collect the driver's electromyogram, electrocardiogram, and galvanic skin response signals through a three-lead DryPad electrode, and set the sampling rate to 2 kHz; collect the longitudinal / lateral acceleration of the chassis at a frequency of 100 Hz through the CAN FD bus protocol; obtain centimeter-level pose accuracy information based on the ASENSING GNSS / RTK combined positioning system, IMU inertial measurement unit, and DTU data transmission unit; transmit the chassis signals and GNSS pose data of the vehicle to the SpeedGoat real-time processing platform through the CAN bus protocol; send 5V TTL trigger pulses based on the timing trigger module to achieve multi-device microsecond-level synchronization; transmit all data to a laptop computer and save it in a storage medium through MATLAB software and physiological signal processing software; S1222. The experimental site is a dedicated vehicle test site, and the driver is a professional test driver selected from the test site; a 5-minute rest period is carried out before the start of each acquisition condition to enable the driver to enter the test session in a calm state to collect baseline physiological signals, and 3 pre-experiments are carried out to calibrate the driving to eliminate equipment interference; the evaluation data includes the statement output of real-time during acquisition and a complete summary after acquisition; The specific method of S13 is as follows: S131. Use the rising edge of the hardware trigger signal as the time reference and implement 100Hz unified resampling through the cubic spline interpolation algorithm; S132. Establish a multi-level filtering processing flow: use Kalman filtering to suppress high-frequency noise for dynamic signals, and implement wavelet threshold denoising combined with an adaptive filtering algorithm for physiological signals; Among them, the state equation and observation equation in the Kalman filter state space model are respectively: ; ; In the formula, is the state transition matrix; is the process noise; is the observation matrix; is the observation noise; is the control input matrix, which maps the control vector to the change of the state variable; is the control vector, indicating the known external input applied to the system at time ; S133. Physiological data includes real-time electrical signals and key physiological features; in terms of physiological feature extraction, physiological feature processing is carried out based on the NeuroKit2 physiological signal analysis framework; (1) for the ECG signal, extract the clear signal after denoising ; calculate the heart rate , reflecting the body's stress level or exercise state; evaluating the quality indicators of ECG data ; and extracting the peak positions of QRS complex waves, P waves, and T waves , and , describing the electrical activity phases of the atria and ventricles throughout the cardiac cycle for detecting cardiac rhythm abnormalities; (2) For the EDA signal, first remove the noise to obtain a smoothed EDA signal ; then separate the slow-varying component , reflecting the tone of the autonomic nervous system; capture the rapidly changing component , related to instantaneous emotional responses; quantify the process of instantaneous conductance increase triggered by stimuli, including the start time , peak time , height , amplitude , rise time , recovery time , used to evaluate an individual's response sensitivity to specific events; (3) For the EMG signal, extract muscle activity features after filtering and envelope detection, including the processed EMG signal ; the intensity of muscle contraction , duration , and the starting point of muscle activity and the ending point .
[0009] Furthermore, the specific method of S2 is as follows: S21. Data correlation analysis; S22. Data feature hierarchical processing; S23. Prediction model architecture design and training.
[0010] Furthermore, the specific method of S21 is as follows: S211. A feature screening framework based on time-delay causal inference, through fusing cross-correlation analysis and Granger causality test, realizes an interpretable mapping model of kinetic features to physiological responses; S212. Use a sliding window cross-correlation algorithm to quantify the dynamic association strength. Let the chassis data be , the physiological data be , and the cross-correlation coefficient with a lag of is defined as: ; By traversing the lag time window of , extract the maximum correlation coefficient and the corresponding lag time , characterizes the optimal alignment of the two signals; through related studies such as human nerve conduction delay and visual persistence, the time Set to 300ms; S213, construct a bivariate vector autoregression model, assuming that the chassis data Physiological data There is a causal effect, and a two-order regression equation is established: ; The residual sum of squares of the constrained model and the unconstrained model are compared by the F test. If the statistic satisfies: ; Then reject the null hypothesis and judge for Granger causes; ADF test is used to verify the stationarity of time series, and first-order difference or logarithmic transformation is performed on non-stationary data; S214, based on the dual threshold mechanism, two types of analysis results are integrated: 1) in the cross-correlation analysis, the feature data with cross-correlation coefficient and lag time within the lag window are retained; 2) in the Granger causality test, the features with only statistical correlation but no causal explanatory power are eliminated; the final selected data are: ECG_rate, ECG_quality, EDA_clean, EDA_phasic, EMG_Amplitude, EMG_Activity; The specific method of S22 is as follows: The initial state of each working condition and peak state As a benchmark, calculate the data extreme value in the time series window The relative change percentage of the data is calculated, and the data are classified according to the relative change percentage; The specific method of S23 is as follows: S231. Design a cascade prediction method based on temporal graph attention network and Transformer model; it includes four parts: graph network conversion, temporal feature compression, graph attention extraction and Transformer prediction; S232. First, construct a fully-connected graph structure to represent the relationships between sensors or measurement points, where each data point of chassis dynamics is regarded as a node in the graph. Secondly, divide the entire time series into time windows of a fixed length and apply graph convolution operations at each time step. Apply noise reduction and smoothing processing within each window. Introduce a graph convolution layer with an attention mechanism to assign learnable attention coefficients to the edges between each pair of nodes, dynamically learning the importance weights between nodes. Finally, the processed compact and semantically rich feature representation is fed into a Transformer-based encoder for further processing. The Transformer uses a self-attention mechanism. Ultimately, the prediction results for each time step are output through a fully-connected layer, completing the end-to-end mapping from input to output.
[0011] Furthermore, the specific method of S3 is as follows: S31. Construction of the token graph framework; S32. Construction of the dataset; S33. Model design and training.
[0012] Furthermore, the specific method of S31 is as follows: Construct a multi-dimensional driving evaluation system based on knowledge graph technology; synchronously record the driver's real-time voice evaluation and final summary feedback under working conditions through an in-vehicle data acquisition system, and combine vehicle dynamics parameters and physiological signals to form an interpretable semantic mapping network; through the induction and standardization processing of evaluation texts, finally form a classification system, including two parts: physiological feelings and chassis responses; physiological feelings include two aspects: riding experience and physical feelings; chassis responses include four parts: power response, steering characteristics, braking performance, and suspension feedback; The specific method of S32 is as follows: S321. Design special prompt words; S322. During the acquisition process, require the driver to output real-time evaluations at each time node and give a final summary after the test; all the collected data need to be time-synchronized to ensure that each record has a unified timestamp; subsequently, perform NLP preprocessing steps on the original text data and map the driver's evaluation data to the corresponding time dimension; use the global time to accurately match the evaluation tokens and time series data; adopt the sliding window method to construct a continuous "prompt word - chassis dynamics - physiological data - evaluation token" combined dataset; the data within each window represents the complete information within a period of time. S323. Introduce a feature enhancement strategy. First, extract statistical features, calculate the mean, variance, maximum, and minimum of the data within the window. Then, calculate the change trend of the data through the moving average method. The specific process is as follows: For the determined normalized multi-dimensional time series data, use 30%, 70%, and 100% of the data length as data breakpoints, and calculate the average SMA of the data in each stage respectively. Consider the SMA change greater than 0.1 as rising, less than -0.1 as falling, and between as stable. Thus, convert the time series data into a semantic description of the change trend, including 9 types: rising, falling, stable, rising then falling, rising then stable, falling then rising, falling then stable, stable then rising, and stable then falling. The specific method of S33 is as follows: S331. First, perform time series feature compression. Compress the multi-dimensional time series data into a 256-dimensional feature vector through a 1D convolutional layer and a global max pooling layer. Then, jointly input it with the structured prompt text into the BERT-base model to generate a 768-dimensional semantic representation. Add 6 independent classification heads after the BERT token output layer, corresponding to body feeling, emotional response, dynamic response, steering characteristics, braking performance, and suspension feedback respectively. Each classification head adds a fully connected layer on top of the last hidden state of BERT and applies the softmax activation function to generate the probability distribution of each category. During the training process, use the cross-entropy loss function as the optimization objective. After the input data is encoded by BERT, it will be passed to their respective classification heads, and then generate the prediction probability of each category. The prediction probability value is used to calculate the loss and guide the gradient update in the backpropagation process. S332. Design a weighted loss function. Based on the occurrence frequency of each category, assign weights to each category. The weights are calculated according to the occurrence frequency of each category in the dataset. The calculation method is as follows: ; In the formula, is the total number of samples in the th subclass; is the number of token categories under this subclass; is the number of times the th token appears under the subclass; S333. Define the AdamW optimizer and adopt a weight decay strategy to avoid overfitting. Introduce a learning rate scheduler with linear warm-up and cosine annealing to help the model converge more stably. Set a maximum gradient norm limit to prevent gradient explosion problems. In each epoch, iterate through the training data loader, batch process the input data and calculate the loss value. Apply gradient accumulation techniques to improve the effect of mini-batch training and regularly evaluate the performance on the validation set. If there is no performance improvement in 3 consecutive batches, that is, the evaluation accuracy no longer rises, then terminate early. Through the design and training process, the BERT model transforms kinetic parameters and physiological signals into token outputs that conform to human cognition through feature fusion and semantic mapping.
[0013] Further, the specific method of S4 is as follows: S41. Select a large language model; S42. Develop knowledge base retrieval enhancement; S43. Design a multi-reflection mechanism.
[0014] Further, the specific method of S41 is as follows: Select the Qwen model as the core LLM model for summarizing evaluation results; The specific method of S42 is as follows: S421. Introduce the RAG retrieval enhancement technology. By combining an external knowledge base and real-time retrieval, help the model obtain more background information and improve the depth of summarization; Collect text materials including the following parts: ① Theoretical literature and technical reports, including vehicle performance, driving experience, and autonomous driving test evaluation; ② A systematic evaluation database, collecting structured driver evaluation data, covering detailed feedback in different driving scenarios; ③ Domain expertise, covering training materials and tutorials, providing professional guidance; ④ Historical test records, detailed reports and representative cases of previous similar tests, used as reference templates; S422. Perform targeted processing on the collected texts in various formats. In the text processing stage, remove irrelevant characters, punctuation marks, and special formats. In the text segmentation stage, set the segmentation block size limit to 200 characters and the number of shared characters between two blocks to 20 to maintain context coherence. Vectorize the segmented texts. Establish a knowledge vector library based on Chroma. For the input question, use the maximum marginal relevance-based similar text evaluation method to search the vectorized texts, increase the diversity of retrieval results, and avoid duplicate information. Combine the retrieved relevant texts with the query prompt template to form a complete prompt, and input it into the language model to organize and output the query result; The specific method of S43 is as follows: S431. Based on the Self-Refine mechanism in the LLM field, design the Multi-Self-Refine mechanism to optimize the output of the LLM. The mechanism includes a summary module, a context awareness module, a short-term memory module, and a long-term memory module. S432. After the input of the token group of driving experience, physical feelings, and emotional responses, first extract keywords for RAG retrieval. Each summary module will generate a test report summary through the prompting words of guided summary and RAG results and immediately store it in the short-term memory module. After one round of retrieval, the context awareness module will receive the output, score the output results from two aspects of semantic coherence and result accuracy, and give specific improvement suggestions. Finally, it is judged whether the execution is completed according to three indicators: "whether the specified generation times are reached", "whether the predetermined quality standard is reached", and "whether the quality change rate of two consecutive rounds is lower than the set value". If the execution is not completed, the summary module is required to adjust according to the suggestions. If the execution is completed, the final evaluation report is feedback and output.
[0015] The beneficial effects of the present invention are as follows: (1) Through multi-condition data collection, physiological signal prediction, evaluation token matching, and multi-token induction and summary, the present invention finally realizes multi-dimensional human-like evaluation output of the key points of the test scenario, which can effectively improve the evaluation accuracy and reduce the data collection and experimental costs in practical applications.
[0016] (2) The present invention designs and proposes a physiological signal prediction method based on temporal GCN. By correlation analysis to screen typical physiological characteristics, through temporal data segmentation, graph structure conversion, and graph attention extraction, the prediction accuracy of the chassis-physiology is significantly improved.
[0017] (3) The present invention designs an evaluation token matching method based on the large language model and designs a multi-classification application method of the BERT model to map complex and disordered physiological and chassis data to highly condensed human evaluation tokens.
[0018] (4) The present invention designs an evaluation token summary method based on the large language model, and realizes the adaptive summary of the evaluation results through knowledge base enhanced retrieval and multi-reflection mechanism. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a flow diagram of the present invention; Figure 2 Schematic diagram of the designed working conditions; Figure 3 Schematic diagram of the built multi-modal heterogeneous data acquisition system; Figure 4 Schematic diagram of the extracted feature data; Figure 5 Schematic diagram of the data classification rules; Figure 6 Schematic diagram of the lemmatization evaluation system; Figure 7 Schematic diagram of the RAG retrieval enhancement technology; Figure 8 Schematic diagram of the evaluation report summary technology. Specific implementation manners
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only parts related to the present invention are shown in the accompanying drawings, rather than all structures.
[0022] Embodiment 1: Refer to Figure 1 , this embodiment provides a human-like evaluation method for intelligent vehicle test scenarios driven by a large model, including the following steps: S1. Data acquisition and processing; Design the acquisition working conditions, build the acquisition equipment, collect data at the test site, and perform preliminary data processing, specifically as follows: S11. Design of the acquisition working conditions; When the driver is in the same initial state, different responses of vehicle dynamics will have different degrees of influence on the driver, which will in turn lead to differences in objective physiological signals and driving feelings. In order to accurately collect the response of the driver's physiological signals and driving feelings under different working conditions, from the perspective of dynamic deconstruction, based on various test standard procedures such as SAE J2944, ISO 7401, and SAE J266, the present invention designs acquisition working conditions with prominent single-dimensional indicators and coupled multi-dimensional indicators.
[0023] S112. The designed working conditions are as Figure 2 shown. The designed working conditions cover the longitudinal, lateral, and vertical motion aspects in vehicle dynamics. The test parameters strictly follow standard documents such as ISO 8855 and ISO 10844. By setting dynamic excitations covering the ISO standard spectrum, the vehicle state migration process from normal driving to critical instability can be completely characterized.
[0024] S12. Construction of the acquisition equipment and on-site acquisition; S121. The multi-modal heterogeneous data acquisition system established in the present invention is as Figure 3 shown. The BioNomadix wireless physiological recorder is used to synchronously collect the driver's electromyogram (EMG), electrocardiogram (ECG), and electrodermal activity (EDA) signals through a three-lead DryPad electrode. The sampling rate is set to 2 kHz (meeting the requirements of the human vibration perception frequency band in ISO 2631-1); various dynamic parameters such as chassis longitudinal / transverse acceleration are collected at a frequency of 100 Hz through the CAN FD bus protocol; centimeter-level pose accuracy information is obtained based on the ASENSING GNSS / RTK integrated positioning system, IMU inertial measurement unit, and DTU data transmission unit; the chassis signals of the vehicle and the GNSS pose data are transmitted to the SpeedGoat real-time processing platform through the CAN bus protocol; a 5V TTL trigger pulse is sent based on the timing trigger module to achieve microsecond-level synchronization of multiple devices; all data is transmitted to a laptop computer and saved in a storage medium through MATLAB software and dedicated physiological signal processing software.
[0025] S122. The experimental site can be the China Hainan Automobile Test Center. The drivers are professional test drivers at the test center who have conducted tests on vehicle performance such as chassis data calibration, new vehicle test drives, and handling stability tests many times and have rich driving experience and scoring experience. To ensure the stability of physiological indicators during data collection, a 5-minute rest period is carried out before each collection condition, enabling the driver to enter the test session in a calm state to collect baseline physiological signals, and three pre-experiments are conducted to calibrate the driving to eliminate equipment interference. The evaluation data includes real-time statement output during collection and a complete summary after collection, which can well assist the subsequent token output work of this research.
[0026] S13. Data processing; S131. To address the timing misalignment problem caused by the difference in sampling frequencies of multi-source data, the rising edge of the hardware trigger signal is used as the time reference, and a cubic spline interpolation algorithm is adopted to achieve unified resampling at 100 Hz, effectively eliminating the time deviation between devices.
[0027] S132. To solve the problem of signal distortion caused by electromagnetic interference and baseline drift, a multi-stage filtering processing flow is established: Kalman filtering is used to suppress high-frequency noise for dynamic signals, and wavelet threshold denoising combined with an adaptive filtering algorithm is implemented for physiological signals to eliminate motion artifacts while retaining effective features.
[0028] Among them, the state equation and observation equation in the Kalman filter state space model are respectively: ; ; Wherein, is the state transition matrix; is the process noise; is the observation matrix; is the observation noise; is the control input matrix, which maps the control vector to the change of the state variable; is the control vector, indicating the known external input applied to the system at time .
[0029] S133. Physiological data not only contains real-time electrical signals, but also contains a large number of in-depth key physiological characteristics. In terms of physiological feature extraction, physiological features are processed based on the NeuroKit2 physiological signal analysis framework. (1) For ECG signals, extract the clear signal after denoising ; calculate the heart rate , which reflects the body's stress level or exercise state; evaluate the quality index of the ECG data; and extract the peak positions of the QRS complex, P wave and T wave , and , which describe the electrical activity stages of the atrium and ventricle during the entire heartbeat cycle and are used to detect cardiac rhythm abnormalities. (2) For EDA signals, first remove the noise to obtain a smoothed EDA signal ; then separate the slow-varying component , which reflects the tone of the autonomic nervous system; capture the rapidly changing component , which is related to the instantaneous emotional response; quantify the process of instantaneous conductance increase triggered by stimuli, including the start time , the peak time , the height , the amplitude , the rise time , and the recovery time , so as to evaluate the reaction sensitivity of an individual to a specific event; (3) For EMG signals, extract muscle activity characteristics after filtering and envelope detection, including the processed EMG signal ; the intensity of muscle contraction , the duration , and the starting point and the ending point of muscle activity. The extracted feature data is as shown in Figure 4 , which can provide a standardized data basis for subsequent driving state analysis.
[0030] S2. Physiological signal prediction; Conduct a correlation analysis on the data, perform hierarchical processing in combination with data characteristics, design the architecture of the physiological signal prediction model, and conduct model training; specifically as follows: S21. Data correlation analysis; S211. Both the collected chassis dynamics data and physiological signal data are high-dimensional and complex time series with interactive coupling characteristics. Therefore, directly using multi-dimensional dynamics data to predict multi-dimensional physiological data may lead to problems such as overfitting and difficulty in convergence. To solve these problems and improve the effectiveness of the prediction model, this study proposes a feature screening framework based on time-delay causal inference. By integrating cross-correlation analysis and Granger causality test, an interpretable mapping model of dynamics features to physiological responses is realized.
[0031] S212. Aiming at the time-asynchronous characteristics of physiological signals and chassis dynamics data, a sliding window cross-correlation algorithm is used to quantify the dynamic correlation strength. Let the chassis data be , and the physiological data be . Its lag cross-correlation coefficient is defined as: ; By traversing the lag time window of , the maximum correlation coefficient and the corresponding lag time are extracted to characterize the best alignment state of the two signals. Through relevant research such as human nerve conduction delay and persistence of vision, the time is set to 300 ms.
[0032] S213. To distinguish statistical correlation from causal dependence, a bivariate vector autoregressive model is constructed. Assuming that the chassis data has a causal influence on the physiological data , a two-order regression equation is established: ; By comparing the sum of squared residuals of the restricted model (only including the lag terms of Y) and the unrestricted model (including the lag terms of X) through the F test, if the statistic satisfies: ; Then the null hypothesis is rejected, and it is determined that is the Granger cause. And the ADF test (Augmented Dickey-Fuller Test) is used to verify the stationarity of the time series, and first-order difference or logarithmic transformation is performed on non-stationary data to avoid spurious causal inference.
[0033] S214. Fusion of two types of analysis results based on a dual-threshold mechanism: 1) During cross-correlation analysis, retain the feature data with cross-correlation coefficients and lag times within the lag window; 2) During Granger causality testing, set the significance level and eliminate features that only have statistical correlation but no causal explanatory power. Through the dual constraints of time-delay alignment and causal inference, the complex interaction between chassis dynamics data and physiological signals can be more comprehensively understood, breaking through the modeling limitations of traditional Pearson correlation for time-varying systems, ensuring that the selected features are not only statistically significant but also have actual causal relationships, providing a high signal-to-noise ratio input space for subsequent physiological prediction models. The finally selected data are: ECG_rate, ECG_quality, EDA_clean, EDA_phasic, EMG_Amplitude, EMG_Activity.
[0034] S22. Hierarchical processing of data features; During the data acquisition process, it was found that different initial states of human drivers would affect the subsequent physiological signal change levels. At the same time, physiological signals have small-range fluctuation phenomena and hierarchical change characteristics, and random data is difficult to accurately predict. To eliminate the baseline differences of individuals under different working conditions, the present invention uses the initial state and peak state under each working condition as a reference to calculate the relative change percentage ( ) of the data extreme values within the time series window (512 points). The data classification rules are as shown. Figure 5 shown.
[0035] S23. Design and training of the prediction model architecture; S231. During the actual acquisition process, both chassis dynamics data and physiological data are continuous time variables, with high dimensions, many disturbances, and complex spatio-temporal dependence relationships. Traditional methods are difficult to effectively handle such problems. To address this challenge, the present invention proposes a cascaded prediction method based on the Temporal Graph Attention Network (T-GAT) and the Transformer model. It includes four core parts: graph network transformation, temporal feature compression, graph attention extraction, and Transformer prediction.
[0036] S232. First, construct a fully connected graph structure to represent the relationships between sensors or measurement points, and each data point of chassis dynamics is regarded as a node in the graph. Second, divide the entire time series into time windows of fixed length, and apply graph convolutional operations (GCN) at each time step to simplify the original high-dimensional time series data and retain important spatio-temporal dependencies. To further improve the data quality, noise reduction and smoothing processes are applied within each window to reduce the impact of noise interference. Furthermore, a graph convolutional layer with an attention mechanism (GAT) is introduced to assign learnable attention coefficients to the edges between each pair of nodes, dynamically learning the importance weights between nodes, thereby effectively capturing the complex relationships in the graph structure. Finally, the compact and semantic-rich feature representations obtained through the above three stages of processing are fed into a Transformer-based encoder for further processing. The Transformer adopts a self-attention mechanism, which can fully explore the hidden correlations without sacrificing the time order, and combines positional encoding to ensure that the inherent order information of the time series is not lost. Ultimately, the prediction results at each time step are output through a fully connected layer, completing the end-to-end mapping from input to output.
[0037] S3. Evaluate token matching; Process the evaluation text data, construct a token graph, perform feature processing on the time series data, design the architecture of the token matching model, and conduct model training; S31. Construct the token graph framework; S311. To systematically analyze the correlation mechanism between driver evaluations and vehicle objective data, the present invention constructs a multi-dimensional driving evaluation system based on knowledge graph technology. By synchronously recording the real-time voice evaluations and final summary feedback of drivers under typical working conditions through an in-vehicle data acquisition system, and combining vehicle dynamics parameters and physiological signals, an interpretable semantic mapping network is formed. Through the induction and standardization processing of the evaluation text (merging synonyms and eliminating ambiguities), the classification system shown in Figure 6 is finally formed, which includes two parts: physiological feelings and chassis responses. Physiological feelings include two aspects: ride experience and body feelings. Chassis responses include four parts: power response, steering characteristics, braking performance, and suspension feedback. It can be found that the index system has different lexical expressions, which can capture the multi-dimensional feelings of drivers during the test, and is more comprehensive and specific than a single scale score.
[0038] S32. Construct the data set; S321. To construct a structured and consistent real-time evaluation data set, ensure sufficient context information during model training, and provide clear training objectives, this study designed a special prompt: "You are an intelligent vehicle test evaluator. Please evaluate according to the chassis dynamics data { } and physiological index data {}, and describe your driving experience, physical sensations, and emotional reactions using words.
[0039] S322. The scenario data is time-sequential. Therefore, during the data collection process, the driver is required to combine their driving experience and provide real-time evaluation outputs at each time node, and give a final summary after the test. All the collected data (including chassis dynamics signals, physiological signals, and the driver's evaluations) need to be precisely time-synchronized to ensure that each record has a unified timestamp. Subsequently, perform NLP preprocessing steps such as word segmentation and stop word removal on the original text data, and map the driver's evaluation data to the corresponding time dimension. Use the global time to precisely match the evaluation tokens and time-sequential data. Adopt the sliding window method to construct a continuous dataset of "prompt words - chassis dynamics - physiological data - evaluation tokens" combinations. The data within each window represents the complete information within a time period, facilitating the model to learn the relationships between different modal data.
[0040] S323. Predicting time-sequential data, especially text data, is challenging. A larger amount of data input and more dimensional data features have been proven beneficial for improving the prediction effect. Therefore, introduce a feature enhancement strategy. First, extract statistical features and calculate the mean, variance, maximum value, and minimum value of the data within the window. Then, calculate the change trend of the data through the moving average method. The specific process is as follows: For the determined standardized multi-dimensional time-sequential data, use 30%, 70%, and 100% of the data length as data breakpoints, and calculate the average value SMA of each stage of the data respectively. Consider SMA change greater than 0.1 as rising, less than -0.1 as falling, and between them as stable, thereby converting the time-sequential data into a semantic description of the change trend, including 9 types: rising, falling, stable, rising first and then falling, rising first and then stable, falling first and then rising, falling first and then stable, stable first and then rising, stable first and then falling.
[0041] S33. Model design and training; S331. When dealing with large-scale data, due to the performance limitations of the language model, long-time series data cannot be directly input. Therefore, time series feature compression is first performed. Through a 1D convolutional layer and a global max pooling layer, multi-dimensional time series data is compressed into a 256-dimensional feature vector. Subsequently, it is jointly input into the BERT-base model (with the underlying parameters frozen) with structured prompt text to generate a 768-dimensional semantic representation. To adapt to the token prediction task of this study, 6 independent classification heads are added after the BERT token output layer, corresponding to body sensations, emotional responses, dynamic responses, steering characteristics, braking performance, and suspension feedback respectively. Each classification head adds a fully connected layer on top of the last hidden state of BERT and applies the softmax activation function to generate the probability distribution of each category. During the training process, the cross-entropy loss function is used as the optimization objective. After the input data is encoded by BERT, it will be passed to their respective classification heads, thereby generating the prediction probabilities of each category. These probability values are used to calculate the loss and guide the gradient update during the backpropagation process.
[0042] S332. During the actual evaluation process, due to the personal habits of the evaluators and the characteristics of the scenario working conditions, words such as "linear" that reflect the basic functions of the system will appear more frequently than other tokens, and the category imbalance of each sub-item. The model will tend to predict high-frequency categories to reduce the loss, resulting in a decline in the recognition performance of minority categories. For this reason, the present invention designs a weighted loss function, assigns weights to each category based on the occurrence frequencies of various categories, so that the model pays more attention to minority categories. The weights are calculated according to the occurrence frequencies of various categories in the dataset, and the calculation method is as follows: ; In the formula, is the total number of samples of the th subclass, is the number of token categories under this subclass, is the number of times the th token appears under this subclass. Through this method, the gradient of the model for low-frequency classes will be enhanced, thus avoiding the model from biasing towards the majority class.
[0043] S333. Define the AdamW optimizer and adopt a weight decay strategy to avoid overfitting. Introduce a learning rate scheduler with linear warm-up and cosine annealing to help the model converge more stably, and set a maximum gradient norm limit to prevent gradient explosion problems. In each epoch, traverse the training data loader, batch process the input data and calculate the loss value, apply the gradient accumulation technique to improve the effect of mini-batch training, and regularly evaluate the performance on the validation set. If there is no performance improvement for 3 consecutive batches, terminate early to save resources. Through the above design and training process, the BERT model can transform dynamic parameters and physiological signals into token outputs that conform to human cognition through feature fusion and semantic mapping.
[0044] S4. Evaluation summary generation; Select a large language model, build a text knowledge base, and develop a multi-reflection mechanism to build a token summary evaluation model, as follows: S41. Selection of large language model; S411. A single token only represents the short-term response in this time series segment. However, the test process is long-term and continuous, and there will be a large number of representative test segments during the process. All prediction results should be integrated and a summary output should be made. The pre-trained large language model performs excellently in terms of understanding ability and generalization. In order to aggregate a large number of evaluation token prediction results into a complete test conclusion, this section selects the Qwen (a large-scale language model developed by Alibaba Cloud) model as the core LLM model for summarizing evaluation results. Qwen performs excellently in the field of Chinese text understanding and generation, is good at capturing complex logical relationships and causal chains, and has efficient integration and reasoning capabilities, making it an ideal choice for aggregating test conclusions.
[0045] S42. Development of enhanced knowledge base retrieval; S421. Although Qwen has powerful language understanding ability and generalization performance, there may still be problems with inaccurate expressions when dealing with complex and ambiguous situations. Therefore, this section introduces the RAG (Retrieval-Augmented Generation) retrieval enhancement technology, which helps the model obtain more background information and improve the depth of summarization by combining external knowledge bases and real-time retrieval, as Figure 7 shown.
[0046] The text materials collected in the present invention include the following parts: ① Theoretical literature and technical reports, including theoretical materials on vehicle performance, driving experience, autonomous driving test evaluation, etc.; ② Systematic evaluation databases, collecting a large amount of structured driver evaluation data, covering detailed feedback in different driving scenarios; ③ Domain expertise, covering expert interviews, training materials and tutorials, providing professional guidance; ④ Historical test records, detailed reports and representative cases of previous similar tests, serving as reference templates.
[0047] S423. The collected text materials are in various formats, including PDF, Word, TXT, Excel, etc. To ensure the effective utilization of the materials, targeted processing is carried out on texts in various formats. In the text processing stage, irrelevant characters, punctuation marks, and special formats are removed. In the text segmentation stage, the size limit of each segmentation block is set to 200 characters, and the number of shared characters between two blocks is set to 20 to maintain context coherence. The segmented text is vectorized, and a knowledge vector library is established based on Chroma. For the input question, a similarity text evaluation method based on Maximal Marginal Relevance (MMR) is used to search the vectorized text, increasing the diversity of retrieval results and avoiding duplicate information. The retrieved relevant texts are combined with the query prompt word template to form a complete prompt word, and the query result can be sorted out and output by inputting it into the language model.
[0048] S43. Design of multi-reflection mechanism; S431. To further enhance the quality of the evaluation results, based on the Self-Refine mechanism in the existing LLM field, the present invention designs a Multi-Self-Refine mechanism to optimize the output of the LLM. This mechanism includes a summary module, a context awareness module, a short-term memory module, and a long-term memory module. The generation of multiple agents helps to improve the diversity of the generated results, explore different expression methods, and generate the best summary text through the "evaluate - select - suggest - optimize" mechanism.
[0049] S432. After the input of the token group of driving experience, physical feelings, and emotional responses, keywords are first extracted for RAG retrieval. Each summary module will generate a test report summary through the prompt words of guided summary and RAG results and immediately store it in the short-term memory module. Through multiple experiments, it is finally determined that when the number of summary modules is 3, the generation effect and resource consumption can be well balanced. After one round of retrieval, the context awareness module will receive the above output and score the output results from two aspects: semantic coherence (evaluating whether the summary text is logically clear, smoothly expressed, and can accurately reflect the driver's overall experience) and result accuracy (checking whether the summary content covers all important details and provides reasonable explanations), and give specific improvement suggestions. Finally, it is judged whether the execution is completed according to three indicators: "whether the specified generation times are reached", "whether the predetermined quality standard is reached", and "whether the quality change rate of two consecutive rounds is lower than the set value". If the execution is not completed, the summary module is required to adjust according to the suggestions. If the execution is completed, the final evaluation report is feedback and output. As Figure 8 shown.
[0050] Example Two: This example conducts experimental verification and evaluation on Example One, specifically as follows: S5. Evaluate the performance of each sub-module and evaluate the key performance of the entire method for intelligent vehicle test scenarios. Design a control group for comprehensive analysis as follows: S51. Comparative verification of physiological signal prediction models; S511. High-accuracy physiological signal prediction is the basis for the successful execution of subsequent tasks. To verify the effectiveness of the physiological signal prediction model based on temporal GCN-Transformer proposed in the present invention for correlation analysis, the physiological signal prediction model based on temporal GCN-Transformer is used as experimental group 1, the algorithm that only uses the Transformer architecture without graph structure is used as control group 1-1, and the physiological signal prediction model based on temporal GCN-Transformer is applied to the original dataset without any feature selection or correlation analysis as control group 1-2.
[0051] S512. To further quantitatively analyze the prediction effects of each model, the mean squared error (MSE) is used as the evaluation index to measure the average of the squares of the differences between the predicted values and the true values.
[0052] S52. Comparative verification of token matching models S521. The core task of the present invention is to use multi-modal data to evaluate test scenarios. Therefore, it is necessary to verify the accuracy of evaluating token prediction by integrating chassis data and predicted physiological data. To verify the influence of feature selection and data type on the prediction results, the dataset containing statistical feature processing is used, and the predicted results and true results are mixed for training as experimental group 2; no feature processing is performed, and the predicted results and true results are mixed for training as control group 2-1, feature processing is performed but only the predicted physiological results are used for training as control group 2-2; feature processing is performed but only the original physiological results are used for training as control group 2-3.
[0053] S523. Further quantify the application effect of the token matching model based on the confusion matrix. For each token level described in the above scenario evaluation, a confusion matrix is constructed respectively, and the accuracy, precision, recall, and F1 score of category matching are calculated. The calculation methods are as follows: The accuracy represents the proportion of all correctly predicted samples in the total number of samples: , the precision represents the proportion of samples predicted as a certain category that are actually of that category: , the recall represents the proportion of samples that are actually of a certain category and are correctly predicted as that category: , the F1 score represents the harmonic mean of the precision and recall: 。
[0054] S53. Evaluation and summary model comparison and verification; S531. Summarizing discrete and sequential tokens into a complete and systematic evaluation text is the ultimate task of the evaluation model in this paper. To verify the effectiveness of the model in generating test conclusions and analyze the impact of the RAG and Multi-Self-Refine mechanisms on the model performance through ablation studies, the following groups are designed: Experimental group 3 uses the complete Qwen-based evaluation result summary model (including the RAG and Multi-Self-Refine mechanisms), Control group 3-1 is the Qwen model with only the Multi-Self-Refine mechanism, Control group 3-2 is the Qwen model with only the RAG, and Control group 3-3 is the Qwen baseline model without the RAG and Multi-Self-Refine.
[0055] S532. During the experiment, generate test reports based on the continuous tokens generated by the above process. To evaluate the generation effects of each group, the following evaluation methods are designed: ① Semantic similarity, which is used to measure the similarity degree of two texts at the semantic level. First, convert the text into a fixed-length vector representation based on the pre-trained language model (Sentence-BERT), and then use the cosine semantic similarity evaluation method in the large model field to calculate the average semantic similarity between the input tokens and the generated report to evaluate whether the model accurately integrates the keywords into the generated text. ② Structural diversity refers to the richness and variability of the generated text in terms of vocabulary and syntax. Texts with high diversity not only avoid monotony and repetition but also can better attract readers and convey more information. It is comprehensively evaluated by vocabulary diversity and the standard deviation of sentence length distribution. ③ Manual evaluation: Evaluate the quality of the generated text by having evaluation experts score it, which mainly includes the following four parts: 1) Logical clarity: Evaluate whether the generated report is logically coherent and whether the information is conveyed clearly. 2) Expressive fluency: Check whether the language expression is natural and fluent, and whether there are grammar errors or incoherences. 3) Content integrity: Ensure that the generated report covers all key tokens and these tokens are reasonably integrated into the report. 4) Result consistency: Evaluate the consistency between the report and the input tokens to ensure that no important information is omitted or irrelevant information is introduced. By designing a scoring form, invite multiple experts to score each dimension (1-10 points), and take the average value as the final score.
[0056] S54. Overall test scenario evaluation comparison and verification; S541. After validating each part of the model, the human-like evaluation method for intelligent vehicle test scenarios driven by the large model proposed in the present invention is compared and validated with the traditional scenario evaluation method, and the following groups are designed: Experimental group 4 is the human-like evaluation method for test scenarios driven by the large model proposed in the present invention; Control group 4-1 relies only on the objective indicators of vehicle dynamics, and defines the criticality of test scenarios by the acceleration limit value; Control group 4-2 relies on the driver's rating scale for driving experience.
[0057] S542. Based on the simulation software to replay a batch of test scenarios, criticality evaluations are carried out by different methods, and the testers make self-judgments to evaluate the differences between different methods in terms of scenario criticality evaluation.
[0058] In summary, through multi-condition data collection, physiological signal prediction, evaluation token matching, and multi-token induction and summary, the present invention finally realizes the multi-dimensional human-like evaluation output of the criticality of test scenarios, which can effectively improve the evaluation accuracy and reduce the data collection and experimental costs in practical applications.
[0059] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A human-like evaluation method for intelligent vehicle test scenarios driven by a large model, characterized in that: The following steps are involved: S1, data collection and processing; Design collection conditions, build collection equipment, collect data at the test site, and perform preliminary data processing; S2, physiological signal prediction; Conduct correlation analysis on the data, perform hierarchical processing based on data features, design the physiological signal prediction model architecture, and conduct model training; S3, evaluate word-unit matching; Process the evaluation text data, construct the word graph, perform feature processing on the time series data, design the word matching model architecture, and perform model training; S4, evaluation summary generation; Select a large language model, build a text knowledge base, develop a multi-reflection mechanism, and build a word-meta summary evaluation model.
2. According to claim 1, a large model-driven human-like evaluation method for intelligent vehicle test scenarios is characterized in that: The specific method of S1 is as follows: S11, collection working condition design; S12, collection equipment construction and field collection; S13. Data processing.
3. According to claim 2, a large model-driven human-like evaluation method for intelligent vehicle test scenarios is characterized in that: The specific method of S11 is as follows: S111. Based on the test standard procedures of SAE J2944, ISO 7401 and SAE J266, design the collection conditions with prominent single-dimensional indicators and coupled multi-dimensional indicators, that is, single and coupled in the horizontal, vertical and vertical aspects; S112. The designed working conditions cover the longitudinal, lateral and vertical movements in vehicle dynamics; the test parameters follow ISO 8855 and ISO 10844 standard documents; The specific method of S12 is as follows: S121. Use BioNomadix wireless physiological recorder to synchronously collect driver's electromyography, electrocardiography and skin conduction signals through three-lead DryPad electrodes, with the sampling rate set to 2kHz; collect chassis longitudinal / lateral acceleration at a frequency of 100Hz through CAN FD bus protocol; Based on the ASENSING GNSS / RTK combined positioning system, IMU inertial measurement unit and DTU data transmission unit, centimeter-level pose accuracy information is obtained; The vehicle's chassis signal and GNSS posture data are transmitted to the SpeedGoat real-time processing platform via the CAN bus protocol; Based on the timing trigger module, 5V TTL trigger pulses are sent to achieve microsecond synchronization of multiple devices; all data are transmitted to a laptop computer and saved in a storage medium through MATLAB software and physiological signal processing software; S1222. The test site is a special vehicle test site, and the driver shall select a professional test driver at the test site; A 5-minute rest period was conducted before each acquisition process, allowing the driver to enter the test in a calm state to collect baseline physiological signals, and three pre-experimental calibration drives were conducted to eliminate equipment interference; the evaluation data included real-time sentence output during acquisition and a complete summary after acquisition; The specific method of S13 is as follows: S131, using the rising edge of the hardware trigger signal as the time reference, and realizing 100 Hz unified resampling through the cubic spline interpolation algorithm; S132, establishing a multi-stage filtering process: using Kalman filtering to suppress high-frequency noise for dynamic signals, and implementing wavelet threshold noise reduction combined with adaptive filtering algorithm for physiological signals; Among them, the state equation and observation equation in the Kalman filter state space model are: ; ; In the formula, is the state transfer matrix; is the process noise; is the observation matrix; is the observation noise; For the control input matrix, the control vector is mapped to the change of the state variable; is the control vector, indicating the known external inputs applied to the system; S133. Physiological data includes real-time electrical signals and key physiological features. In terms of physiological feature extraction, physiological feature processing is performed based on the NeuroKit2 physiological signal analysis framework. (1) For ECG signals, extract the clear signal after denoising. ;Calculate heart rate , reflecting the body's stress level or exercise status; quality indicators for evaluating ECG data ; and extract the peak positions of the QRS complex, P wave, and T wave , and , describing the electrical activity phase of the atria and ventricles during the entire heart cycle, and is used to detect abnormal cardiac rhythm; (2) For the EDA signal, first remove the noise to obtain a smooth EDA signal ; Then separate the slow-varying component , reflecting the tone of the autonomic nervous system; capturing rapidly changing components , which is related to transient emotional responses; quantifies the transient increase in conductance triggered by a stimulus, including its onset time , peak time ,high ,amplitude , rise time , recovery time , used to assess the individual's sensitivity to specific events; (3) For EMG signals, muscle activity features are extracted after filtering and envelope detection, including the processed EMG signal The strength of muscle contraction , duration , and the starting point of muscle activity and end point .
4. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 1 is characterized in that: The specific method of S2 is as follows: S21, data correlation analysis; S22, data feature classification processing; S23. Prediction model architecture design and training.
5. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 4 is characterized in that: The specific method of S21 is as follows: S211, a feature screening framework based on time-delay causal inference, which integrates cross-correlation analysis and Granger causality test to achieve interpretable mapping modeling of dynamic features to physiological responses; S212, using sliding window cross-correlation algorithm to quantify dynamic correlation strength, assuming chassis data is , the physiological data is , hysteresis The cross-correlation coefficient is defined as: ; By traversing The maximum correlation coefficient is extracted from the lag time window and the corresponding lag time , characterizes the optimal alignment of the two signals; through related studies such as human nerve conduction delay and visual persistence, the time Set to 300ms; S213, construct a bivariate vector autoregression model, assuming that the chassis data Physiological data There is a causal effect, and a two-order regression equation is established: ; The residual sum of squares of the constrained model and the unconstrained model are compared by the F test. If the statistic satisfies: ; Then reject the null hypothesis and judge for Granger causes; ADF test is used to verify the stationarity of time series, and first-order difference or logarithmic transformation is performed on non-stationary data; S214, based on the dual threshold mechanism, two types of analysis results are integrated: 1) in the cross-correlation analysis, the feature data with cross-correlation coefficient and lag time within the lag window are retained; 2) in the Granger causality test, the features with only statistical correlation but no causal explanatory power are eliminated; the final selected data are: ECG_rate, ECG_quality, EDA_clean, EDA_phasic,EMG_Amplitude, EMG_Activity; The specific method of S22 is as follows: The initial state of each working condition and peak state As a benchmark, calculate the data extreme value in the time series window The relative change percentage of the data is calculated, and the data are classified according to the relative change percentage; The specific method of S23 is as follows: S231. Design a cascade prediction method based on temporal graph attention network and Transformer model; it includes four parts: graph network conversion, temporal feature compression, graph attention extraction and Transformer prediction; S232. First, a fully connected graph structure is constructed to characterize the relationship between sensors or measurement points, and each chassis dynamics data point is regarded as a node in the graph; second, the entire time series is divided into time windows of fixed length, and graph convolution operations are applied at each time step; noise reduction and smoothing are applied within each window; a graph convolution layer with an attention mechanism is introduced to assign a learnable attention coefficient to each pair of edges between nodes, and dynamically learn the importance weights between nodes; finally, the processed compact and semantically rich feature representation is sent to the Transformer-based encoder for further processing; the Transformer adopts a self-attention mechanism; finally, the prediction result of each time step is output through the fully connected layer, completing the end-to-end mapping from input to output.
6. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 1 is characterized in that: The specific method of S3 is as follows: S31, word-gram framework construction; S32, dataset construction; S33. Model design and training.
7. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 6 is characterized in that: The specific method of S31 is as follows: A multi-dimensional driving evaluation system is built based on knowledge graph technology. The real-time voice evaluation and final summary feedback of the driver under typical working conditions are recorded synchronously through the vehicle data acquisition system, and combined with vehicle dynamics parameters and physiological signals to form an interpretable semantic mapping network. Through the induction and standardization of the evaluation text, a classification system is finally formed, which includes two parts: physiological feelings and chassis response. Physiological feelings include driving experience and physical feelings. Chassis response includes four parts: power response, steering characteristics, braking performance and suspension feedback. The specific method of S32 is as follows: S321. Design special prompt words; S322. During the collection process, the driver is required to make real-time evaluation output at each time point and give a final summary after the test. All collected data needs to be time-synchronized to ensure that each record has a unified timestamp; then, the raw text data is subjected to NLP preprocessing steps and the driver’s evaluation data is mapped to the corresponding time dimension; The global time is used to accurately match the evaluation word unit and the time series data; the sliding window method is used to construct a continuous "prompt word-chassis dynamics-physiological data-evaluation word unit" combined data set; the data in each window represents the complete information within a time period; S323, introduce feature enhancement strategy, first extract statistical features, calculate the mean, variance, maximum and minimum values of the data in the window; then calculate the change trend of the data by moving average method, the specific process is as follows: for the standardized multi-dimensional time series data, take 30%, 70% and 100% of the data length as the data breakpoints, calculate the average value SMA of the data in each stage respectively, and consider the SMA change greater than 0.1 as rising, less than -0.1 as falling, and between as stable, so as to convert the time series data into the semantic description of the change trend, including rising, falling, stable, rising first and then falling, rising first and then stable, falling first and then rising, falling first and then stable, stable first and then rising, stable first and then falling; The specific method of S33 is as follows: S331. First, time series feature compression is performed. The multi-dimensional time series data is compressed into a 256-dimensional feature vector through a 1D convolution layer and a global maximum pooling layer. Then, it is input into the BERT-base model together with the structured prompt text to generate a 768-dimensional semantic representation. Six independent classification heads are added after the BERT tag output layer, corresponding to physical feelings, emotional reactions, power responses, steering characteristics, braking performance, and suspension feedback. Each classification head adds a fully connected layer on top of the last hidden state of BERT, and applies the softmax activation function to generate the probability distribution of each category. During the training process, the cross entropy loss function is used as the optimization target. After the input data is encoded by BERT, it will be passed to the respective classification heads to generate the predicted probability of each category. The predicted probability value is used to calculate the loss and guide the gradient update during the back-propagation process. S332. Design a weighted loss function to assign weights to each category based on the frequency of occurrence of each category. The weights are calculated based on the frequency of occurrence of each category in the data set. The calculation method is as follows: ; In the formula, For the The total number of samples in each subclass; is the number of word categories under this subcategory; Subcategory The number of times a word appears; S333. Define the AdamW optimizer and adopt the weight decay strategy to avoid overfitting. Introduce the learning rate scheduler of linear warm-up and cosine annealing to help the model converge more stably. Set the maximum gradient norm limit to prevent the gradient explosion problem. In each epoch, traverse the training data loader, batch process the input data and calculate the loss value. Apply the gradient accumulation technology to improve the effect of small batch training, and regularly evaluate the performance on the validation set. If there is no performance improvement for three consecutive batches, terminate early. Through the design and training process, the BERT model converts kinetic parameters and physiological signals into word outputs that conform to human cognition through feature fusion and semantic mapping.
8. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 1 is characterized in that: The specific method of S4 is as follows: S41, Large Language Model Selection; S42, knowledge base retrieval enhancement development; S43. Multi-reflection mechanism design.
9. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 8, characterized in that: The specific method of S41 is as follows: The Qwen model was selected as the core LLM model for summarizing the evaluation results; The specific method of S42 is as follows: S421, introduce RAG retrieval enhancement technology, which helps the model obtain more background information and improve the depth of summary by combining external knowledge base and real-time retrieval; S422. Collect text materials including the following: ① Theoretical literature and technical reports, including vehicle performance, driving experience, and autonomous driving test evaluation; ② A systematic evaluation database, which collects structured driver evaluation data and covers detailed feedback in different driving scenarios; ③Domain expertise, including training materials and tutorials, provides professional guidance; ④Historical test records, detailed reports of similar tests in the past and representative cases, as reference templates; S423, process the collected texts in various formats in a targeted manner. In the text processing stage, remove irrelevant characters, punctuation marks and special formats. In the text segmentation stage, set the segmentation block size limit to 200 characters and the number of characters shared between two blocks to 20 to maintain contextual coherence; vectorize the segmented text; establish a knowledge vector library based on Chroma; for input problems, search the vectorized text based on the similar text evaluation method of maximum marginal relevance to increase the diversity of retrieval results and avoid duplicate information; combine the retrieved relevant text with the query prompt word template to form a complete prompt word, and input the language model to organize and output the query results; The specific method of S43 is as follows: S431. Based on the Self-Refine mechanism in the LLM field, a Multi-Self-Refine mechanism is designed to optimize the output of LLM; the mechanism includes a summary module, a context-aware module, a short-term memory module, and a long-term memory module; S432. After the word tuples of driving experience, physical feelings and emotional reactions are input, keywords are first extracted for RAG retrieval; each summary module will generate a test report summary through the prompt words of the guided summary and RAG results, and immediately store it in the short-term memory module. After a round of retrieval, the context perception module will receive the output, score the output results from the two aspects of semantic coherence and result accuracy, and give specific improvement suggestions; finally, based on the three indicators of "whether the specified number of generations is reached", "whether the predetermined quality standards are met" and "whether the quality change rate of two consecutive rounds is lower than the set value", it is judged whether the execution is completed. If the execution is not completed, the summary module is required to make adjustments according to the suggestions. If the execution is completed, the final evaluation report is fed back and output.
Citation Information
Patent Citations
Human-like value alignment method and system based on multi-dimensional feedback reinforcement learning
CN118013016A
Machine dialogue capability improving method and system based on social simulation system
CN118468930A
Intelligent automobile virtual simulation test method based on large language model
CN118586281A
Examination test question generation method based on large model
CN118839003A
Lightweight process form intelligent prediction method and system
CN119443760A
Cited By
Method for processing multi-modal physiological data of driver
CN122175024A