A large-scale model-driven human-like evaluation method for intelligent vehicle test scenarios

Through a large-model driven human-like evaluation method for intelligent vehicle test scenarios, combined with multi-working condition data collection and physiological signal prediction, the multi-classification application of the BERT model and knowledge graph technology are designed to solve the high cost and low precision problems of existing evaluation methods, and realize low-cost, high-precision multi-dimensional evaluation, reducing data collection costs and improving evaluation accuracy.

CN120145204BActive Publication Date: 2025-09-23JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510630192.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-23
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing smart car test scenario evaluation methods are difficult to fully reflect multi-dimensional nonlinear information, and high-precision physiological data acquisition equipment and professional evaluation personnel bring problems of high cost and low efficiency.

Method used

A large-scale model-driven human-like evaluation method for intelligent vehicle test scenarios is adopted. Through multi-working condition data collection, physiological signal prediction, evaluation word unit matching and multi-word unit induction and summary, combined with a multimodal heterogeneous data acquisition system, using the BioNomadix wireless physiological recorder and the SpeedGoat real-time processing platform, a physiological signal prediction model based on time series GCN and a multi-classification application method of the BERT model are designed to construct an evaluation system for knowledge graph technology.

Benefits of technology

It achieves low-cost, high-precision multi-dimensional human-like evaluation, improves the accuracy of critical evaluation of test scenarios, reduces data collection and experimental costs, and realizes adaptive summary of evaluation results through the multi-reflection mechanism of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145204B_ABST
    Figure CN120145204B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of autonomous vehicle testing technology, specifically a large-scale model-driven human-like evaluation method for intelligent vehicle test scenarios. It includes the following steps: S1: data acquisition and processing; S2: physiological signal prediction; S3: evaluation word unit matching; and S4: evaluation summary generation. Through multi-operational data acquisition, physiological signal prediction, evaluation word unit matching, and multi-word unit summarization, this method ultimately achieves a multi-dimensional, human-like evaluation output of key test scenarios. In practical applications, this method can effectively improve evaluation accuracy and reduce data acquisition and experimental costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous vehicle testing, and specifically provides a large-model-driven human-like evaluation method for intelligent vehicle test scenarios. Background Art

[0002] Testing and evaluation technologies are the foundation and prerequisite for the industrialization of intelligent vehicles. Compared to traditional mileage-based testing methods, scenario-based testing abstracts the real world into fragmented test scenarios. Through parameterized design, it enables efficient world description and test execution, becoming a core method for intelligent vehicle performance evaluation. Key scenarios, in particular, can expose potential defects in the system under test and possess extremely high testing and application value. Assessing key scenarios is a hot topic in current research and serves as the foundation for subsequent research in scenario generation and automated testing.

[0003] Research on critical evaluation of intelligent vehicle test scenarios currently focuses on objective data evaluation and human evaluation. In the objective data evaluation area, objective threshold evaluation methods based on vehicle dynamics were first proposed and widely used. To further enhance the interpretability of the evaluation process, researchers have introduced the concept of risk potential fields to visualize the potential risk distribution of vehicles in complex traffic environments. The ultimate goal of intelligent vehicle development is to become intelligent tools that fully serve humanity. Their commodity nature dictates that market acceptance is highly dependent on actual user experience. Therefore, establishing a comprehensive evaluation system that reflects human-like characteristics is crucial. Regarding human evaluation, basic approaches based on questionnaires or scales have been widely used. By collecting direct user feedback on the driving experience, they provide important insights for improving intelligent vehicle performance. Furthermore, in the field of human-robot co-driving, researchers are incorporating occupant physiological signals into decision-making strategies to better meet human driving needs. However, existing research still has some shortcomings: existing subjective evaluation methods only evaluate from a deterministic numerical dimension, and the resulting subjective scale values ​​are difficult to fully reflect the multidimensional nonlinear information during the test; physiological data can reflect human psychological characteristics, but high-precision physiological data acquisition equipment, professional evaluation personnel and comprehensive testing sites have brought significant cost and efficiency problems to the evaluation process.

[0004] Therefore, there is an urgent need for an evaluation system that comprehensively considers multi-dimensional information, finds effective alternatives to physiological data, and achieves low-cost, high-precision quantitative evaluation of key test scenarios. Summary of the Invention

[0005] To solve the above problems, the present invention provides a large-model-driven human-like evaluation method for intelligent automobile test scenarios. Through multi-working condition data collection, physiological signal prediction, evaluation word unit matching, and multi-word unit induction and summary, it ultimately achieves a critical multi-dimensional human-like evaluation output of the test scenario. In practical applications, it can effectively improve evaluation accuracy and reduce data collection and experimental costs.

[0006] The technical solution of the present invention is described as follows in conjunction with the accompanying drawings:

[0007] The present invention provides a large-model driven human-like evaluation method for intelligent vehicle test scenarios, comprising the following steps:

[0008] S1, data collection and processing;

[0009] Design data collection conditions, build data collection equipment, collect data at the test site, and perform preliminary data processing;

[0010] S2, physiological signal prediction;

[0011] Conduct correlation analysis on the data, perform hierarchical processing based on data features, design a physiological signal prediction model architecture, and conduct model training;

[0012] S3, evaluate word-gram matching;

[0013] Process the evaluation text data, construct a lemma graph, perform feature processing on time series data, design a lemma matching model architecture, and perform model training;

[0014] S4, evaluation summary generation;

[0015] Select a large language model, build a text knowledge base, develop a multi-reflection mechanism, and build a word-meta summary and evaluation model.

[0016] Furthermore, the specific method of S1 is as follows:

[0017] S11, collection working condition design;

[0018] S12, collection equipment setup and field collection;

[0019] S13. Data processing.

[0020] Furthermore, the specific method of S11 is as follows:

[0021] S111. Based on the test standard procedures of SAE J2944, ISO 7401 and SAE J266, design the collection conditions with prominent single-dimensional indicators and coupled multi-dimensional indicators, i.e. single and coupled in the horizontal, vertical and vertical aspects;

[0022] S112. The designed working conditions cover the longitudinal, lateral and vertical motions in vehicle dynamics; the test parameters comply with ISO8855 and ISO 10844 standards.

[0023] The specific method of S12 is as follows:

[0024] S121. Use a BioNomadix wireless physiological recorder to synchronously collect the driver's electromyographic, electrocardiographic, and electrodermal signals via three-lead DryPad electrodes, with a sampling rate set to 2kHz. Chassis longitudinal and lateral accelerations are collected at a frequency of 100Hz via the CAN FD bus protocol. Centimeter-level position accuracy is obtained using the ASENSING GNSS / RTK combined positioning system, IMU inertial measurement unit, and DTU data transmission unit. The vehicle's chassis signals and GNSS position data are transmitted to the SpeedGoat real-time processing platform via the CAN bus protocol. A timing trigger module sends 5V TTL trigger pulses to achieve microsecond-level synchronization of multiple devices. All data are transmitted to a laptop computer and stored on a storage medium using MATLAB software and physiological signal processing software.

[0025] S1222. The test site is a dedicated vehicle proving ground, and the driver is a professional test driver from the proving ground. A 5-minute rest period is conducted before each acquisition cycle to allow the driver to enter the test in a calm state and collect baseline physiological signals. Three pre-experimental calibration drives are also conducted to eliminate equipment interference. Evaluation data includes real-time sentence output during acquisition and a complete summary after acquisition.

[0026] The specific method of S13 is as follows:

[0027] S131, using the rising edge of the hardware trigger signal as the time reference, and implementing 100 Hz uniform resampling through the cubic spline interpolation algorithm;

[0028] S132. Establish a multi-stage filtering process: use Kalman filtering to suppress high-frequency noise on dynamic signals, and implement wavelet threshold noise reduction combined with adaptive filtering algorithm on physiological signals;

[0029] Among them, the state equation and observation equation in the Kalman filter state space model are:

[0030] ;

[0031] ;

[0032] Where, is the state transfer matrix; is the process noise; is the observation matrix; is the observation noise; To control the input matrix, the control vector is mapped to the change of the state variable; is the control vector, indicating the time Known external inputs applied to the system;

[0033] S133, physiological data includes real-time electrical signals and key physiological features; in terms of physiological feature extraction, physiological feature processing is performed based on the NeuroKit2 physiological signal analysis framework; (1) For ECG signals, extract the clear signal after denoising ;Calculate heart rate , reflecting the body's stress level or exercise state; quality indicators for evaluating ECG data ; and extract the peak positions of the QRS complex, P wave, and T wave 、 and , describing the electrical activity phase of the atria and ventricles during the entire cardiac cycle, and is used to detect cardiac rhythm abnormalities; (2) For EDA signals, first remove the noise to obtain a smooth EDA signal ; Then separate the slow-varying component , reflecting the tension of the autonomic nervous system; capturing rapidly changing components , related to transient emotional responses; quantifies the transient conductance increase triggered by a stimulus, including its onset time , peak time ,high ,amplitude , rise time , recovery time , used to assess the individual's sensitivity to specific events; (3) For EMG signals, muscle activity features are extracted after filtering and envelope detection, including the processed EMG signal The intensity of muscle contraction , duration , and the starting point of muscle activity and end point .

[0034] Furthermore, the specific method of S2 is as follows:

[0035] S21, data correlation analysis;

[0036] S22, data feature hierarchical processing;

[0037] S23. Prediction model architecture design and training.

[0038] Furthermore, the specific method of S21 is as follows:

[0039] S211, a feature screening framework based on time-delay causal inference, which integrates cross-correlation analysis and Granger causality test to achieve interpretable mapping modeling of dynamic features to physiological responses;

[0040] S212, using sliding window cross-correlation algorithm to quantify dynamic correlation strength, assuming chassis data is , physiological data are , hysteresis The cross-correlation coefficient is defined as:

[0041] ;

[0042] By traversing The maximum correlation coefficient is extracted from the lag time window and the corresponding lag time , characterizes the optimal alignment of the two signals; through relevant research on human nerve conduction delay, visual persistence, etc., the time Set to 300ms;

[0043] S213, build a two-variable vector autoregressive model, assuming that the chassis data Physiological data There is a causal effect, and a two-order regression equation is established:

[0044] ;

[0045] Compare the residual sum of squares of the constrained model and the unconstrained model through the F test. If the statistic satisfies:

[0046] ;

[0047] Then reject the null hypothesis and determine for Granger cause; and use ADF test to verify the stationarity of time series, and perform first-order difference or logarithmic transformation on non-stationary data;

[0048] S214. Two types of analysis results are integrated using a dual-threshold mechanism: 1) In cross-correlation analysis, feature data with cross-correlation coefficients and lag times within the lag window are retained; 2) In Granger causality testing, features with only statistical correlation but no causal explanatory power are eliminated. The final selected data are: ECG_rate, ECG_quality, EDA_clean, EDA_phasic, EMG_Amplitude, and EMG_Activity.

[0049] The specific method of S22 is as follows:

[0050] The initial state of each working condition and peak state As a benchmark, calculate the extreme value of the data in the time series window The relative percentage change of the data is calculated, and the data are graded according to the relative percentage change;

[0051] The specific method of S23 is as follows:

[0052] S231. Design a cascade prediction method based on temporal graph attention network and Transformer model; it includes four parts: graph network conversion, temporal feature compression, graph attention extraction and Transformer prediction;

[0053] S232. First, a fully connected graph structure is constructed to characterize the relationship between sensors or measurement points. Each chassis dynamics data point is regarded as a node in the graph. Second, the entire time series is divided into time windows of fixed length, and graph convolution operations are applied at each time step. Noise reduction and smoothing are applied within each window. A graph convolution layer with an attention mechanism is introduced to assign a learnable attention coefficient to the edge between each pair of nodes, and dynamically learn the importance weights between nodes. Finally, the compact and semantically rich feature representation obtained is sent to the Transformer-based encoder for further processing. The Transformer adopts a self-attention mechanism. Finally, the prediction result of each time step is output through the fully connected layer, completing the end-to-end mapping from input to output.

[0054] Furthermore, the specific method of S3 is as follows:

[0055] S31, word-gram framework construction;

[0056] S32, dataset construction;

[0057] S33. Model design and training.

[0058] Furthermore, the specific method of S31 is as follows:

[0059] A multi-dimensional driving evaluation system is constructed based on knowledge graph technology. The on-board data acquisition system simultaneously records the driver's real-time voice evaluation and final summary feedback under driving conditions, and combines them with vehicle dynamics parameters and physiological signals to form an interpretable semantic mapping network. Through the induction and standardization of the evaluation text, a classification system is ultimately formed, which includes two parts: physiological sensation and chassis response. Physiological sensation includes driving experience and physical sensation. Chassis response includes power response, steering characteristics, braking performance, and suspension feedback.

[0060] The specific method of S32 is as follows:

[0061] S321. Design special prompt words;

[0062] S322. During the data collection process, the driver is required to provide real-time evaluation output at each time point and provide a final summary after the test. All collected data must be time-synchronized to ensure that each record has a unified timestamp. Subsequently, the raw text data is subjected to NLP preprocessing steps, and the driver's evaluation data is mapped to the corresponding time dimension. The global time is used to accurately match the evaluation word unit and the time series data. A sliding window method is used to construct a continuous "prompt word-chassis dynamics-physiological data-evaluation word unit" combined data set. The data in each window represents the complete information within a time period.

[0063] S323. Introduce a feature enhancement strategy. First, extract statistical features and calculate the mean, variance, maximum, and minimum values ​​of the data in the window. Then, calculate the data change trend through the moving average method. The specific process is as follows: for the standardized multi-dimensional time series data, use 30%, 70%, and 100% of the data length as the data breakpoints, and calculate the average value SMA of the data in each stage respectively. An SMA change greater than 0.1 is considered to be an increase, less than -0.1 is considered to be a decrease, and the range is considered to be stable. In this way, the time series data is converted into a semantic description of the change trend, including 9 types: increase, decrease, stability, increase first and then decrease, increase first and then stabilize, decrease first and then increase, decrease first and then stabilize, stabilize first and then increase, and stabilize first and then decrease.

[0064] The specific method of S33 is as follows:

[0065] S331. First, time series feature compression is performed. The multi-dimensional time series data is compressed into a 256-dimensional feature vector through a 1D convolutional layer and a global maximum pooling layer. It is then combined with the structured prompt text and input into the BERT-base model to generate a 768-dimensional semantic representation. Six independent classification heads are added after the BERT labeled output layer, corresponding to physical sensation, emotional reaction, power response, steering characteristics, braking performance, and suspension feedback. Each classification head adds a fully connected layer on top of the last hidden state layer of BERT and applies the softmax activation function to generate the probability distribution of each category. During training, the cross-entropy loss function is used as the optimization target. After the input data is encoded by BERT, it is passed to the respective classification head to generate the predicted probability of each category. The predicted probability value is used to calculate the loss and guide the gradient update during the backpropagation process.

[0066] S332. Design a weighted loss function to assign weights to each category based on the frequency of occurrence of each category. The weights are calculated based on the frequency of occurrence of each category in the dataset. The calculation method is as follows:

[0067] ;

[0068] Where, For the The total number of samples in each subclass; is the number of word categories under this subcategory; For subclass The number of times a word appears;

[0069] S333. Define the AdamW optimizer and adopt a weight decay strategy to avoid overfitting. Introduce a learning rate scheduler with linear warm-up and cosine annealing to help the model converge more stably. Set a maximum gradient norm limit to prevent the gradient explosion problem. In each epoch, traverse the training data loader, batch process the input data and calculate the loss value. Apply the gradient accumulation technique to improve the effect of small-batch training, and regularly evaluate the performance on the validation set. If there is no performance improvement for three consecutive batches, that is, the evaluation accuracy no longer increases, terminate early. Through the design and training process, the BERT model converts dynamic parameters and physiological signals into word outputs that conform to human cognition through feature fusion and semantic mapping.

[0070] Furthermore, the specific method of S4 is as follows:

[0071] S41, Large Language Model Selection;

[0072] S42, knowledge base retrieval enhancement development;

[0073] S43. Multi-reflection mechanism design.

[0074] Furthermore, the specific method of S41 is as follows:

[0075] The Qwen model was selected as the core LLM model for summarizing the evaluation results;

[0076] The specific method of S42 is as follows:

[0077] S421. Introducing RAG retrieval enhancement technology, which helps the model obtain more background information and improve the depth of summary by combining external knowledge base and real-time retrieval;

[0078] S422. Collect textual materials, including the following: ① Theoretical literature and technical reports covering vehicle performance, driving experience, and autonomous driving test evaluations; ② A systematic evaluation database, collecting structured driver evaluation data and providing detailed feedback from different driving scenarios; ③ Domain expertise, including training materials and tutorials, providing professional guidance; ④ Historical test records, including detailed reports and representative cases of similar tests in the past, serving as reference templates;

[0079] S423. Targeted processing is performed on the collected text in various formats. During the text processing phase, irrelevant characters, punctuation marks, and special formats are removed. During the text segmentation phase, the segmentation block size is limited to 200 characters, and the number of characters shared between two blocks is set to 20 to maintain contextual coherence. The segmented text is vectorized. A knowledge vector library is established based on Chroma. For input problems, the vectorized text is searched based on a similar text evaluation method based on maximum marginal relevance to increase the diversity of search results and avoid duplication of information. The retrieved relevant text is combined with the query prompt word template to form a complete prompt word, which is input into the language model to organize and output the query results.

[0080] The specific method of S43 is as follows:

[0081] S431. Based on the Self-Refine mechanism in the LLM field, a Multi-Self-Refine mechanism is designed to optimize the output of LLM. The mechanism includes a summary module, a context-aware module, a short-term memory module, and a long-term memory module.

[0082] S432. After the word tuples of driving experience, physical feelings and emotional reactions are input, keywords are first extracted for RAG search. Each summary module will generate a test report summary through the prompt words of the guided summary and RAG results, and immediately store it in the short-term memory module. After one round of retrieval, the context perception module will receive the output, score the output results in terms of semantic coherence and result accuracy, and give specific improvement suggestions. Finally, based on the three indicators of "whether the specified number of generations is reached", "whether the predetermined quality standards are met" and "whether the quality change rate of two consecutive rounds is lower than the set value", it is judged whether the execution is completed. If the execution is not completed, the summary module is required to make adjustments based on the suggestions. If the execution is completed, the final evaluation report is output as feedback.

[0083] The beneficial effects of the present invention are:

[0084] (1) The present invention achieves a multi-dimensional human-like evaluation output that is critical to the test scenario through multi-condition data collection, physiological signal prediction, evaluation word unit matching, and multi-word unit summarization. In practical applications, it can effectively improve the evaluation accuracy and reduce the data collection and experimental costs.

[0085] (2) This paper proposes a physiological signal prediction method based on time series GCN, which screens typical physiological features through correlation analysis, and significantly improves the chassis-physiology prediction accuracy through time series data segmentation, graph structure conversion, and graph attention extraction.

[0086] (3) The present invention designs an evaluation word unit matching method based on a large language model and a multi-classification application method of the BERT model to map complex and disordered physiological and chassis data to highly concentrated human evaluation words.

[0087] (4) The present invention designs an evaluation word-meta summary method based on a large language model, and realizes adaptive summary of evaluation results through knowledge base enhanced retrieval and multi-reflection mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0089] Figure 1 It is a schematic diagram of the process of the present invention;

[0090] Figure 2 Schematic diagram of the designed working condition;

[0091] Figure 3 Schematic diagram of the multimodal heterogeneous data acquisition system built;

[0092] Figure 4 Schematic diagram of feature data after extraction;

[0093] Figure 5 This is a schematic diagram of data classification rules;

[0094] Figure 6 This is a schematic diagram of the word unit evaluation system;

[0095] Figure 7 Schematic diagram of RAG retrieval enhancement technology;

[0096] Figure 8 Summarize the technical diagram for the evaluation report. DETAILED DESCRIPTION

[0097] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0098] Example 1:

[0099] See Figure 1 This embodiment provides a large-model driven human-like evaluation method for intelligent vehicle test scenarios, including the following steps:

[0100] S1, data collection and processing;

[0101] Design the data collection conditions, build the data collection equipment, collect data at the test site, and perform preliminary data processing, as follows:

[0102] S11, collection working condition design;

[0103] S111. When the driver is in the same initial state, different responses of vehicle dynamics will have different degrees of impact on the driver, which will lead to differences in objective physiological signals and driving experience. In order to accurately collect the driver's physiological signal responses and driving experience under different working conditions, the present invention designs a collection condition that emphasizes single-dimensional indicators and couples multi-dimensional indicators from the perspective of dynamic deconstruction based on various test standards and procedures such as SAEJ2944, ISO 7401 and SAE J266.

[0104] S112、Designed working conditions are as follows Figure 2 As shown, the designed operating conditions cover the longitudinal, lateral, and vertical motions of vehicle dynamics. Test parameters strictly adhere to standards such as ISO 8855 and ISO 10844. By setting dynamic excitations covering the ISO standard spectrum, the vehicle state transition process from normal driving to critical instability can be fully characterized.

[0105] S12, collection equipment setup and field collection;

[0106] S121, the multi-modal heterogeneous data acquisition system constructed by the present invention is as follows Figure 3 As shown in the figure, a BioNomadix wireless physiological recorder was used to synchronously collect the driver's electromyographic (EMG), electrocardiographic (ECG), and electrodermal (EDA) signals via three-lead DryPad electrodes, with a sampling rate set to 2kHz (compliant with the human vibration perception frequency band requirements of ISO 2631-1). Various dynamic parameters, such as chassis longitudinal and lateral acceleration, were collected at a 100Hz frequency via the CAN FD bus protocol. Centimeter-level pose accuracy was obtained using the ASENSING GNSS / RTK combined positioning system, IMU inertial measurement unit, and DTU data transmission unit. The vehicle's chassis signals and GNSS pose data were transmitted to the SpeedGoat real-time processing platform via the CAN bus protocol. A timed trigger module sent 5V TTL trigger pulses to achieve microsecond-level synchronization of multiple devices. All data was transmitted to a laptop and stored on a storage medium using MATLAB software and dedicated physiological signal processing software.

[0107] S122. The experimental site can be the Hainan Automobile Proving Ground in China. The driver should be a professional test driver from the proving ground. They have extensive driving and scoring experience, having conducted numerous tests to assess vehicle performance, including chassis data calibration, new car test drives, and handling stability testing. To ensure the stability of physiological indicators during data collection, a 5-minute resting period is performed before each collection cycle, allowing the driver to enter the test in a calm state to collect baseline physiological signals. Three pre-experimental calibration drives are also conducted to eliminate equipment interference. The evaluation data includes real-time sentence output during collection and a complete summary after collection, which can be used to effectively assist with subsequent word output in this study.

[0108] S13, data processing;

[0109] S131. To address the timing misalignment problem caused by differences in sampling frequencies of multi-source data, the rising edge of the hardware trigger signal is used as the time reference, and a cubic spline interpolation algorithm is used to achieve 100Hz unified resampling, effectively eliminating time deviations between devices.

[0110] S132. In order to solve the signal distortion problem caused by electromagnetic interference and baseline drift, a multi-stage filtering process is established: Kalman filtering is used to suppress high-frequency noise for dynamic signals, and wavelet threshold noise reduction combined with adaptive filtering algorithm is implemented for physiological signals to eliminate motion artifacts while retaining effective features.

[0111] Among them, the state equation and observation equation in the Kalman filter state space model are:

[0112] ;

[0113] ;

[0114] Where, is the state transfer matrix; is the process noise; is the observation matrix; is the observation noise; To control the input matrix, the control vector is mapped to the change of the state variable; is the control vector, indicating the time A known external input applied to a system.

[0115] S133, physiological data not only contains real-time electrical signals, but also contains a large number of in-depth key physiological features. In terms of physiological feature extraction, physiological feature processing is performed based on the NeuroKit2 physiological signal analysis framework. (1) For ECG signals, extract the clear signal after denoising ;Calculate heart rate , reflecting the body's stress level or exercise state; quality indicators for evaluating ECG data ; and extract the peak positions of the QRS complex, P wave, and T wave 、 and , describes the electrical activity phase of the atria and ventricles during the entire cardiac cycle and is used to detect cardiac rhythm abnormalities. (2) For EDA signals, first remove the noise to obtain a smooth EDA signal ; Then separate the slow-varying component , reflecting the tension of the autonomic nervous system; capturing rapidly changing components , related to transient emotional responses; quantifies the transient conductance increase triggered by a stimulus, including its onset time , peak time ,high ,amplitude , rise time , recovery time , used to assess the individual's sensitivity to specific events; (3) For EMG signals, muscle activity features are extracted after filtering and envelope detection, including the processed EMG signal The intensity of muscle contraction , duration , and the starting point of muscle activity and end point The extracted feature data is as follows: Figure 4 As shown, it can provide a standardized data basis for subsequent driving status analysis.

[0116] S2, physiological signal prediction;

[0117] Conduct correlation analysis on the data, perform hierarchical processing based on data features, design a physiological signal prediction model architecture, and conduct model training; the details are as follows:

[0118] S21, data correlation analysis;

[0119] S211. The collected chassis dynamics data and physiological signal data are both high-dimensional and complex time series with interactive coupling characteristics. Therefore, directly using multidimensional dynamic data to predict multidimensional physiological data can lead to problems such as overfitting and difficulty in convergence. To address these issues and improve the effectiveness of the prediction model, this study proposes a feature screening framework based on time-delayed causal inference. By integrating cross-correlation analysis and Granger causality testing, it achieves interpretable mapping modeling of dynamic features to physiological responses.

[0120] S212, according to the time asynchronous characteristics of physiological signals and chassis dynamics data, the sliding window cross-correlation algorithm is used to quantify the dynamic correlation strength. Assume that the chassis data is , physiological data are , its hysteresis The cross-correlation coefficient is defined as:

[0121] ;

[0122] By traversing The maximum correlation coefficient is extracted from the lag time window and the corresponding lag time , characterizes the optimal alignment of the two signals. Through the research on human nerve conduction delay, visual persistence, etc., the time Set to 300ms.

[0123] S213, in order to distinguish statistical correlation from causal dependence, a bivariate vector autoregressive model is constructed, assuming that the chassis data Physiological data There is a causal effect, and a two-order regression equation is established:

[0124] ;

[0125] Compare the residual sum of squares of the constrained model (containing only the Y lag term) and the unconstrained model (containing the X lag term) through the F test. If the statistic satisfies:

[0126] ;

[0127] Then reject the null hypothesis and determine for The ADF test (Augmented Dickey-Fuller Test) is used to verify the stationarity of the time series, and the non-stationary data are processed by first-order difference or logarithmic transformation to avoid pseudo-causal inference.

[0128] S214. A dual-threshold mechanism is used to fuse two types of analysis results: 1) In cross-correlation analysis, features with cross-correlation coefficients and lag times within the lag window are retained; 2) In Granger causality testing, a significance level is set to eliminate features that are only statistically correlated but lack causal power. This dual constraint of time-delay alignment and causal inference allows for a more comprehensive understanding of the complex interactions between chassis dynamics data and physiological signals, transcending the limitations of traditional Pearson correlation in modeling time-varying systems. This ensures that the selected features are not only statistically significant but also demonstrate causal relationships, providing a high signal-to-noise ratio input space for subsequent physiological prediction models. The final selected data sets are: ECG_rate, ECG_quality, EDA_clean, EDA_phasic, EMG_Amplitude, and EMG_Activity.

[0129] S22, data feature classification processing;

[0130] During the data collection process, it was found that different initial states of human drivers will affect the subsequent level of physiological signal changes. At the same time, physiological signals have small-scale fluctuations and graded change characteristics, and random data are difficult to accurately predict. In order to eliminate the baseline differences of individuals under different working conditions, the present invention uses the initial state of each working condition as the starting point. and peak state As a benchmark, calculate the data extreme value in the time series window (512 points) The relative percentage change ( ), data classification rules such as Figure 5 shown.

[0131] S23, prediction model architecture design and training;

[0132] S231. In actual data collection, chassis dynamics and physiological data are both continuous-time variables with high dimensionality, multiple perturbations, and complex spatiotemporal dependencies. Traditional methods struggle to effectively address these challenges. To address this challenge, this paper proposes a cascaded prediction method based on a Temporal Graph Attention Network (T-GAT) and a Transformer model. This method comprises four core components: graph network conversion, temporal feature compression, graph attention extraction, and Transformer prediction.

[0133] S232. First, a fully connected graph structure is constructed to represent the relationships between sensors or measurement points, with each chassis dynamics data point treated as a node in the graph. Second, the entire time series is divided into fixed-length time windows, and a graph convolution (GCN) operation is applied at each time step to simplify the original high-dimensional time series data while preserving important spatiotemporal dependencies. To further improve data quality, noise reduction and smoothing are applied within each window to reduce the impact of noise interference. Furthermore, a graph convolutional layer (GAT) with an attention mechanism is introduced to assign a learnable attention coefficient to each pair of node edges, dynamically learning the importance weights between nodes, effectively capturing the complex relationships in the graph structure. Finally, the compact and semantically rich feature representation obtained through the previous three stages is fed into a Transformer-based encoder for further processing. The Transformer utilizes a self-attention mechanism to fully exploit hidden correlations without sacrificing temporal order, and combines it with positional encoding to ensure that the inherent sequential information of the time series is not lost. Finally, a fully connected layer outputs the prediction result for each time step, completing the end-to-end mapping from input to output.

[0134] S3, evaluate word-gram matching;

[0135] Process the evaluation text data, construct a lemma graph, perform feature processing on time series data, design a lemma matching model architecture, and perform model training;

[0136] S31, word-gram framework construction;

[0137] S311. In order to systematically analyze the correlation mechanism between driver evaluation and vehicle objective data, the present invention builds a multi-dimensional driving evaluation system based on knowledge graph technology. The vehicle-mounted data acquisition system synchronously records the driver's real-time voice evaluation and final summary feedback under typical working conditions, and combines vehicle dynamic parameters and physiological signals to form an interpretable semantic mapping network. By summarizing and standardizing the evaluation text (merging synonyms and eliminating ambiguity), the following is finally formed. Figure 6 The classification system shown here encompasses both physiological sensations and chassis response. Physiological sensations encompass both the driving experience and physical sensations. Chassis response encompasses power response, steering characteristics, braking performance, and suspension feedback. As you can see, this index system, with its diverse vocabulary, captures the driver's multidimensional experience during the test, providing a more comprehensive and specific assessment than a single scale.

[0138] S32, dataset construction;

[0139] S321. In order to build a structured and consistent real-time evaluation dataset and ensure sufficient context information for model training and provide clear training objectives, this study designed a special prompt: "You are an intelligent vehicle test evaluator. Please evaluate the chassis dynamics data based on the chassis dynamics data. } and physiological index data { }, use words to describe your driving experience, physical sensations and emotional reactions."

[0140] S322. Scenario data is time-series, so during the collection process, the driver is required to combine his or her own driving experience to make real-time evaluation outputs at each time node, and give a final summary after the test. All collected data (including chassis dynamics signals, physiological signals and driver evaluations) need to be accurately time-synchronized to ensure that each record has a unified timestamp. Subsequently, the original text data is subjected to NLP pre-processing steps such as word segmentation and stop word removal, and the driver's evaluation data is mapped to the corresponding time dimension. Global time is used to accurately match evaluation words and time series data. A sliding window method is used to construct a continuous "prompt word-chassis dynamics-physiological data-evaluation word" combined data set. The data in each window represents complete information within a time period, which facilitates the model to learn the relationship between different modal data.

[0141] S323. Prediction of time series data, especially text data, is challenging. Larger amounts of data input and more dimensional data features have been shown to be beneficial for improving prediction results. Therefore, a feature enhancement strategy is introduced. First, statistical features are extracted and the mean, variance, maximum, and minimum values ​​of the data within the window are calculated. Then, the data trend is calculated using the moving average method. The specific process is as follows: for the standardized multi-dimensional time series data, 30%, 70%, and 100% of the data length are used as data breakpoints. The average SMA of the data at each stage is calculated respectively. An SMA change greater than 0.1 is considered to be rising, less than -0.1 is considered to be falling, and the SMA change between them is considered to be stable. In this way, the time series data is converted into a semantic description of the change trend, including rising, falling, stable, rising first and then falling, rising first and then stable, falling first and then rising, falling first and then stable, stable first and then rising, and stable first and then falling.

[0142] S33, model design and training;

[0143] S331. When processing large-scale data, long time series data cannot be directly input due to the performance limitations of language models. Therefore, time series feature compression is first performed. 1D convolutional layers and global max pooling layers are used to compress the multi-dimensional time series data into 256-dimensional feature vectors. This is then combined with structured prompt text and fed into the BERT-base model (with the underlying parameters frozen) to generate a 768-dimensional semantic representation. To adapt to the word prediction task of this study, six independent classification heads are added after the BERT labeling output layer, corresponding to physical sensation, emotional response, power response, steering characteristics, braking performance, and suspension feedback. Each classification head adds a fully connected layer on top of the last hidden state layer of BERT and applies a softmax activation function to generate a probability distribution for each class. During training, the cross-entropy loss function is used as the optimization objective. After the input data is encoded by BERT, it is passed to each classification head, which generates predicted probabilities for each class. These probabilities are used to calculate the loss and guide gradient updates during backpropagation.

[0144] S332. In the actual evaluation process, due to the personal habits of the evaluators and the characteristics of the scene working conditions, the frequency of occurrence of words such as "linear" that reflect the basic functions of the system will be higher than that of other words. The categories of each sub-item are unbalanced, and the model will tend to predict the high-frequency category to reduce the loss, resulting in a decrease in the recognition performance of the minority class. To this end, the present invention designs a weighted loss function. Based on the frequency of occurrence of each category, a weight is assigned to each category, so that the model pays more attention to the minority class. The weight is calculated based on the frequency of occurrence of each category in the data set. The calculation method is as follows:

[0145] ;

[0146] Where, For the The total number of samples in each subclass, is the number of word categories under this subcategory, For this subcategory By using this method, the gradient of the model for low-frequency classes will be enhanced, thus preventing the model from being biased towards the majority class.

[0147] S333. Define the AdamW optimizer and adopt a weight decay strategy to avoid overfitting. Introduce a learning rate scheduler with linear warmup and cosine annealing to help the model converge more stably. Set a maximum gradient norm limit to prevent gradient explosion. In each epoch, traverse the training data loader, batch process the input data and calculate the loss value. Apply gradient accumulation techniques to improve the effect of small-batch training, and regularly evaluate performance on the validation set. If there is no performance improvement after three consecutive batches, terminate early to save resources. Through the above design and training process, the BERT model can transform dynamic parameters and physiological signals into word-unit outputs that are consistent with human cognition through feature fusion and semantic mapping.

[0148] S4, evaluation summary generation;

[0149] Select a large language model, build a text knowledge base, develop a multi-reflection mechanism, and build a word-meta summary and evaluation model. The details are as follows:

[0150] S41, Large Language Model Selection;

[0151] S411. A single word only represents a short-term response within that timeframe. However, the testing process is long-term and continuous, involving numerous representative test segments. All prediction results should be synthesized and summarized. Pre-trained large language models excel in comprehension and generalization. To aggregate the prediction results of a large number of evaluation word units into a comprehensive test conclusion, this section uses Qwen (a large-scale language model developed by Alibaba) as the core LLM for summarizing evaluation results. Qwen excels in Chinese text understanding and generation, excels at capturing complex logical relationships and causal chains, and possesses efficient integration and reasoning capabilities, making it an ideal choice for summarizing test conclusions.

[0152] S42, knowledge base retrieval enhancement development;

[0153] S421. Although Qwen has strong language understanding and generalization capabilities, it may still have problems with imprecise expressions when dealing with complex and ambiguous situations. To this end, this section introduces RAG (Retrieval-Augmented Generation) retrieval enhancement technology. By combining external knowledge bases and real-time retrieval, it helps the model obtain more background information and improve the depth of summary, such as Figure 7shown.

[0154] S422. The text materials collected by the present invention include the following parts: ① theoretical literature and technical reports, including theoretical materials on vehicle performance, driving experience, autonomous driving test evaluation, etc.; ② a systematic evaluation database, which collects a large amount of structured driver evaluation data and covers detailed feedback in different driving scenarios; ③ domain expertise, which covers expert interviews, training materials and tutorials, and provides professional guidance; ④ historical test records, including detailed reports and representative cases of previous similar tests, which serve as reference templates.

[0155] S423. The collected text materials come in various formats, including PDF, Word, TXT, and Excel. To ensure effective use of the materials, targeted processing is performed on each format. During the text processing phase, irrelevant characters, punctuation, and special formatting are removed. During the text segmentation phase, the segmented text is limited to 200 characters, with the number of characters shared between two segments limited to 20 to maintain contextual coherence. The segmented text is vectorized, and a knowledge vector library is established based on Chroma. Regarding input, the vectorized text is searched using a similar text evaluation method based on Maximum Marginal Relevance (MMR), increasing the diversity of search results and avoiding duplication. The retrieved relevant text is combined with a query prompt template to form a complete prompt. This is then fed into a language model to organize and output query results.

[0156] S43, multi-reflection mechanism design;

[0157] S431. To further enhance the quality of evaluation results, this paper designs a Multi-Self-Refine mechanism, building on the existing Self-Refine mechanism in the LLM field, to optimize LLM output. This mechanism comprises a summary module, a context-awareness module, a short-term memory module, and a long-term memory module. The generation of multiple agents helps increase the diversity of generated results, explore different expression methods, and generate the optimal summary text through an "evaluate-select-suggest-optimize" mechanism.

[0158] S432. After the word tuples of driving experience, physical feelings and emotional reactions are input, keywords are first extracted for RAG retrieval. Each summary module will generate a test report summary through the prompt words of the guided summary and RAG results, and immediately store it in the short-term memory module. After multiple experiments, it is finally determined that when the number of summary modules is 3, the generation effect and resource consumption can be well balanced. After a round of retrieval, the context perception module will receive the above output, and score the output results from two aspects: semantic coherence (evaluating whether the summary text is logically clear, fluent in expression, and whether it can accurately reflect the driver's overall experience) and result accuracy (checking whether the summary content covers all important details and whether a reasonable explanation is provided), and give specific improvement suggestions. Finally, the execution is judged based on the three indicators of "whether the specified number of generations is reached", "whether the predetermined quality standards are met" and "whether the quality change rate of two consecutive rounds is lower than the set value". If the execution is not completed, the summary module is required to make adjustments according to the suggestions. If the execution is completed, the final evaluation report is fed back. Figure 8 shown.

[0159] Example 2:

[0160] This example conducts experimental verification and evaluation on Example 1, as follows:

[0161] S5. Evaluate the performance of each submodule and the criticality of the entire method to the smart car test scenario, and design a control group for comprehensive analysis, as follows:

[0162] S51, comparative validation of physiological signal prediction models;

[0163] S511. High-accuracy physiological signal prediction is the basis for the successful execution of subsequent tasks. In order to verify the effectiveness of the physiological signal prediction model based on the time series GCN-Transformer proposed in the present invention, correlation analysis was completed, and the physiological signal prediction model based on the time series GCN-Transformer was used as the experimental group 1, the algorithm that only used the Transformer architecture and did not include the graph structure was used as the control group 1-1, and the original data set was used without any feature selection or correlation analysis, and the physiological signal prediction model based on the time series GCN-Transformer was used as the control group 1-2.

[0164] S512. In order to further quantitatively analyze the prediction effect of each model, the mean squared error (MSE) is used as the evaluation indicator to measure the average of the squares of the differences between the predicted values ​​and the true values.

[0165] S52. Comparative Verification of Word-Month Matching Models

[0166] S521. The core task of the present invention is to use multimodal data to realize the evaluation of test scenarios. Therefore, the accuracy of word prediction evaluated by integrating chassis data and predicted physiological data is the key point that needs to be verified. In order to verify the influence of feature selection and data type on the prediction results, the experimental group 2 is trained by using a data set containing statistical feature processing and a mixture of predicted results and real results; the control group 2-1 is trained by using a mixture of predicted results and real results without feature processing; the control group 2-2 is trained by feature processing but only using predicted physiological results; the control group 2-3 is trained by feature processing but only using original physiological results.

[0167] S523. Further quantify the application effect of the word-gram matching model based on the confusion matrix. For each word-gram level described in the aforementioned scenario evaluation, respectively construct a confusion matrix and calculate the accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score of the category matching. The calculation method is as follows: Accuracy represents the proportion of all correctly predicted samples to the total number of samples: , the precision rate indicates the proportion of samples predicted to be of a certain category that are actually of that category: , the recall rate indicates the proportion of samples that are actually of a certain category that are correctly predicted to be of that category: , F1 Score represents the harmonic mean of precision and recall: .

[0168] S53, comparative verification of evaluation summary model;

[0169] S531. Summarizing discrete and sequential word units into complete and systematic evaluation texts is the final task of the evaluation model in this paper. In order to verify the effectiveness of the model in generating test conclusions and analyze the impact of RAG and Multi-Self-Refine mechanisms on model performance through ablation studies, the following groups are designed: Experimental Group 3 uses a complete Qwen-based evaluation result summary model (including RAG and Multi-Self-Refine mechanisms), Control Group 3-1 is a Qwen model that only contains the Multi-Self-Refine mechanism, Control Group 3-2 is a Qwen model that only contains RAG, and Control Group 3-3 is a Qwen baseline model that does not contain RAG and Multi-Self-Refine.

[0170] S532. During the experiment, a test report was generated based on the continuous tokens generated by the above process. To evaluate the generation performance of each group, the following evaluation methods were designed: ① Semantic similarity, which measures the semantic similarity between two texts. First, the texts were converted into fixed-length vector representations based on a pre-trained language model (Sentence-BERT). Then, the cosine semantic similarity evaluation method, used in the field of large models, was used to calculate the average semantic similarity between the input tokens and the generated report. This method assesses whether the model accurately incorporates keywords into the generated text. ② Structural diversity, which refers to the richness and variety of vocabulary and syntax in the generated text. Highly diverse text not only avoids monotonous repetition but also better engages readers and conveys more information. This was comprehensively evaluated using lexical diversity and the standard deviation of sentence length distribution. ③ Human evaluation, in which expert evaluators scored the generated texts to assess their quality. This evaluation mainly includes the following four components: 1) Logical clarity: This evaluates whether the generated report is logically coherent and clearly conveys information. 2) Fluency: This examines whether the language flow is natural and fluent, and whether there are any grammatical errors or incoherence. 3) Content Completeness: Ensure that the generated report covers all key terms and that these terms are properly integrated into the report. 4) Result Consistency: Evaluate the consistency between the report and the input terms to ensure that no important information is omitted or irrelevant information is introduced. A scoring table was designed, and multiple experts were invited to score each dimension (1-10 points). The average score was taken as the final score.

[0171] S54, overall test scenario evaluation and comparison verification;

[0172] S541. After verifying each part of the model, the large-model-driven human-like evaluation method for intelligent vehicle test scenarios proposed in the present invention is compared and verified with the traditional scenario evaluation method, and the following groups are designed: Experimental Group 4 is the large-model-driven human-like evaluation method for test scenarios proposed in the present invention; Control Group 4-1 relies only on objective indicators of vehicle dynamics and defines the criticality of the test scenario with acceleration limit values; Control Group 4-2 relies on the driver's rating scale for driving experience.

[0173] S542. Replay batch test scenarios based on simulation software, conduct criticality evaluation using different methods, and have test personnel make self-judgments to evaluate the differences between different methods in scenario criticality evaluation.

[0174] In summary, the present invention realizes multi-dimensional human-like evaluation output that is critical to the test scenario through multi-condition data collection, physiological signal prediction, evaluation word unit matching, and multi-word unit induction and summary. In practical applications, it can effectively improve the evaluation accuracy and reduce data collection and experimental costs.

[0175] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A large-scale model-driven human-like evaluation method for intelligent vehicle test scenarios, characterized by: The following steps are involved: S1, data collection and processing; Design data collection conditions, build data collection equipment, collect data at the test site, and perform preliminary data processing; S2, physiological signal prediction; Conduct correlation analysis on the data, perform hierarchical processing based on data features, design a physiological signal prediction model architecture, and conduct model training; the details are as follows: S21. Data correlation analysis; specifically: S211, a feature screening framework based on time-delay causal inference, which integrates cross-correlation analysis and Granger causality test to achieve interpretable mapping modeling of dynamic features to physiological responses; S212. Use the sliding window cross-correlation algorithm to quantify the dynamic correlation strength. Let the chassis data be X(t) and the physiological data be Y(t). The cross-correlation coefficient with lag τ is defined as: By traversing the lag time window of τ∈[-T,T], the maximum correlation coefficient ρ is extracted max and the corresponding lag time τ max , representing the optimal alignment state of the two signals; through relevant research on human nerve conduction delay, visual persistence, etc., the time T is set to 300ms; S213. Construct a two-variable vector autoregressive model. Assuming that chassis data X(t) has a causal effect on physiological data Y(t), establish a two-order regression equation: Compare the residual sum of squares of the constrained model and the unconstrained model through the F test. If the statistic satisfies: Then reject the null hypothesis and determine that X(t) is the Granger cause of Y(t); and use the ADF test to verify the stationarity of the time series, and perform first-order difference or logarithmic transformation on non-stationary data; S214. Two types of analysis results are integrated based on a dual-threshold mechanism: 1) In cross-correlation analysis, feature data with cross-correlation coefficients and lag times within the lag window are retained; 2) In Granger causality testing, features with only statistical correlation but no causal explanatory power are eliminated. The final selected data are: ECG_rate, ECG_quality, EDA_clean, EDA_phasic, EMG_Amplitude, EMG_Activity; S22. Data feature classification processing; specifically: Take the initial state X under each working condition initial and peak state X max As a benchmark, calculate the data extreme value X in the time series window peak The relative percentage change of the data is calculated, and the data are graded according to the relative percentage change; S23, prediction model architecture design and training; Specifically: S231. Design a cascade prediction method based on temporal graph attention network and Transformer model; it includes four parts: graph network conversion, temporal feature compression, graph attention extraction and Transformer prediction; S232. First, a fully connected graph structure is constructed to represent the relationship between sensors or measurement points. Each chassis dynamics data point is regarded as a node in the graph. Second, the entire time series is divided into time windows of fixed length, and a graph convolution operation is applied at each time step. Noise reduction and smoothing are applied within each window. A graph convolution layer with an attention mechanism is introduced to assign a learnable attention coefficient to each pair of edges between nodes, dynamically learning the importance weights between nodes. Finally, the resulting compact and semantically rich feature representation is fed into a Transformer-based encoder for further processing. The Transformer uses a self-attention mechanism. Finally, the prediction result for each time step is output through a fully connected layer, completing the end-to-end mapping from input to output. S3, evaluate word-gram matching; Process the review text data, construct a lemma graph, perform feature processing on time series data, design a lemma matching model architecture, and perform model training. The details are as follows: S31. Term graph framework construction; specifically: A multi-dimensional driving evaluation system is constructed based on knowledge graph technology. The on-board data acquisition system simultaneously records the driver's real-time voice evaluation and final summary feedback under typical driving conditions, and combines them with vehicle dynamics parameters and physiological signals to form an interpretable semantic mapping network. Through the induction and standardization of the evaluation text, a classification system is ultimately formed, which includes two parts: physiological sensation and chassis response. Physiological sensation includes driving experience and physical sensation. Chassis response includes power response, steering characteristics, braking performance, and suspension feedback. S32, data set construction; specifically: S321. Design special prompt words; S322. During the data collection process, the driver is required to provide real-time evaluation output at each time point and give a final summary after the test is completed; All collected data needs to be time-synchronized to ensure that each record has a consistent timestamp. Subsequently, the raw text data is subjected to NLP preprocessing steps, and the driver evaluation data is mapped to the corresponding time dimension. Global time is used to accurately match evaluation tokens with time series data. A sliding window method is used to construct a continuous "prompt word-chassis dynamics-physiological data-evaluation token" combined dataset. The data within each window represents complete information within a time period. S323. Introduce a feature enhancement strategy. First, extract statistical features and calculate the mean, variance, maximum, and minimum values ​​of the data in the window. Then, calculate the data change trend using the moving average method. The specific process is as follows: for the standardized multidimensional time series data, use 30%, 70%, and 100% of the data length as data breakpoints, and calculate the average value (SMA) of the data in each stage. An SMA change greater than 0.1 is considered an increase, less than -0.1 is considered a decrease, and any changes between them are considered stable. This converts the time series data into a semantic description of the change trend, including nine types: increase, decrease, stability, increase first and then decrease, increase first and then stabilize, decrease first and then increase, decrease first and then stabilize, stabilize first and then increase, and stabilize first and then decrease. S33. Model design and training; specifically: S331. First, time series feature compression is performed. The multi-dimensional time series data is compressed into a 256-dimensional feature vector through a 1D convolutional layer and a global maximum pooling layer. It is then combined with the structured prompt text and input into the BERT-base model to generate a 768-dimensional semantic representation. Six independent classification heads are added after the BERT labeled output layer, corresponding to physical sensation, emotional reaction, power response, steering characteristics, braking performance, and suspension feedback. Each classification head adds a fully connected layer on top of the last hidden state layer of BERT and applies the softmax activation function to generate the probability distribution of each category. During training, the cross-entropy loss function is used as the optimization target. After the input data is encoded by BERT, it is passed to the respective classification head to generate the predicted probability of each category. The predicted probability value is used to calculate the loss and guide the gradient update during the backpropagation process. S332. Design a weighted loss function to assign weights to each category based on the frequency of occurrence of each category. The weights are calculated based on the frequency of occurrence of each category in the dataset. The calculation method is as follows: Where N i is the total number of samples of the i-th subclass; C i N is the number of word categories under this subcategory; ij is the number of occurrences of the j-th word in the subcategory; S333. Define the AdamW optimizer and adopt a weight decay strategy to avoid overfitting. Introduce a learning rate scheduler with linear warmup and cosine annealing to help the model converge more stably. Set a maximum gradient norm limit to prevent gradient explosion. In each epoch, traverse the training data loader, batch process the input data and calculate the loss value. Apply gradient accumulation techniques to improve the effect of small-batch training, and regularly evaluate the performance on the validation set. If there is no performance improvement after three consecutive batches, terminate early. Through the design and training process, the BERT model converts dynamic parameters and physiological signals into word-unit outputs that are consistent with human cognition through feature fusion and semantic mapping. S4, evaluation summary generation; Select a large language model, build a text knowledge base, develop a multi-reflection mechanism, and build a word-meta summary and evaluation model.

2. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 1 is characterized in that: The specific method of S1 is as follows: S11, collection working condition design; S12, collection equipment setup and field collection; S13. Data processing.

3. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 2 is characterized in that: The specific method of S11 is as follows: S111. Based on the test standard procedures of SAE J2944, ISO 7401 and SAE J266, design the collection conditions with prominent single-dimensional indicators and coupled multi-dimensional indicators, i.e. single and coupled in the horizontal, vertical and vertical aspects; S112. Designed operating conditions cover longitudinal, lateral, and vertical motion in vehicle dynamics; test parameters comply with ISO 8855 and ISO 10844 standards. The specific method of S12 is as follows: S121. Use a BioNomadix wireless physiological recorder to synchronously collect the driver's electromyographic, electrocardiographic, and electrodermal signals via three-lead DryPad electrodes with a sampling rate of 2 kHz. Also, collect chassis longitudinal and lateral acceleration at a frequency of 100 Hz via the CAN FD bus protocol. Based on the ASENSING GNSS / RTK combined positioning system, IMU inertial measurement unit and DTU data transmission unit, centimeter-level pose accuracy information is obtained; The vehicle's chassis signals and GNSS posture data are transmitted to the SpeedGoat real-time processing platform via the CAN bus protocol; A timing trigger module sends 5V TTL trigger pulses to achieve microsecond-level synchronization of multiple devices. All data is transmitted to a laptop and stored on a storage medium using MATLAB software and physiological signal processing software. S1222. The test site is a special vehicle testing ground, and the driver shall select a professional test driver from the testing ground; Before each acquisition cycle, a 5-minute rest period was conducted to allow the driver to enter the test in a calm state to collect baseline physiological signals. Three pre-experimental calibration drives were also conducted to eliminate equipment interference. Evaluation data included real-time sentence output during acquisition and a complete summary after acquisition. The specific method of S13 is as follows: S131, using the rising edge of the hardware trigger signal as the time reference, and implementing 100 Hz uniform resampling through the cubic spline interpolation algorithm; S132. Establish a multi-stage filtering process: use Kalman filtering to suppress high-frequency noise on dynamic signals, and implement wavelet threshold noise reduction combined with adaptive filtering algorithm on physiological signals; Among them, the state equation and observation equation in the Kalman filter state space model are: x k =Fx k-1 +This k +w k z k =Hx k +v k Where F is the state transfer matrix; w k is the process noise; H is the observation matrix; v k is the observation noise; B is the control input matrix, which maps the control vector to the change of the state variable; u k is the control vector, which represents the known external input applied to the system at time k; S133, physiological data includes real-time electrical signals and key physiological features; in terms of physiological feature extraction, physiological feature processing is performed based on the NeuroKit2 physiological signal analysis framework; (1) For ECG signals, the clear signal ECG after denoising is extracted clean ;Calculate heart rate ECG rate , reflecting the body's stress level or exercise state; evaluating the quality index of ECG data ECG quality ; and extract the peak positions of the QRS complex, P wave and T wave ECG r_peaks , ECG p_peaks and ECG t_peaks , describing the electrical activity phase of the atria and ventricles in the entire cardiac cycle, and is used to detect cardiac rhythm abnormalities; (2) For the EDA signal, first remove the noise to obtain a smooth EDA signal EDA clean ; Then separate the slow-varying component EDA tonic , reflecting the tension of the autonomic nervous system; capturing the rapidly changing component EDA phasic , related to transient emotional responses; quantifies the transient conductance increase triggered by a stimulus, including the onset time SCR onsets , peak time SCR peaks , high SCR height , amplitude SCR amplitude , rise time SCR risetime , recovery time SCR recoverytime , used to assess the individual's sensitivity to specific events; (3) For EMG signals, muscle activity features are extracted after filtering and envelope detection, including the processed EMG signal EMG clean ; The intensity of muscle contraction EMG amplitude , duration EMG activity , and the starting point of muscle activity EMG onsets and end point EMG offsets .

4. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 1 is characterized in that: The specific method of S4 is as follows: S41, Large Language Model Selection; S42, knowledge base retrieval enhancement development; S43. Multi-reflection mechanism design.

5. The method for human-like evaluation of intelligent vehicle test scenarios driven by a large model according to claim 4 is characterized in that: The specific method of S41 is as follows: The Qwen model was selected as the core LLM model for summarizing the evaluation results; The specific method of S42 is as follows: S421. Introducing RAG retrieval enhancement technology, which helps the model obtain more background information and improve the depth of summary by combining external knowledge base and real-time retrieval; S422. Collect textual materials, including the following: ① Theoretical literature and technical reports covering vehicle performance, driving experience, and autonomous driving test evaluations; ② A systematic evaluation database, collecting structured driver evaluation data, including detailed feedback from different driving scenarios; ③Domain expertise, including training materials and tutorials, provides professional guidance; ④ Historical test records, including detailed reports and representative cases of similar tests in the past, as reference templates; S423. Targeted processing is performed on the collected text in various formats. During the text processing phase, irrelevant characters, punctuation marks, and special formats are removed. During the text segmentation phase, the segmentation block size is limited to 200 characters, and the number of characters shared between two blocks is set to 20 to maintain contextual coherence. The segmented text is vectorized. A knowledge vector library is established based on Chroma. For input problems, the vectorized text is searched based on a similar text evaluation method based on maximum marginal relevance to increase the diversity of search results and avoid duplication of information. The retrieved relevant text is combined with the query prompt word template to form a complete prompt word, which is input into the language model to organize and output the query results. The specific method of S43 is as follows: S431. Based on the Self-Refine mechanism in the LLM field, a Multi-Self-Refine mechanism is designed to optimize the output of LLM. The mechanism includes a summary module, a context-aware module, a short-term memory module, and a long-term memory module. S432. After inputting word tuples representing driving experience, physical sensation, and emotional response, keywords are first extracted for RAG retrieval. Each summary module generates a test report summary using prompt words that guide the summary and the RAG results, and immediately stores the summary in the short-term memory module. After one round of retrieval, the context-aware module receives the output, scores the output results based on semantic coherence and result accuracy, and provides specific improvement suggestions. Finally, the completion of the test is determined based on three indicators: whether the specified number of generation times has been reached, whether the predetermined quality standard has been met, and whether the rate of change in quality over two consecutive rounds is lower than the set value. If the test has not been completed, the summary module is required to make adjustments based on the suggestions. If the test has been completed, the final evaluation report is output as feedback.

Citation Information

Patent Citations

  • Machine dialogue capability improving method and system based on social simulation system

    CN118468930A

  • Intelligent automobile virtual simulation test method based on large language model

    CN118586281A