An automatic driving trust state early warning method and system based on multi-modal features, a terminal, and a storage medium

By acquiring and fusing driver eye-tracking data, driving behavior, and individual characteristics, and using a trust prediction model to predict and provide feedback on trust status, the problem of the inability to accurately predict driver trust status in existing technologies is solved, thus achieving real-time safety improvement.

CN120697779BActive Publication Date: 2025-11-21SHENZHEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511212528.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-21
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

In existing technologies, subjective scale tools cannot accurately predict the driver's trust level, resulting in an inability to adapt to rapidly changing driving environments in real time and provide effective warnings.

Method used

By acquiring the driver's eye movement data, driving behavior data, and individual characteristic information, feature extraction and fusion are performed. A trust prediction model is used to predict the trust status, and dynamic feedback is provided based on the prediction results.

Benefits of technology

It enables accurate prediction and dynamic assessment of the driver's trust status, provides real-time feedback and reminders, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120697779B_ABST
    Figure CN120697779B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of automatic driving and discloses an automatic driving trust state early warning method and system based on multi-modal features, a terminal and a storage medium, the method comprises the following steps: acquiring eye movement data signals, driving behavior data and individual feature information of a driver, performing feature extraction on the eye movement data signals, the driving behavior data and the individual feature information to obtain eye movement features, driving features and individual features; performing feature fusion on the eye movement features, the driving features and the individual features to obtain to-be-tested multi-modal features; inputting the to-be-tested multi-modal features into a trust prediction model for prediction to obtain a trust state prediction result; and performing dynamic feedback on the driver according to a trust level of the trust state prediction result and the individual features to obtain a dynamic feedback reminding result. The trust state is accurately predicted through the trust prediction model, the dynamic evaluation of the trust state of the driver is realized, and the driver is fed back according to the trust state.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic driving, in particular to an automatic driving trust state early warning method and system based on multi-modal features, a terminal and a computer readable storage medium. BACKGROUND

[0002] Intelligent driving technology can enable vehicles to autonomously complete core driving operations such as acceleration, deceleration, steering, and path planning in various road situations. Although intelligent driving technology has developed rapidly in recent years, it is still a long way to achieve fully automated driving in all working conditions due to the uncertainty of complex traffic environments and the dynamic changes of various driving tasks. Therefore, in the foreseeable future, intelligent driving systems will remain in the human-machine collaborative control stage. In this stage, the driver and the intelligent system need to jointly participate in vehicle control, share control rights, and achieve optimal configuration of information processing, environmental perception, and decision-making processes to ensure the safety of driving tasks. Under this collaborative framework, the driver's trust level in the automatic driving system is a key factor affecting system safety and interaction efficiency, directly affecting the driver's behavior intervention, risk decision-making, and emergency response capability. Existing research shows that if the driver's trust level in the system deviates, it may cause potential safety hazards. Specifically, excessive trust can lead to the driver's disengagement from the perception of the surrounding environment, reducing their ability to respond to emergencies, while insufficient trust can lead to excessive intervention, continuous monitoring, and increased cognitive load on the system, leading to fatigue accumulation and even traffic accidents caused by human operation errors. Currently, trust assessment tools mainly rely on traditional subjective scale tools. Although subjective scale tools are widely used in psychology research, they have poor real-time performance, strong subjectivity, and limited application scenarios, making it difficult to adapt to rapidly changing driving environments.

[0003] Currently, trust assessment tools mainly rely on traditional subjective scale tools, but there are problems such as poor real-time performance, and the trust assessment results obtained are difficult to adapt to rapidly changing driving environments, making it impossible to accurately predict the user's trust state through subjective scale tools and to warn the driver's operation behavior according to the trust state, which is a problem that needs to be solved urgently.

[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY

[0005] The main purpose of the present application is to provide an automatic driving trust state early warning method and system based on multi-modal features, a terminal and a computer readable storage medium, which aims to solve the problem that the user's trust state cannot be accurately predicted through subjective scale tools in the prior art, and the operation behavior of the driver is warned according to the trust state.

[0006] In order to achieve the above object, the application provides a multi-modal feature-based automatic driving trust state early warning method, which comprises the following steps:

[0007] Obtaining eye movement data signals, driving behavior data and individual characteristic information of a driver, extracting features from the eye movement data signals, the driving behavior data and the individual characteristic information to obtain eye movement features, driving features and individual features;

[0008] Fusing the eye movement features, the driving features and the individual features to obtain to-be-tested multi-modal features, inputting the to-be-tested multi-modal features into a trust prediction model for prediction to obtain a trust state prediction result;

[0009] According to the trust level of the trust state prediction result and the individual features, dynamically feeding back to the driver to obtain a dynamic feedback reminder result.

[0010] Optionally, the multi-modal feature-based automatic driving trust state early warning method, wherein the obtaining of the eye movement data signals, the driving behavior data and the individual characteristic information of the driver, the extracting of features from the eye movement data signals, the driving behavior data and the individual characteristic information to obtain the eye movement features, the driving features and the individual features specifically comprises:

[0011] Obtaining eye movement data signals of the driver collected by an eye movement tracking device, driving behavior data of the driver collected by a vehicle-mounted controller and individual characteristic information of the driver collected in advance;

[0012] Processing the eye movement data signals in sequence to obtain eye movement features, including eye movement point processing, fixation point extraction, pupil processing and blink recognition processing;

[0013] Performing interpolation filling and outlier rejection on the driving behavior data to obtain target driving behavior data, and performing time sequence alignment on the eye movement data signals and the target driving behavior data to obtain driving features;

[0014] Performing normalization processing and one-hot encoding processing on the individual characteristic information to obtain individual features;

[0015] The individual characteristic information includes gender information, age information, driving experience information, emotional state information and individual trait information.

[0016] Optionally, the multi-modal feature-based automatic driving trust state early warning method, wherein the eye movement point processing comprises interpolation processing and smoothing processing;

[0017] The eye movement features include fixation time, total fixation times, average pupil diameter and saccade times.

[0018] The eye movement data signal is sequentially subjected to eye movement point processing, gaze point extraction, pupil processing and blink recognition processing to obtain eye movement features, specifically including:

[0019] The eye movement data signal is subjected to interpolation processing and smoothing processing according to a linear interpolation formula and a preset interval length to obtain a gaze time;

[0020] The eye movement data signal is subjected to gaze point extraction according to a preset angular velocity threshold to obtain a total number of gazes;

[0021] The eye movement data signal is subjected to pupil processing according to a preset pupil diameter to obtain an average pupil diameter;

[0022] The eye movement data signal is subjected to blink recognition processing according to a preset time threshold to obtain a number of saccades.

[0023] Optionally, the automatic driving trust state early warning method based on multi-modal features, wherein the eye movement features, the driving features and the individual features are fused to obtain a to-be-tested multi-modal feature, and the to-be-tested multi-modal feature is input into a trust prediction model for prediction to obtain a trust state prediction result, specifically including:

[0024] The eye movement features, the driving features and the individual features are fused to obtain multi-modal features and a to-be-tested multi-modal feature;

[0025] The multi-modal features are divided according to a preset division ratio to obtain a training set and a test set, and a machine learning model is trained according to the training set to obtain a trust prediction model;

[0026] The to-be-tested multi-modal feature is input into the trust prediction model for prediction to obtain a trust state prediction result;

[0027] The test set is used for performance evaluation of the trust prediction model.

[0028] Optionally, the automatic driving trust state early warning method based on multi-modal features, wherein the machine learning model includes any one of a random forest model, an extreme random tree model, a limit gradient boosting model and a lightweight gradient boosting machine model;

[0029] The trust prediction model includes a first trust prediction model, a second trust prediction model, a third trust prediction model and a fourth trust prediction model;

[0030] The multi-modal features are divided according to a preset division ratio to obtain a training set and a test set, and a machine learning model is trained according to the training set to obtain a trust prediction model, specifically including:

[0031] divide the multi-modal features according to a preset division ratio to obtain a training set and a test set;

[0032] extract the training set multiple times through a bootstrap resampling method to obtain multiple target training sets, determine multiple decision trees of the multiple target training sets according to a CART algorithm, generate a random forest model according to the multiple decision trees, train the random forest model according to the multiple target training sets, and obtain a first trust prediction model;

[0033] determine multiple feature sets of the training set according to a preset double randomness mechanism, generate an extreme random tree model according to the multiple feature sets, train the extreme random tree model according to the training set, and obtain a second trust prediction model;

[0034] train the extreme gradient boosting model according to the training set and a preset node division strategy to obtain a third trust prediction model;

[0035] train the light gradient boosting machine model according to a preset model parameter and the training set to obtain a fourth trust prediction model;

[0036] test the performance of the first trust prediction model, the second trust prediction model, the third trust prediction model, and the fourth trust prediction model to obtain a trust prediction model.

[0037] Optionally, the automatic driving trust state early warning method based on multi-modal features, wherein the dynamic feedback reminder result includes voice and graphical interface feedback result, atmosphere lamp feedback result, smell prompt feedback result, and tactile feedback result.

[0038] The trust level includes any one of a first trust, a second trust, and a third trust.

[0039] The dynamic feedback of the driver according to the trust level of the trust state prediction result and the individual characteristics obtains a dynamic feedback reminder result, specifically including:

[0040] divide the trust state prediction result according to a preset threshold strategy to obtain any one of the first trust, the second trust, and the third trust.

[0041] The dynamic feedback of the driver according to any one of the first trust, the second trust, and the third trust, and the gender information, the age information, the driving experience information, the emotional state information, and the individual trait information obtains voice and graphical interface feedback results, atmosphere lamp feedback results, smell prompt feedback results, and tactile feedback results.

[0042] Optionally, the automatic driving trust state early warning method based on multi-modal features, wherein the dynamic feedback to the driver according to any one of the first trust, the second trust and the third trust, and the gender information, the age information, the driving experience information, the emotional state information and the individual trait information, obtains voice and graphical interface feedback results, atmosphere lamp feedback results, smell prompt feedback results and tactile feedback results, and then further comprises:

[0043] obtaining a trust level of the driver, if the trust level is in the first trust, then feedback adjusting the trust level according to the voice and graphical interface feedback results and the smell prompt feedback results, to obtain a target trust level;

[0044] if the trust level is in the third trust, then reminding the driver according to the tactile feedback results and the atmosphere lamp feedback results, so as to drive safely.

[0045] In addition, to achieve the above-mentioned purpose, the present application also provides an automatic driving trust state early warning system based on multi-modal features, wherein the automatic driving trust state early warning system based on multi-modal features:

[0046] a feature extraction module for obtaining eye movement data signals, driving behavior data and individual characteristic information of the driver, extracting features of the eye movement data signals, the driving behavior data and the individual characteristic information, to obtain eye movement features, driving features and individual features;

[0047] a trust state prediction module for fusing the eye movement features, the driving features and the individual features, to obtain a to-be-tested multi-modal feature, inputting the to-be-tested multi-modal feature into a trust prediction model for prediction, to obtain a trust state prediction result;

[0048] a trust state feedback module for dynamically feeding back to the driver according to the trust level of the trust state prediction result and the individual characteristics, to obtain a dynamic feedback reminder result.

[0049] In addition, to achieve the above-mentioned purpose, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores an automatic driving trust state early warning program based on multi-modal features, and the automatic driving trust state early warning program based on multi-modal features, when executed by a processor, realizes the steps of the automatic driving trust state early warning method based on multi-modal features as described above.

[0050] In the present application, the eye movement data signal, driving behavior data and individual characteristic information of the driver are obtained, the eye movement data signal, driving behavior data and individual characteristic information are feature extracted to obtain eye movement features, driving features and individual characteristics; the eye movement features, the driving features and the individual characteristics are feature fused to obtain a to-be-tested multi-modal feature, the to-be-tested multi-modal feature is input into a trust prediction model for prediction to obtain a trust state prediction result; the driver is dynamically fed back according to the trust level of the trust state prediction result and the individual characteristics to obtain a dynamic feedback reminder result. The trust state is accurately predicted by the trust prediction model, the dynamic evaluation of the trust state of the driver is realized, and the driver is fed back according to the trust state. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 is a flowchart of a preferred embodiment of the automatic driving trust state early warning method based on multi-modal features of the present application;

[0052] Figure 2 is a structure diagram of a preferred embodiment of the automatic driving trust state early warning system based on multi-modal features of the present application;

[0053] Figure 3 is a structure diagram of a preferred embodiment of the terminal of the equipment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical scheme and advantages of the present application more clear and definite, the present application will be further described in detail below with reference to the drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0055] Currently, the tool for trust evaluation mainly relies on the traditional subjective scale tool, but there is poor real-time performance, and the trust evaluation results obtained are difficult to adapt to the rapidly changing driving environment, which leads to the inability to accurately predict the trust state of the user through the subjective scale tool, and to warn the operation behavior of the driver according to the trust state, therefore, a kind of automatic driving trust state early warning method based on multi-modal features is needed, the trust state is accurately predicted by trust prediction model, the dynamic evaluation of the trust state of the driver is realized and the driver is fed back according to the trust state.

[0056] The automatic driving trust state early warning method based on multi-modal features according to the preferred embodiment of the present application, as shown in Figure 1 The automatic driving trust state early warning method based on multi-modal features includes the following steps:

[0057] The step S10 comprises:

[0058] The step S10 comprises:

[0059] The step S11 comprises:

[0060] The step S12 comprises:

[0061] The step S13 comprises:

[0062] The step S14 comprises:

[0063] Specifically, the eye movement data signal of the driver collected by the eye movement tracking device (using a head-mounted eye tracker or using an infrared-based depth camera, which can be collected through a camera device integrated in the vehicle in actual application, such as the gaze behavior of the driver to the road environment, the human-vehicle interaction interface and other driving related information during driving, and the signal is transmitted to the data collection device in real time through Bluetooth or vehicle-mounted sensor network), the driving behavior data of the driver collected by the vehicle-mounted controller (including but not limited to steering wheel angle, acceleration and deceleration behavior, brake operation, lane keeping and other control information) and the individual characteristic information of the driver collected in advance (for example: 1, gender information (male and female); 2, age information, indicating the natural age of the user (unit: years); 3, driving experience information, including driving time, driving frequency and common driving scene types (such as urban roads, highways or rural roads, etc.); 4, emotional state information, obtained through a self-rating scale or a non-contact facial expression recognition module, divided into three categories of positive emotion, neutral emotion and negative emotion; 5, individual trait information, divided into cautious type and adventurous type, representing the basic tendency of the user to risk avoidance or risk taking in driving behavior, which can be obtained by measuring the individual personality characteristics scale), the eye movement data signal is sequentially subjected to eye movement point processing, fixation point extraction, pupil processing and blink recognition processing to obtain eye movement characteristics, the driving behavior data is subjected to interpolation filling and outlier rejection to obtain target driving behavior data, the eye movement data signal and the target driving behavior data are time-aligned to obtain driving characteristics (driving behavior data is collected simultaneously in the vehicle operating system, including steering wheel angle, accelerator and brake operation, vehicle speed change, lane deviation and takeover frequency, after interpolation filling, outlier rejection and time alignment, the stable driving characteristics are extracted), the individual characteristic information is subjected to normalization processing and one-hot encoding processing to obtain individual characteristics, wherein the individual characteristic information includes gender information, age information, driving experience information, emotional state information and individual trait information.

[0064] In the embodiment, the individual characteristic information is processed before the user participates in the intelligent driving task to ensure that the static variables can participate in the modeling process of the trust state in the subsequent model efficiently and stably. Specifically, the gender is converted into a binary vector by one-hot encoding as a binary classification variable; the age is normalized by the minimum-maximum method as a continuous variable to reduce the influence of its numerical scale on the model weight learning; the driving experience is composed of driving years, driving frequency and typical driving scenarios, among which the years information is normalized, the frequency and scenario type are respectively vectorized by one-hot encoding; the emotional state is obtained by self-assessment questionnaire or facial expression recognition module and divided into positive, neutral and negative categories, which are also processed by one-hot encoding to avoid category feature sorting errors; the individual trait information is evaluated by a composite psychological scale (such as the widely used Impulsivity Scale and Stimulus Seeking Scale) and divided into cautious and adventurous types, which are represented by binary one-hot encoding. Through the above encoding and normalization operations, the system uniformly converts the multi-source heterogeneous individual characteristic information into a numerical vector form, which not only improves the modeling ability of the model for individual differences, but also enhances the prediction performance and generalization stability.

[0065] The step S12 comprises:

[0066] Step S121, performing interpolation processing and smoothing processing on the eye movement data signal according to a linear interpolation formula and a preset interval length to obtain a fixation time;

[0067] Step S122, extracting a fixation point from the eye movement data signal according to a preset angular velocity threshold to obtain a total fixation count;

[0068] Step S123, performing pupil processing on the eye movement data signal according to a preset pupil diameter to obtain an average pupil diameter;

[0069] Step S124, performing blink recognition processing on the eye movement data signal according to a preset time threshold to obtain a saccade count.

[0070] Specifically, the eye movement data signal is subjected to interpolation processing and smoothing processing (noise reduction processing) according to a linear interpolation formula and a preset interval length, to obtain a fixation time, the eye movement data signal is subjected to fixation point extraction according to a preset angular velocity threshold, to obtain a total fixation number, the eye movement data signal is subjected to pupil processing (linear interpolation processing using a linear interpolation method after abnormal value elimination, and signal noise reduction using a sliding mean filter method) according to a preset pupil diameter, to obtain an average pupil diameter, the eye movement data signal is subjected to blink recognition processing according to a preset time threshold (including a maximum time threshold and a minimum time threshold), to obtain a saccade number, the eye movement point processing includes interpolation processing and smoothing processing, and the eye movement features include the fixation time, the total fixation number, the average pupil diameter, and the saccade number.

[0071] In the embodiment, the maximum interval length is set to 75 ms, and a signal less than the maximum interval length is missing, data compensation is performed using a linear interpolation method, and data noise reduction is completed using a sliding mean filter method, wherein the linear interpolation formula is as follows:

[0072] ;

[0073] wherein, x0 represents the abscissa of the first end point of the two end points of the interpolation interval, y0 represents the ordinate of the first end point of the two end points of the interpolation interval, x1 represents the abscissa of the second end point of the two end points of the interpolation interval, y1 represents the ordinate of the second end point of the two end points of the interpolation interval, xi represents the abscissa of an arbitrary interpolation point of the interpolation interval, yi represents the ordinate of the arbitrary interpolation point of the interpolation interval, the purpose is to use a straight line formed by the two end points of the interpolation interval to approximately replace the point that needs to be interpolated, and a sliding mean filter is used for smoothing processing of an input discrete time signal to reduce noise interference and signal fluctuation, wherein the implementation steps of the sliding mean filter are as follows: the number of input sample points is set to n, the output is y, and a sliding mean filter with a factor of m is used, the sliding filter first calculates the mean value of the first n sample points as the initial filter output:

[0074] ;

[0075] wherein, y(k) represents the output value at the kth sample point, x0 represents the first sample point, x1 represents the second sample point, i represents the index of the sample point, x(k) represents the abscissa of the kth sample point.​​​​ sampling points, This represents the size of the sliding window, and then the filter is updated (smoothed) to obtain the gaze time for each new input signal point. Update the filter output using a recursive formula:

[0076] ;

[0077] in, Indicates the filter at the 1st The output value at each sampling point express The output value at each sampling point Indicates the input signal at the th The value of each sampling point, Indicates the input signal at the th The value of each sampling point.

[0078] As an example, fixation point extraction: The angular velocity calculation window length is set to 20ms, and the angular velocity threshold is 30° / s. Angular velocities greater than 30° / s are classified as saccades, and those less than or equal to 30° / s are classified as fixations. The maximum time threshold between fixations is set to 75ms, and the maximum angular velocity threshold between fixations is set to 0.5° / s. When the interval between two adjacent fixations is less than 75ms or the interval angle is less than 0.5° / s, the two adjacent fixations are merged into one. The minimum fixation time threshold (preset pupil diameter) is set to 60ms. Fixations with fixation durations less than the minimum fixation time threshold are deleted, yielding the total number of fixations. Pupil processing: This includes data interpolation and signal denoising. The minimum pupil diameter is set to 2mm. Data smaller than the minimum pupil diameter are defined as outliers and removed. After outlier removal, linear interpolation is performed, followed by moving average filtering to denoise the signal, yielding the average pupil diameter. Blink processing: The maximum time threshold is set to 350ms and the minimum time threshold is 75ms. When both eyes close simultaneously and the closing time is between 75ms and 350ms (preset time threshold), it is classified as a blink. After the feature extraction module completes the preprocessing of the eye movement data signal, it extracts eye movement features and uses the correlation coefficient method to obtain eye movement features that are highly correlated with mental workload. Finally, the eye movement features include the fixation time, the total number of fixations, the average pupil diameter, and the number of saccades.

[0079] Step S20: The eye-tracking features, driving features, and individual features are fused to obtain the multimodal features to be tested. The multimodal features to be tested are then input into the trust prediction model for prediction to obtain the trust state prediction result.

[0080] Step S20 includes:

[0081] Step S21, the eye movement features, the driving features and the individual features are fused to obtain multi-modal features and to-be-tested multi-modal features;

[0082] Step S22, the multi-modal features are divided according to a preset division ratio to obtain a training set and a test set, a machine learning model is trained according to the training set to obtain a trust prediction model;

[0083] Step S23, the to-be-tested multi-modal features are input into the trust prediction model for prediction to obtain a trust state prediction result; the test set is used for performance evaluation of the trust prediction model.

[0084] Specifically, the eye movement features, the driving features and the individual features are fused to obtain multi-modal features and to-be-tested multi-modal features (multi-modal features (sample data) are constructed by using variables such as eye movement features, driving behavior features and individual features screened from a feature processing module, for training and testing of a model), the multi-modal features are divided according to a preset division ratio (sample data is divided into a training set and a test set according to a 7:3 ratio, and a 5-fold cross-validation strategy is used to verify the training set), to obtain the training set and the test set, a machine learning model (a random forest model (random forest), an extreme random tree model (Extra Trees), an extreme gradient boosting model (XGBoost, eXtreme Gradient Boosting) and a light gradient boosting machine (lightgbm, Light Gradient Boosting Machine) is trained according to the training set to obtain a trust prediction model, the to-be-tested multi-modal features are input into the trust prediction model for prediction to obtain a trust state prediction result; the test set is used for performance evaluation of the trust prediction model.

[0085] In this embodiment, in a machine learning task, accuracy, recall rate and precision rate and the like can be used to measure the performance of the trust prediction model in a classification task. A confusion matrix describes the comparison between the prediction results of the trust prediction model and the true labels, and contains four parts of true positive examples (TP), true negative examples (TN), false positive examples (FP) and false negative examples (FN). Accuracy: the overall classification accuracy of the model is measured by calculating the proportion of the number of samples correctly classified by the trust prediction model to the total number of samples. The higher the accuracy, the more consistent the classification results of the model with the true labels. The calculation formula of the accuracy (Acc) is as follows:

[0086] ​​​​ ;

[0087] Recall: the ratio of the number of samples correctly predicted as positive by the model to the number of true positive samples, used to measure the model's ability to identify positive samples, with a value range of 0 to 1, the higher the value, the more comprehensive and accurate the trust prediction model can find all true positive samples, and reduce the missed cases, the calculation formula of the recall rate (Recall) is as follows:

[0088] ;

[0089] Precision: the proportion of true positive samples in the results predicted as positive samples by the model, reflecting the accuracy of the model in predicting positive samples. The value is between 0 and 1, the closer to 1, the lower the misjudgment rate of the model in predicting positive samples, and the more accurate the identification of positive samples, the calculation formula of the precision rate (Precision) is as follows:

[0090] ;

[0091] The step S22 comprises:

[0092] Step S221, dividing the multi-modal features according to a preset division ratio to obtain a training set and a test set;

[0093] Step S222, multiple extraction of the training set by a bootstrap resampling method to obtain a plurality of target training sets, determination of a plurality of decision trees of the plurality of target training sets according to a CART algorithm, generation of a random forest model according to the plurality of decision trees, training of the random forest model according to the plurality of target training sets, and obtaining of a first trust prediction model;

[0094] Step S223, determining a plurality of feature sets of the training set according to a preset double randomness mechanism, generating an extreme random tree model according to the plurality of feature sets, training the extreme random tree model according to the training set, and obtaining a second trust prediction model;

[0095] Step S224, training the extreme gradient boosting model according to the training set and a preset node division strategy to obtain a third trust prediction model;

[0096] Step S225, training the light gradient boosting machine model according to a preset model parameter and the training set to obtain a fourth trust prediction model;

[0097] Step S226, performance testing of the first trust prediction model, the second trust prediction model, the third trust prediction model and the fourth trust prediction model to obtain a trust prediction model. ​​

[0098] Specifically, the multi-modal features are divided according to a preset division ratio to obtain a training set and a test set, multiple target training sets are obtained by multiple extraction of the training set through a bootstrap resampling method, multiple decision trees of the multiple target training sets are determined according to a CART (Classification and Regression Trees) algorithm (a Gini (Gini Impurity) index based on the CART algorithm is introduced, a split node is performed according to a classification logic with the smallest numerical value of the Gini index until a threshold value or a sample size is insufficient to support a resampling technique), a random forest model is generated according to the multiple decision trees (each target training set can generate a complete decision tree, and multiple decision trees constitute a random forest), the random forest model is trained according to the multiple target training sets to obtain a first trust prediction model, multiple feature sets of the training set are determined according to a preset double randomness mechanism (in each split node, a subset is randomly selected from the training set to constitute multiple feature sets), an extremely random tree model is generated according to the multiple feature sets, the extremely random tree model is trained according to the training set to obtain a second trust prediction model, a third trust prediction model is obtained by training the extreme gradient boosting model according to the training set and a preset node division strategy, a fourth trust prediction model is obtained by training the light gradient boosting machine model according to a preset model parameter and the training set (a small batch training method is adopted, and the training set is divided into multiple small batches), performance tests are performed on the first trust prediction model, the second trust prediction model, the third trust prediction model and the fourth trust prediction model to obtain a trust prediction model (the performances of four mainstream machine learning models, i.e., the random forest model, the extremely random tree model, the extreme gradient boosting model and the light gradient boosting machine model, in the driver trust state prediction task are systematically evaluated to obtain the trust prediction model), the machine learning model includes any one of the random forest model, the extremely random tree model, the extreme gradient boosting model and the light gradient boosting machine model, and the trust prediction model includes the first trust prediction model, the second trust prediction model, the third trust prediction model and the fourth trust prediction model.

[0099] In this embodiment, the training process of the random forest model: through bootstrap resampling technology, for example, randomly drawing samples from the number of training sets to generate target training sets, each target training set corresponds to a decision tree, and generally the number of randomly selected target training sets is much smaller than To ensure that the data can not participate in training error calculation, then, in each node split decision tree, from the attribute set A, m attributes are randomly selected to form a subset of attributes, the optimal attribute in each subset is used to split the node, in the specific construction of each decision tree, the Gini index based on (CART, Classification and Regression Trees) algorithm is introduced, the classification logic with the minimum Gini index value is used to split the node, until the threshold or the sample size is insufficient to support the resampling technique, and the first trust prediction model is obtained, wherein the calculation formula of Gini is as follows:

[0100] ;

[0101] Wherein, Gini coefficient, Total number of categories, The number of samples of the first Class.

[0102] As an example, the training process of the extreme random tree model: in the partition process of each tree node, the extreme random tree model introduces a double randomness mechanism, feature random selection: at each split node, a subset of features is randomly selected from all features to form a plurality of feature sets, and a partition threshold is randomly generated: for each of the candidate features, instead of statistically optimizing the partition point according to the distribution of the training set, a plurality of candidate partition thresholds are randomly sampled within the value range of the feature, and then, for each candidate feature and its corresponding random threshold combination, the classification effect is calculated, and the Gini Impurity or information gain is usually used as an evaluation index, and the optimal partition method is selected to complete the split of the current node. According to the plurality of feature sets, an extreme random tree model is generated, the extreme random tree model is trained according to the training set, and a second trust prediction model is obtained. The training process of the extreme random tree model further enhances the diversity and robustness of the model, effectively alleviating the risk of overfitting. The growth process of each tree in the extreme random tree model is controlled by a plurality of hyperparameters, such as: the maximum depth of the decision tree (max_depth), which is used to prevent the tree structure from expanding too much, causing the model complexity to expand, and is usually set to 30 or an appropriate value, the minimum number of samples contained in an internal node (min_samples_split), if the number of samples is less than the value, the partition is not performed, the minimum number of samples required on the leaf node (min_samples_leaf), which prevents the tree from generating too many isolated nodes and improves the stability of the model, the number of trees that constitute the model (n_estimators), the more the number of trees, the more stable the model, and the value is usually in the range of 100-500. These parameters jointly determine the structural complexity and modeling ability of each tree. In specific implementation, cross-validation can be used to automatically optimize the parameters to balance the prediction performance and computational efficiency.

[0103] Further, the training process of the extreme gradient boosting model adopts an additive modeling method to construct a prediction function. The model continuously fits the residual of the previous round of model through iterative learning to form a prediction function (f(x) = f0+ f1(x) + f2(x) +... + fn(x)) , the calculation formula of the prediction function is as follows:

[0104] ;

[0105] Wherein, represents a regression tree obtained in the i-th round of learning, represents a function space of CART (Classification and Regression Trees), represents the total number of categories;

[0106] ​The extreme gradient boosting model introduces a second-order derivable objective function design. On the basis of traditional loss functions (such as square loss and logarithmic loss), a model complexity control term is introduced. The regularization term is used to suppress the excessive complexity of the tree structure, effectively alleviating the problem of overfitting. To improve the optimization efficiency, the extreme gradient boosting model uses a second-order Taylor expansion to approximate the loss function. The first-order gradient (Gradient) and the second-order derivative (Hessian) are introduced for optimization update, so that the algorithm can achieve a good balance between accuracy and efficiency. The extreme gradient boosting model introduces a node partition strategy based on information gain (Gain). The position of the split point is determined by evaluating the contribution of the loss function after feature partitioning. The specific process is as follows: calculate the gradient gain of all possible split points for each candidate feature of the training set; select the feature and split point with the maximum gain as the optimal partition of the current node; repeat the above process until the termination condition (such as maximum depth, minimum sample size, etc.) is met to generate a complete tree structure. The node partition strategy combined with engineering optimization methods (such as block structure caching, feature parallelism, candidate partition point sorting, etc.) significantly improves the splitting efficiency and model performance. The third trust prediction model is obtained by training the extreme gradient boosting model according to the candidate features of the training set and the preset node partition strategy.

[0107] After training, the extreme gradient boosting model combines the regression trees constructed in each round into the optimal prediction extreme gradient boosting model. For a given input of the training set, the model will output the prediction value through each sub-tree in turn and perform cumulative summation to output the final trust state prediction result.

[0108] In this embodiment, the lightweight gradient boosting machine model training process: the model is constructed by iteratively training multiple decision trees. In terms of model parameter (preset model parameter) setting, the preset model parameters include: setting the number of trees (num_boost_round) to 200 to balance the model complexity and training efficiency; the maximum depth of the tree (max_depth) is limited to 6 to avoid overfitting of the model to the training data, while ensuring that the model can effectively capture data features, and using histogram algorithm and one-side gradient sampling (GOSS, Gradient-based One-Side Sampling) technology to reduce computational load and memory consumption, improve model generalization ability, and set a small learning rate (learning_rate=0.05) to make the lightweight gradient boosting machine model more robustly converge during training. The early stopping method (early_stopping_rounds=20) is used to monitor the validation set indicators. If the evaluation indicators on the validation set do not improve for 20 consecutive rounds, the training is stopped to prevent overfitting. A small batch training method is used to divide the training data into multiple small batches, which are input into the model for training in turn. In each round of training, the lightweight gradient boosting machine model calculates the gradient according to the difference between the current model prediction result and the true label, and then builds a new decision tree to fit the gradient, constantly updating the model parameters and gradually improving the model performance to obtain the fourth trust prediction model. During training, multi-classification evaluation indicators (such as multi-classification accuracy and macro average score) are used to monitor model performance, and the loss function is used to measure the difference between the multi-classification prediction probability and the true label.

[0109] Step S30, dynamically feedback the driver according to the trust level of the trust state prediction result and the individual characteristics, to obtain a dynamic feedback reminder result.

[0110] The step S30 includes:

[0111] Step S31, dividing the trust state prediction result according to a preset threshold strategy to obtain any one of the first trust, the second trust and the third trust;

[0112] Step S32, dynamically feedback the driver according to any one of the first trust, the second trust and the third trust, and the gender information, the age information, the driving experience information, the emotional state information and the individual trait information, to obtain voice and graphical interface feedback results, atmosphere lamp feedback results, smell prompt feedback results and tactile feedback results.

[0113] Specifically, the trust state prediction result is divided according to a preset threshold strategy, to obtain any one of the first trust (low trust), the second trust (moderate trust), and the third trust (excessive trust), and the driver is dynamically fed back according to any one of the first trust, the second trust, and the third trust and the gender information, the age information, the driving experience information, the emotional state information, and the individual trait information, to obtain a voice and graphical interface feedback result, an atmosphere lamp feedback result, an odor prompt feedback result, and a tactile feedback result (for example, in a high-risk situation such as tunnel lane changing and congestion obstacle avoidance, the system can increase the “third trust” determination sensitivity to strengthen the risk prompt capability; in a low-risk situation such as regular cruising, the system can relax the tolerance to avoid redundant intervention), the dynamic feedback reminder result includes the voice and graphical interface feedback result, the atmosphere lamp feedback result, the odor prompt feedback result, and the tactile feedback result, and the trust level includes any one of the first trust, the second trust, and the third trust.

[0114] In the embodiment, the voice and graphical interface feedback result: the system can broadcast prompt content through the vehicle-mounted voice system, and the language style is flexibly adjusted according to the user emotion and cognitive load, for example, an encouraging tone is used to guide a low-trust (first trust) driver to establish confidence, or a serious tone is used to warn an excessive reliance, and at the same time, risk information, system state, or intervention suggestion can be synchronously displayed through a graphic interface such as a center control screen and a head-up display (HUD), the atmosphere lamp feedback result: situational visual guidance is realized through the in-vehicle atmosphere lamp, for example, when the third trust state is detected, the light can be adjusted to a cool tone (such as blue light) and low-frequency flickering is used to prompt alertness; if the low-trust state is detected, warm tone stable lighting (such as orange light) can be enabled to relieve the tension of the driver and improve his acceptance of the system, the odor prompt feedback result: the system can integrate a micro odor release device to release a fragrance with a mood regulating effect at a key state node, for example, when the driver is in a high anxiety or the system trust value is too low, a lavender type fragrance with a soothing effect can be released; in a fatigue state or a low trust level but requiring a quick response, a refreshing type fragrance such as mint and citrus can be released to improve attention, and the tactile feedback result: steering wheel vibration and seat vibration as a low-interference and strong-perceptible prompt method can be started when the system identifies that the driver's attention is distracted or the trust is deviated, to quickly establish a feedback loop and guide the driver to pay attention to the interactive task, for example, when the trust is too high and the hand is not in contact with the steering wheel, the steering wheel vibration is used to remind the driver to keep alert and prepare to take over, for example, if the system detects that the driver has not paid attention to the road for a long time (for example, the line of sight deviates and the steering wheel is not held), the seat vibration warning is triggered, and if the driver still does not respond, the vibration amplitude and frequency are increased, accompanied by visual and sound prompts.

[0115] Further, the step S30 further comprises obtaining a trust level of the driver, if the trust level is in a first trust state, the trust level is adjusted according to the voice and graphical interface feedback result and the smell prompt feedback result, to obtain a target trust level (for example, when the driver is in a high anxiety state or the system trust value is too low, a lavender smell with a soothing effect can be released; when in a fatigue state or the trust level is low but a quick response is required, a refreshing smell such as mint or citrus can be released to improve attention), if the trust level is in a third trust state, the driver is reminded according to the tactile feedback result and the atmosphere lamp feedback result, so that the driver can drive safely (for example, when the trust level is too high and the driver's hand is not in contact with the steering wheel, the driver is reminded to keep alert and prepare to take over by the steering wheel vibration), so that the driver can drive safely.

[0116] Further, as shown in the above Figure 2 Based on the above automatic driving trust state early warning method based on multi-modal features, the present application also correspondingly provides an automatic driving trust state early warning system based on multi-modal features, wherein the automatic driving trust state early warning system based on multi-modal features comprises:

[0117] A feature extraction module 51 is configured to obtain eye movement data signals, driving behavior data and individual feature information of a driver, and extract features from the eye movement data signals, the driving behavior data and the individual feature information to obtain eye movement features, driving features and individual features.

[0118] A trust state prediction module 52 is configured to fuse the eye movement features, the driving features and the individual features to obtain a to-be-tested multi-modal feature, and input the to-be-tested multi-modal feature into a trust prediction model for prediction to obtain a trust state prediction result.

[0119] A trust state feedback module 53 is configured to dynamically feedback the driver according to a trust level of the trust state prediction result and the individual features to obtain a dynamic feedback reminder result.

[0120] Further, as shown in the above Figure 3 Based on the above automatic driving trust state early warning method based on multi-modal features and system, the present application also correspondingly provides a terminal, which comprises a processor 10, a memory 20 and a display 30. Figure 3 Only part of the components of the terminal are shown, but it should be understood that all the shown components are not required to be implemented, and more or less components can be alternatively implemented.

[0121] The memory 20 can be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 can also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the memory 20 can include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software and various data installed on the terminal, such as program codes of the installed terminal, etc. The memory 20 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 20 stores the automatic driving trust state early warning program based on multi-modal features 40, which can be executed by the processor 10 to implement the automatic driving trust state early warning method based on multi-modal features in the present application.

[0122] The processor 10 can be a central processing unit (CPU), a microprocessor or other data processing chip in some embodiments, used to run program codes or process data stored in the memory 20, such as to execute the automatic driving trust state early warning method based on multi-modal features, etc.

[0123] The display 30 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visualized user interface. The terminal communicates with each other through a system bus.

[0124] In an embodiment, the following steps are implemented when the processor 10 executes the automatic driving trust state early warning program based on multi-modal features 40 in the memory 20:

[0125] Obtaining eye movement data signals of a driver, driving behavior data and individual characteristic information, performing feature extraction on the eye movement data signals, the driving behavior data and the individual characteristic information to obtain eye movement features, driving features and individual features;

[0126] Performing feature fusion on the eye movement features, the driving features and the individual features to obtain to-be-tested multi-modal features, inputting the to-be-tested multi-modal features into a trust prediction model for prediction to obtain a trust state prediction result;

[0127] According to the trust level of the trust state prediction result and the individual characteristics, dynamic feedback is performed on the driver, and a dynamic feedback reminding result is obtained.

[0128] The eye movement data signal of the driver, the driving behavior data of the driver, and the individual characteristic information of the driver are acquired, and feature extraction is performed on the eye movement data signal, the driving behavior data, and the individual characteristic information to obtain eye movement features, driving features, and individual characteristics, specifically including:

[0129] The eye movement data signal of the driver collected by the eye movement tracking device, the driving behavior data of the driver collected by the vehicle-mounted controller, and the individual characteristic information of the driver collected in advance are acquired.

[0130] The eye movement data signal is sequentially subjected to eye movement point processing, fixation point extraction, pupil processing, and blink recognition processing to obtain eye movement features.

[0131] The driving behavior data is subjected to interpolation filling and outlier removal to obtain target driving behavior data, and the eye movement data signal is time-aligned with the target driving behavior data to obtain driving features.

[0132] The individual characteristic information is subjected to normalization processing and one-hot encoding processing to obtain individual characteristics.

[0133] The individual characteristic information includes gender information, age information, driving experience information, emotional state information, and individual trait information.

[0134] The eye movement point processing includes interpolation processing and smoothing processing.

[0135] The eye movement features include fixation time, total fixation times, average pupil diameter, and saccade times.

[0136] The eye movement data signal is sequentially subjected to eye movement point processing, fixation point extraction, pupil processing, and blink recognition processing to obtain eye movement features, specifically including:

[0137] The eye movement data signal is subjected to interpolation processing and smoothing processing according to a linear interpolation formula and a preset interval length to obtain fixation time.

[0138] The eye movement data signal is subjected to fixation point extraction according to a preset angular velocity threshold to obtain total fixation times.

[0139] The eye movement data signal is subjected to pupil processing according to a preset pupil diameter to obtain average pupil diameter.

[0140] The eye movement data signal is subjected to blink recognition processing according to a preset time threshold to obtain saccade times.

[0141] The eye movement feature, the driving feature and the individual feature are fused to obtain a to-be-tested multi-modal feature, and the to-be-tested multi-modal feature is input into a trust prediction model for prediction to obtain a trust state prediction result.

[0142] The eye movement feature, the driving feature and the individual feature are fused to obtain a to-be-tested multi-modal feature, and the to-be-tested multi-modal feature is input into a trust prediction model for prediction to obtain a trust state prediction result.

[0143] The multi-modal feature is divided according to a preset division ratio to obtain a training set and a test set, the machine learning model is trained according to the training set to obtain a trust prediction model;

[0144] The to-be-tested multi-modal feature is input into the trust prediction model for prediction to obtain a trust state prediction result; and the test set is used for performance evaluation of the trust prediction model.

[0145] The machine learning model includes any one of a random forest model, an extreme random tree model, a limit gradient boosting model and a light gradient boosting machine model;

[0146] The trust prediction model includes a first trust prediction model, a second trust prediction model, a third trust prediction model and a fourth trust prediction model.

[0147] The multi-modal feature is divided according to a preset division ratio to obtain a training set and a test set, the machine learning model is trained according to the training set to obtain a trust prediction model, and the trust prediction model includes:

[0148] The multi-modal feature is divided according to a preset division ratio to obtain a training set and a test set.

[0149] A plurality of target training sets are obtained by multiple extraction of the training set through a bootstrap resampling method, a plurality of decision trees of the plurality of target training sets are determined according to a CART algorithm, a random forest model is generated according to the plurality of decision trees, the random forest model is trained according to the plurality of target training sets to obtain a first trust prediction model;

[0150] Or a plurality of feature sets of the training set are determined according to a preset double randomness mechanism, an extreme random tree model is generated according to the plurality of feature sets, the extreme random tree model is trained according to the training set to obtain a second trust prediction model;

[0151] Or the limit gradient boosting model is trained according to the training set and a preset node division strategy to obtain a third trust prediction model.

[0152] or according to the preset model parameters and the training set, the light gradient boosting machine model is trained to obtain a fourth trust prediction model;

[0153] The first trust prediction model, the second trust prediction model, the third trust prediction model and the fourth trust prediction model are tested for performance to obtain a trust prediction model.

[0154] The dynamic feedback reminder result includes voice and graphical interface feedback result, atmosphere lamp feedback result, smell prompt feedback result and tactile feedback result.

[0155] The trust level includes any one of first trust, second trust and third trust.

[0156] The trust level according to the trust state prediction result and the individual characteristics is dynamically fed back to the driver to obtain a dynamic feedback reminder result, which specifically includes:

[0157] According to a preset threshold strategy, the trust state prediction result is divided to obtain any one of the first trust, the second trust and the third trust.

[0158] According to any one of the first trust, the second trust and the third trust, and the gender information, the age information, the driving experience information, the emotional state information and the individual trait information, the driver is dynamically fed back to obtain voice and graphical interface feedback result, atmosphere lamp feedback result, smell prompt feedback result and tactile feedback result.

[0159] According to any one of the first trust, the second trust and the third trust, and the gender information, the age information, the driving experience information, the emotional state information and the individual trait information, the driver is dynamically fed back to obtain voice and graphical interface feedback result, atmosphere lamp feedback result, smell prompt feedback result and tactile feedback result, and then includes:

[0160] The trust level of the driver is obtained, and if the trust level is in the first trust, the trust level is fed back and adjusted according to the voice and graphical interface feedback result and the smell prompt feedback result to obtain a target trust level.

[0161] If the trust level is in the third trust, the driver is reminded according to the tactile feedback result and the atmosphere lamp feedback result so as to drive safely.

[0162] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores an automatic driving trust state early warning program based on multi-modal features, and the automatic driving trust state early warning program based on multi-modal features, when executed by a processor, implements the steps of the automatic driving trust state early warning method based on multi-modal features.

[0163] To sum up, the application provides an automatic driving trust state early warning method, system, terminal and storage medium based on multi-modal features, the method comprising: obtaining eye movement data signals, driving behavior data and individual feature information of a driver, performing feature extraction on the eye movement data signals, driving behavior data and individual feature information to obtain eye movement features, driving features and individual features; performing feature fusion on the eye movement features, the driving features and the individual features to obtain to-be-tested multi-modal features, inputting the to-be-tested multi-modal features into a trust prediction model for prediction to obtain a trust state prediction result; and performing dynamic feedback on the driver according to a trust level of the trust state prediction result and the individual features to obtain a dynamic feedback reminder result. The application accurately predicts the trust state through the trust prediction model, realizes dynamic evaluation of the trust state of the driver and feedback on the driver according to the trust state.

[0164] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles or terminal systems including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles or terminal systems. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or terminal system including the element.

[0165] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware (such as a processor, a controller, etc.) to complete, and the program can be stored in a computer readable computer readable storage medium, and the program can include the processes of the above-mentioned method embodiments when executed. The computer readable storage medium can be a memory, a disk, an optical disk, etc.

[0166] It should be understood that the application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all these improvements and changes should be within the protection scope of the claims of the application.

Claims

1. A method for early warning of trust status in autonomous driving based on multimodal features, characterized in that, The autonomous driving trust status early warning method based on multimodal features includes: The driver's eye movement data signals, driving behavior data, and individual characteristic information are acquired. Feature extraction is performed on the eye movement data signals, driving behavior data, and individual characteristic information to obtain eye movement features, driving features, and individual features. Eye-tracking features, driving features, and individual features are fused to obtain the multimodal features to be tested. The multimodal features to be tested are then input into the trust prediction model for prediction to obtain the trust state prediction result. Based on the trust level and individual characteristics of the trust status prediction results, dynamic feedback is given to the driver to obtain dynamic feedback reminder results; Eye-tracking features, driving features, and individual features are fused to obtain the multimodal features to be tested. These multimodal features are then input into a trust prediction model for prediction, yielding a trust state prediction result, which specifically includes: Eye-tracking features, driving features, and individual features are fused to obtain multimodal features and the multimodal features to be tested; The multimodal features are divided according to a preset division ratio to obtain a training set and a test set. The machine learning model is trained based on the training set to obtain a trust prediction model. The multimodal features to be tested are input into the trust prediction model for prediction, and the trust state prediction result is obtained. The test set is used to evaluate the performance of the trust prediction model; Machine learning models include any one of the following: random forest model, extreme random tree model, extreme gradient boosting model, and lightweight gradient boosting machine model; Trust prediction models include the first trust prediction model, the second trust prediction model, the third trust prediction model, and the fourth trust prediction model; The multimodal features are divided according to a preset partitioning ratio to obtain a training set and a test set. The machine learning model is then trained using the training set to obtain a trust prediction model, specifically including: The multimodal features are divided according to a preset division ratio to obtain a training set and a test set; Multiple target training sets are obtained by extracting multiple samples from the training set using the self-service resampling method. Multiple decision trees are determined for the multiple target training sets using the CART algorithm. A random forest model is generated based on the multiple decision trees. The random forest model is trained based on the multiple target training sets to obtain the first trust prediction model. Multiple feature sets of the training set are determined based on a preset dual randomness mechanism. An extreme random tree model is generated based on the multiple feature sets. The extreme random tree model is trained based on the training set to obtain the second trust prediction model. The extreme gradient boosting model is trained based on the training set and a preset node partitioning strategy to obtain the third trust prediction model. The lightweight gradient booster model is trained based on the preset model parameters and training set to obtain the fourth trust prediction model. Performance tests were conducted on the first, second, third, and fourth trust prediction models to obtain the trust prediction model.

2. The autonomous driving trust status early warning method based on multimodal features according to claim 1, characterized in that, The process of acquiring the driver's eye movement data signals, driving behavior data, and individual characteristic information, and extracting features from the eye movement data signals, driving behavior data, and individual characteristic information to obtain eye movement features, driving features, and individual features specifically includes: Acquire driver eye movement data signals collected by an eye-tracking device, driver driving behavior data collected by an onboard controller, and pre-collected individual characteristic information of the driver; The eye movement data signal is sequentially processed by eye movement point processing, fixation point extraction, pupil processing, and blink recognition processing to obtain eye movement features; The driving behavior data is interpolated, filled, and outliers are removed to obtain target driving behavior data. The eye-tracking data signal is then time-aligned with the target driving behavior data to obtain driving features. The individual feature information is normalized and one-hot encoded to obtain the individual features; The individual characteristic information includes gender information, age information, driving experience information, emotional state information, and individual trait information.

3. The autonomous driving trust status early warning method based on multimodal features according to claim 2, characterized in that, The eye-tracking point processing includes interpolation processing and smoothing processing; The eye movement characteristics include fixation time, total number of fixations, average pupil diameter, and number of saccades; The eye movement data signal is sequentially processed by eye movement point processing, fixation point extraction, pupil processing, and blink recognition to obtain eye movement features, specifically including: The eye movement data signal is interpolated and smoothed according to a linear interpolation formula and a preset interval length to obtain the fixation time. The eye movement data signal is subjected to fixation point extraction based on a preset angular velocity threshold to obtain the total number of fixations; The eye movement data signal is processed according to a preset pupil diameter to obtain the average pupil diameter; The blink count is obtained by performing blink recognition processing on the eye movement data signal according to a preset time threshold.

4. The autonomous driving trust status early warning method based on multimodal features according to claim 2, characterized in that, The dynamic feedback reminder results include voice and graphical interface feedback results, ambient light feedback results, odor prompt feedback results, and tactile feedback results; The trust level includes any one of first trust, second trust, and third trust. The step of providing dynamic feedback to the driver based on the trust level predicted by the trust status and the individual characteristics to obtain a dynamic feedback reminder result specifically includes: The trust state prediction results are divided according to a preset threshold strategy to obtain any one of the first trust, the second trust, and the third trust. Based on any one of the first trust, the second trust, and the third trust, as well as the gender information, the age information, the driving experience information, the emotional state information, and the individual trait information, dynamic feedback is provided to the driver to obtain voice and graphical interface feedback results, ambient light feedback results, odor cues feedback results, and tactile feedback results.

5. The autonomous driving trust status early warning method based on multimodal features according to claim 4, characterized in that, The process involves providing dynamic feedback to the driver based on any one of the first trust, the second trust, and the third trust, as well as the gender information, the age information, the driving experience information, the emotional state information, and the individual trait information, to obtain voice and graphical interface feedback results, ambient light feedback results, odor cue feedback results, and tactile feedback results. This process further includes: The driver's trust level is obtained. If the trust level is at the first level, the trust level is adjusted based on the feedback results from the voice and graphical interface and the odor prompts to obtain the target trust level. If the trust level is at the third level, the driver will be reminded based on the tactile feedback and the ambient light feedback to ensure safe driving.

6. An autonomous driving trust status early warning system based on multimodal features, characterized in that, The autonomous driving trust state warning system based on multimodal features is applied to the autonomous driving trust state warning method based on multimodal features according to any one of claims 1-5, wherein the autonomous driving trust state warning system based on multimodal features includes: The feature extraction module is used to acquire the driver's eye movement data signals, driving behavior data and individual feature information, and to extract features from the eye movement data signals, driving behavior data and individual feature information to obtain eye movement features, driving features and individual features; The trust state prediction module is used to fuse the eye movement features, the driving features, and the individual features to obtain the multimodal features to be tested, and input the multimodal features to be tested into the trust prediction model for prediction to obtain the trust state prediction result. The trust status feedback module is used to provide dynamic feedback to the driver based on the trust level of the trust status prediction result and the individual characteristics, and obtain dynamic feedback reminder results.

7. A terminal, characterized in that, The terminal includes: a memory, a processor, and an autonomous driving trust status warning program based on multimodal features stored in the memory and executable on the processor. When the autonomous driving trust status warning program based on multimodal features is executed by the processor, it implements the steps of the autonomous driving trust status warning method based on multimodal features as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an autonomous driving trust status warning program based on multimodal features, which, when executed by a processor, implements the steps of the autonomous driving trust status warning method based on multimodal features as described in any one of claims 1-5.

Citation Information

Patent Citations

  • System and method for improving driving credibility based on multi-modal intelligent interaction

    CN118918558A