Driver decision confidence estimation method and system based on multi-modal data
By integrating eye movement, EEG and driving operation data, combined with LSTM and category distance loss function, a decision confidence estimation model is constructed, which solves the problem of single modal data and ignoring sequence in traditional methods, and accurately evaluates driver decision confidence and improves estimation accuracy.
Patent Information
- Application Number
- CN202510651605.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to accurately evaluate driver's decision-making confidence in dynamic driving scenarios. Traditional methods rely mostly on single modal data and ignore the order of decision-making confidence, resulting in insufficient estimation accuracy.
By collecting driver's eye movement data, EEG data and driving operation data, combining long and short-term memory networks (LSTMs) for feature fusion, and using a weighted loss function based on category distance, a decision confidence estimation model is constructed.
It realizes accurate estimation of driver's decision-making confidence, improves estimation accuracy, adapts to the needs of dynamic driving scenarios, and provides new ideas for the optimization of intelligent driving systems.
Smart Images

Figure CN120284272A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of artificial intelligence and cognitive science, and particularly relates to a method and system for estimating driver decision confidence based on multi - modal data. Background Art
[0002] Driver decision confidence refers to the degree of confidence of a driver in the correctness of the decisions made during driving, and decision confidence has been proven to be highly correlated with the correctness of decisions.
[0003] With the rapid development of intelligent driving technology, intelligent driving systems have made remarkable progress in perception, decision - making, and control. However, in semi - autonomous driving or human - machine co - driving scenarios, the decision - making behavior of drivers remains a key factor affecting driving safety and system performance. When drivers face complex driving scenarios (such as intersection decision - making, emergency avoidance, etc.), their decision confidence directly reflects the accuracy of the current decision. Insufficient decision confidence may lead to hesitation or incorrect operations, thus increasing the risk of traffic accidents. Therefore, accurately estimating driver decision confidence not only helps to understand the driver's cognitive state but also provides an important reference for the design of intelligent driving systems to achieve safer and more efficient driver - machine interaction.
[0004] In the prior art, the research on driver state mainly focuses on fatigue detection, attention analysis, and emotion recognition, etc. For example, the eye - movement tracking technology is used to detect the driver's fixation area, or the electroencephalogram signal is analyzed to detect the driver's fatigue level. However, these methods mostly focus on the evaluation of physiological states and do not involve the evaluation of the driver's cognitive state (such as confidence) during the decision - making process. Traditional confidence evaluation methods are usually based on simple visual discrimination tasks in the laboratory and cannot meet the requirements of dynamic driving scenarios. In addition, the analysis of single - modal data (such as only using electroencephalogram data) in existing research is difficult to comprehensively reflect the decision - confidence state and also ignores the sequential characteristics of the decision - confidence variable, resulting in insufficient estimation accuracy of decision confidence.
[0005] Therefore, there is an urgent need for a method and system that can accurately estimate driver decision confidence to precisely capture the confidence state of drivers during the decision - making process and provide new methods and ideas for the optimization of intelligent driving systems. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention provides a method and system for estimating driver decision confidence based on multi-modal data. By collecting eye movement data, electroencephalogram data, and driving operation data of the driver under decision-making events in the driving task, preprocessing and feature fusion are carried out, and then combined with a long short-term memory network (LSTM) for modeling. At the same time, a weighted loss function based on class distance is used to consider the sequentiality of decision confidence, aiming to achieve accurate estimation of driver decision confidence and provide new methods and ideas for the optimization of intelligent driving systems.
[0007] The specific technical solution is as follows:
[0008] A method for estimating driver decision confidence based on multi-modal data, comprising the following steps:
[0009] S1: Collect multi-modal data of the driver under decision-making events in the driving task, including eye movement data obtained by an eye tracker, electroencephalogram data obtained by an electroencephalogram device, and driving operation data obtained by vehicle sensors, and preprocess the multi-modal data;
[0010] S2: Collect confidence data of the driver under decision-making events in the driving task as the data label for training the decision confidence estimation model;
[0011] S3: Extract features related to decision confidence from the preprocessed multi-modal data, including eye movement features, electroencephalogram features, and driving operation features, and generate a multi-modal feature vector through windowing and feature fusion;
[0012] S4: Build a neural network model based on the long short-term memory network (LSTM), train it using the multi-modal feature vector and confidence data, and optimize the network parameters by using a weighted loss function based on class distance to consider the sequentiality of decision confidence, and obtain a decision confidence estimation model for estimating driver decision confidence;
[0013] S5: Apply the trained neural network model to test data and output an estimated value of driver decision confidence based on the input multi-modal feature vector.
[0014] A system for estimating driver decision confidence based on multi-modal data, used to implement the above method for estimating driver decision confidence based on multi-modal data, including:
[0015] Multi-modal data acquisition module: Used to collect multi-modal data and information data of the driver under decision-making events in the driving task. The multi-modal data includes eye movement data, electroencephalogram data, and vehicle motion state data;
[0016] Multi-modal data preprocessing module: Used to preprocess the collected multi-modal data;
[0017] Feature extraction and fusion module: It is used to extract features related to decision confidence from the preprocessed multi-modal data, including eye movement features, electroencephalogram features, and driving operation features, and form a multi-modal feature vector through windowing and feature fusion;
[0018] Model training module: It is used to construct a long short-term memory network, train it using the multi-modal feature vector and confidence data, and optimize the network using a weighted loss function based on class distance;
[0019] Decision confidence estimation module: It is used to apply the trained long short-term memory network to test data and output an estimated value of the driver's decision confidence.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] The present invention first realizes the accurate estimation of the driver's decision confidence by fusing multi-modal data such as eye movement, electroencephalogram, and driving operation, and combining a long short-term memory network (LSTM) and a weighted loss function based on class distance, overcomes the limitations of traditional methods that rely on single-modal data, ignore the sequentiality of confidence, and cannot adapt to dynamic driving scenarios, thereby significantly improving the accuracy of estimating the driver's decision confidence and providing new methods and ideas for optimizing intelligent driving systems. Description of the drawings
[0022] Figure 1 It is a flowchart of a method for estimating a driver's decision confidence based on multi-modal data according to the present invention;
[0023] Figure 2 It is a structural diagram of a system for estimating a driver's decision confidence based on multi-modal data according to the present invention.
[0024] Multi-modal data acquisition module 21, multi-modal data preprocessing module 22, feature extraction and fusion module 23, model training module 24, decision confidence estimation module 25. Detailed implementation manners
[0025] In order to make the objectives, technical solutions, and effects of the present invention clearer, the preferred embodiments of the present invention are described below with reference to the drawings of the specification. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0026] The present invention provides a method for estimating a driver's decision confidence based on multi-modal data. Figure 1 The flowchart of the method is shown. As Figure 1 shown, the method specifically includes the following steps:
[0027] S1: Collect multimodal data of drivers in decision-making events during driving tasks, including eye movement data obtained by an eye tracker, electroencephalogram (EEG) data obtained by an EEG device, and driving operation data obtained by vehicle sensors, and preprocess the multimodal data;
[0028] Specifically, the eye movement data is collected by a Pupil Core eye tracker. The eye movement data records the pupil size and fixation point position of the driver, and the sampling frequency of the Pupil Core eye tracker is 100 Hz.
[0029] Specifically, the EEG data is collected by a Starstim 8 EEG device. Based on the international 10-20 system, channels Fpz, AF7, AF8, F3, Fz, F4, Cz, and Pz are selected, and the sampling frequency is 100 Hz.
[0030] Specifically, the driving operation data is collected by a Logitech G29 driving sensor. The driving operation data includes the steering wheel angle, accelerator pedal value, and brake pedal value, and the sampling frequency of the Logitech G29 driving sensor is 100 Hz.
[0031] Specifically, the driving task is to complete a 120-minute drive in an urban road scenario, including various decision-making events (such as lane changes, turns, overtaking, etc.). A total of 30 drivers participated in the driving task, and multimodal data corresponding to 200 decision-making events were recorded for each driver.
[0032] Specifically, the preprocessing of the multimodal data includes the following sub-steps:
[0033] S11: Use linear interpolation to align the timestamps of the eye movement data, EEG data, and driving operation data to ensure the time synchronization of the multimodal data. The formula for timestamp alignment is:
[0034] ,
[0035] where, is the unified timestamp after alignment, represents the minimum value, , , are the original timestamps of the eye movement data, EEG data, and driving operation data respectively, represents the sampling point index, is the sampling time interval corresponding to the data with the highest sampling rate, ;
[0036] S12: Perform artifact removal processing on the EEG data. Use the independent component analysis (ICA) method to decompose the original EEG signal into multiple independent components and remove the artifacts. Among them, the original EEG signal Expressed as multiple independent source signals Through a mixing matrix The linear combination of, the formula is:
[0037] ,
[0038] By calculating the demixing matrix For Perform demixing to obtain the estimated independent source signals :
[0039] ,
[0040] Among them, Is the multi-channel raw EEG signal, Is the mixing matrix, representing the linear superposition relationship between artifacts and real EEG signals, Is the independent source signal, including real EEG signals and artifact components; the mixing matrix And the demixing matrix Are obtained by calculating based on the statistical independence of the signals through the ICA algorithm. The ICA algorithm includes, but is not limited to, optimization methods based on information maximization or negentropy maximization; by identifying and removing the independent components related to artifacts such as eye movements and EMG, the artifact-free EEG signal Is obtained.
[0041] S13: Filter the EEG data: Perform band-pass filtering on the artifact-free EEG signal to retain the frequency band related to decision confidence (0.5 Hz to 40 Hz), while filtering out low-frequency drift and high-frequency noise. The filtering parameters are set as follows: the upper limit of the filtering frequency , the lower limit of the filtering frequency , the center frequency , the bandwidth ;
[0042] The filtering formula is:
[0043] ,
[0044] Among them, Is the impulse response function of the band-pass filter, designed by Fourier transform, and its frequency domain response Is the passband within the filtering frequency range To , and other frequencies are stopbands, The expression of is:
[0045] ,
[0046] Among them, Is the center frequency, .
[0047] S14: Smooth the eye movement data: Smooth the pupil size and fixation point position data using moving average filtering;
[0048] Smoothing the eye movement data can reduce noise. The smoothed eye movement data is calculated by the formula:
[0049] ,
[0050] where, is the original eye movement data, is the total number of sampling points in the filtering window, is the sampling time interval corresponding to the data with the highest sampling rate, represents the sampling point index. In one embodiment, the window length .
[0051] S15: Normalize the EEG data, eye movement data, and driving operation data. The normalization formula is:
[0052] ,
[0053] where, represents the data of any modality (the filtered EEG data , the smoothed eye movement data , or the driving operation data ), and are the minimum and maximum values of the corresponding modality data, is the data after normalization processing.
[0054] S2: Collect the confidence data of the driver in the decision-making event during the driving task as the data label for training the decision confidence estimation model. The confidence data can be obtained through the driver's subjective report after the decision-making event, using a standardized five-level Likert Scale. The specific method is: After each decision-making event, the driver reports the confidence value in the range of 1 to 5 by pressing a button (1 = completely no confidence, 5 = completely confident, and the confidence levels represented by 1 to 5 increase in sequence). Each driver completes 200 decision-making events, and the reliability of the confidence data collection is verified through a repeatability test (Cronbach's alpha = 0.85) after the experiment.
[0055] S3: Extract the features related to decision confidence from the preprocessed multi-modal data, including eye movement features, EEG features, and driving operation features, and generate a multi-modal feature vector through windowing and feature fusion. Specifically, it includes the following sub-steps:
[0056] S31: Extract the pupil size change rate and the fixation point dispersion from the eye movement data. and the fixation point dispersion , the calculation formula for the pupil size change rate is:
[0057] ,
[0058] where is the normalized pupil size data, is the sampling time interval corresponding to the data with the highest sampling rate, ;
[0059] The calculation formula for the fixation point dispersion is:
[0060] ,
[0061] where is the normalized fixation point coordinate, is the mean coordinate of the fixation points within the window, is the number of fixation points within, is the time window length, ;
[0062] S32: Extract the frequency domain features from the EEG data, including the power spectral density and the differential entropy , the calculation formula for the power spectral density is:
[0063] ,
[0064] where is the time window length, , is the frequency of the EEG signal, ranging from 0.5 Hz to 50 Hz, is the normalized EEG signal;
[0065] The calculation formula for the differential entropy is:
[0066] ,
[0067] where is the variance of the signal within the frequency band, and are the frequency ranges of the band-pass filter, ranging from 0.5 Hz to 50 Hz; represents the natural logarithm;
[0068] S33: Extract the steering wheel angle change rate from the driving operation data The change rate of the accelerator pedal value and the change rate of the brake pedal value , and their calculation formulas are respectively:
[0069] ,
[0070] ,
[0071] wherein, , and are respectively the normalized steering wheel angle, accelerator pedal value and brake pedal value, is the sampling time interval, .
[0072] S34: According to the decision-making events of the driver, the extracted feature data is segmented into multiple time windows, each time window corresponding to a decision-making event, and the window segmentation formula is:
[0073] ,
[0074] wherein, is the time window of the th decision-making event, is the start time of the time window of the th decision-making event, is the end time of the time window of the th decision-making event, is the time when the th decision-making event occurs, and are respectively the preset time ranges before and after the occurrence of the decision-making event, , ;
[0075] S35: The feature data under each decision-making event window is further processed by the sliding window method to retain the temporal information between features. The length of the sliding window is , and the sliding step is (i.e., there is a 50% overlapping window);
[0076] S36: The features within each sliding window are fused into a multi-modal feature vector , and the expression of
[0077] ,
[0078] wherein, represents the a sliding window represents the number of sliding windows within the decision event window.
[0079] S4: Construct a neural network model based on a long short-term memory network (LSTM), train it using the multi-modal feature vectors and confidence data, and optimize the network parameters by using a weighted loss function based on class distance to consider the sequentiality of decision confidence, thereby obtaining a trained model for estimating driver decision confidence.
[0080] S4 includes the following sub-steps:
[0081] S41: Construct a long short-term memory network (LSTM) to process the temporal characteristics of the multi-modal feature vectors: Use the windowed multi-modal feature vector sequence as the input, where the time step of the LSTM is equal to the number of windows ; the LSTM adopts a two-layer structure; each layer updates the hidden state and cell state through an input gate, a forget gate, and an output gate to generate temporal features, where the output of the first layer is used as the input of the second layer, and the hidden state at the final time step of the second layer serves as the temporal feature representation of the entire sequence;
[0082] Among them, in the two-layer structure of the long short-term memory network (LSTM), the first layer contains 128 hidden units, and the second layer contains 64 hidden units.
[0083] S42: Use the hidden state at the final time step of the second layer of the LSTM as the input and output the confidence estimate through a fully connected layer , and the calculation formula of the fully connected layer is:
[0084] ,
[0085] where, is the weight matrix of the fully connected layer, is the bias vector;
[0086] S43: Optimize the network parameters by using a weighted loss function based on class distance to consider the sequentiality of decision confidence. The loss function is:
[0087] ,
[0088] where, is the true confidence label, is the confidence estimate of the model, is the number of samples, is the weight function based on class distance, and its calculation formula is:
[0089] ,
[0090] Among them, is the distance between the true label and the predicted label, is a tuning parameter. In one embodiment, .
[0091] The sequentiality of decision confidence means that the confidence values (from 1 to 5) themselves have a magnitude order relationship. Using a weighted loss function based on class distance to optimize the network parameters can make the loss weight higher when the prediction deviation is larger, thus taking into account the sequentiality of decision confidence. For example, the loss when the predicted value is 3 and the true value is 5 is higher than the loss when the predicted value is 4 and the true value is 5. This design can improve the model's ability to distinguish the order of confidence levels.
[0092] During the training process, the Adam optimizer is used for iteration, with 500 iterations and a learning rate of 0.001. The training set and the test set are divided in a ratio of 8:2.
[0093] S5: Apply the trained neural network model to the test data, and output an estimated value of the driver's decision confidence based on the input multi-modal feature vector.
[0094] The present invention also provides a driver decision confidence estimation system based on multi-modal data. Figure 2 The block diagram showing the driver decision confidence estimation system 20 is as follows. Figure 2 As shown, the system includes:
[0095] Multi-modal data acquisition module 21: used to collect multi-modal data and information data of the driver in decision-making events during driving tasks. The multi-modal data includes eye movement data, electroencephalogram data, and vehicle motion state data;
[0096] Multi-modal data preprocessing module 22: used to preprocess the collected multi-modal data, including filtering, smoothing, removing artifacts, timestamp alignment, and normalization of relevant data;
[0097] Feature extraction and fusion module 23: used to extract features related to decision confidence from the preprocessed multi-modal data, including eye movement features, electroencephalogram features, and driving operation features, and form a multi-modal feature vector through windowing and feature fusion;
[0098] Model training module 24: used to construct a long short-term memory network (LSTM), train using the multi-modal feature vector and confidence data, and optimize the network using a weighted loss function based on class distance;
[0099] Decision confidence estimation module 25: used to apply the trained long short-term memory network to the test data and output an estimated value of the driver's decision confidence.
[0100] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for estimating driver decision-making confidence based on multi-modal data, characterized in that, It includes the following steps: S1: Collect multi-modal data of the driver under decision-making events during the driving task, including eye movement data obtained by an eye tracker, electroencephalogram (EEG) data obtained by an EEG device, and driving operation data obtained by vehicle sensors, and preprocess the multi-modal data; S2: Collect confidence data of the driver under decision-making events during the driving task as the data label for training the decision confidence estimation model; S3: Extract features related to decision confidence from the preprocessed multi-modal data, including eye movement features, EEG features, and driving operation features, and generate a multi-modal feature vector through windowing and feature fusion; S4: Construct a neural network model based on the long short-term memory network (LSTM), train it using the multi-modal feature vector and confidence data, and optimize the network parameters by using a weighted loss function based on class distance to consider the sequentiality of decision confidence, so as to obtain a decision confidence estimation model for estimating the driver's decision confidence; S5: Apply the trained neural network model to test data and output an estimated value of the driver's decision confidence based on the input multi-modal feature vector.
2. The driver decision confidence estimation method based on multi-modal data according to claim 1, wherein The eye movement data in step S1 includes pupil size and fixation point position. The EEG data is collected according to the international 10-20 system, and channels Fpz, AF7, AF8, F3, Fz, F4, Cz, and Pz are selected. The driving operation data includes steering wheel angle, accelerator pedal value, and brake pedal value.
3. A method for estimating a driver's decision-making confidence based on multimodal data according to claim 1, characterized in that, The preprocessing in step S1 includes the following sub-steps: S11: Use linear interpolation to align the timestamps of the eye movement data, EEG data, and driving operation data to ensure the time synchronization of the multi-modal data. The formula for timestamp alignment is: , Among them, is the unified timestamp after alignment, represents the minimum value, , , are the original timestamps of eye movement data, EEG data, and driving operation data respectively, represents the sampling point index, is the sampling time interval corresponding to the data with the highest sampling rate; S12: Perform artifact removal processing on the EEG data. Use the independent component analysis method to decompose the original EEG signal into multiple independent components and remove the artifacts. Among them, the original EEG signal is expressed as multiple independent source signals through the linear combination of the mixing matrix The formula is: , By calculating the demixing matrix Demix to obtain the estimated independent source signals : , wherein, is a multi-channel raw EEG signal, is a mixing matrix, representing the linear superposition relationship between artifacts and real EEG signals, is an independent source signal, including real EEG signals and artifact components; the mixing matrix and the demixing matrix are obtained by calculating based on the statistical independence of the signals through the ICA algorithm, and the ICA algorithm includes optimization methods based on information maximization or negative entropy maximization; by identifying and removing the independent components related to eye movement and EMG artifacts, the artifact-free EEG signal is obtained; S13: Perform band-pass filtering on the artifact-removed EEG signal to retain the frequency bands related to decision confidence. The filtering formula is: , Among them, is the impulse response function of the band-pass filter, which is designed by Fourier transform, and its frequency-domain response in the filtering frequency range to is the passband, and other frequencies are the stopband. The expression of , Among them, is the center frequency, ; S14: Smooth the eye movement data to reduce noise. The moving average filtering method is adopted. The smoothed eye movement data is calculated by the following formula: , Among them, is the original eye movement data, is the total number of sampling points in the filtering window, is the sampling time interval corresponding to the data with the highest sampling rate, represents the sampling point index; S15: Normalize the EEG data, eye movement data, and driving operation data. The normalization formula is: , Among them, represents the data of any mode, and are the minimum and maximum values of the corresponding modal data, is the data after normalization.
4. The driver decision confidence estimation method based on multimodal data according to claim 3, wherein, Upper limit of filtering frequency , Lower limit of filtering frequency , Center frequency , Bandwidth .
5. The driver decision confidence estimation method based on multimodal data according to claim 1, characterized in that The confidence data in step S2 is obtained through the driver's subjective report after the decision-making event. The subjective report uses a five-level Likert scale, ranging from 1 to 5, where 1 represents complete lack of confidence and 5 represents complete confidence.
6. The driver decision confidence estimation method based on multimodal data according to claim 1, characterized in that The extraction of features related to decision confidence in step S3 includes the following sub-steps: S31: Extract the pupil size change rate and the fixation point dispersion from the eye movement data and the fixation point dispersion , and the calculation formula for the pupil size change rate is as follows: , Among them, is the normalized pupil size data, is the sampling time interval corresponding to the data with the highest sampling rate; Gaze point dispersion The calculation formula is as follows: , Among them, is the normalized fixation point coordinate, is the mean coordinate of the fixation points within the window, is the number of fixation points within, is the time window length; S32: Extract frequency domain features from the EEG data, including power spectral density and differential entropy , the calculation formula of power spectral density is as follows: , Among them, is the time window length, is the EEG signal frequency, is the normalized EEG signal; Differential entropy The calculation formula is as follows: , Among them, is the variance of the in-band signal, and is the frequency range of the band-pass filter; represents the natural logarithm; S33: Extract the steering wheel angle change rate from the driving operation data , the accelerator pedal value change rate and the brake pedal value change rate , and the calculation formulas are respectively as follows: , , Among them, , and are the normalized steering wheel angle, accelerator pedal value, and brake pedal value respectively, is the sampling time interval corresponding to the data with the highest sampling rate; S34: According to the driver's decision-making event, divide the extracted feature data into multiple time windows, each time window corresponding to a decision-making event. The window division formula is: , Among them, is the time window of the th decision event, is the starting moment of the time window of the th decision event, is the ending moment of the time window of the th decision event, is the moment when the th decision event occurs, and are respectively the preset time ranges before and after the occurrence of the decision event; S35: Further process the feature data under each decision event window using the sliding window method to retain the temporal information between features. The length of the sliding window is , and the sliding step is ; S36: Fuse the features within each sliding window into a multi-modal feature vector , The expression of which is: , Among them, represents the th sliding window within the decision event window, and represents the number of sliding windows within the decision event window.
7. The method for estimating a driver's decision-making confidence based on multimodal data according to claim 6, wherein and ranging from 0.5 Hz to 50 Hz, , .
8. A method for estimating a driver's decision-making confidence based on multimodal data according to claim 6, characterized in that, Step S4 includes the following sub-steps: S41: Construct a long short-term memory network to process the temporal characteristics of the multi-modal feature vectors: The sequence of windowed multi-modal feature vectors is used as the input, and the time step of the long short-term memory network is equal to the number of windows ; The LSTM adopts a two-layer structure; Each layer updates the hidden state and cell state through the input gate, forget gate, and output gate to generate temporal features, where the output of the first layer is used as the input of the second layer, and the hidden state of the last time step of the second layer is used as the temporal feature representation of the entire sequence; S42: Use the output features of the last time step of the second layer of LSTM as input, and output the confidence estimate through the fully connected layer , and the calculation formula of the fully connected layer is: , Among them, is the weight matrix of the fully connected layer, is the bias vector; S43: Use a weighted loss function based on class distance to optimize the network parameters to consider the sequentiality of decision confidence. The loss function is: , Among them, is the true confidence label, is the model confidence estimate, is the number of samples, is the weight function based on class distance, and the calculation formula is: , Among them, is the distance between the true label and the predicted label, is the adjustment parameter.
9. A method for estimating a driver's decision-making confidence based on multi-modal data according to claim 8, characterized in that, In the double-layer structure of the short-term memory network, the first layer contains 128 hidden units, and the second layer contains 64 hidden units; during the training process, the Adam optimizer is used for iteration, the number of iterations is 500 times, and the learning rate is 0.
001.
10. A driver decision confidence estimation system based on multimodal data, characterized in that, For implementing the driver decision confidence estimation method based on multi-modal data according to any one of claims 1-9, it includes: Multimodal Data Acquisition Module: It is used to collect multimodal data and information data of the driver under decision-making events during driving tasks. The multimodal data includes eye movement data, electroencephalogram data, and driving operation data; Multimodal Data Preprocessing Module: It is used to preprocess the collected multimodal data; Feature Extraction and Fusion Module: It is used to extract features related to decision confidence from the preprocessed multimodal data, including eye movement features, electroencephalogram features, and driving operation features, and form a multimodal feature vector through windowing and feature fusion; Model Training Module: It is used to construct a long short-term memory network, train it using the multimodal feature vector and confidence data, and optimize the network using a weighted loss function based on class distance; Decision Confidence Estimation Module: It is used to apply the trained long short-term memory network to test data and output an estimated value of the driver's decision confidence.
Citation Information
Cited By
Automatic driving trust state early warning method and system based on multi-modal features, terminal and storage medium
CN120697779A