A driving state recognition method and system based on multimodal fusion
Through multimodal data fusion and dynamic weighting model, the problems of low driving state recognition accuracy and insufficient real-time performance in the prior art are solved, and more accurate and efficient driving state recognition is achieved, and real-time monitoring and accident prevention are supported.
Patent Information
- Application Number
- CN202510012081.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art has problems in driving state recognition with poor feature correlation, high invasiveness of acquisition equipment, single data source and low feature fusion accuracy, resulting in poor recognition effect.
The driving state recognition method based on multimodal fusion is adopted, and the recognition accuracy and real-timeness are improved by comprehensively collecting psychological signal data, physiological signal data, driving performance data and wrist movement information, and feature fusion and weighting are performed through feature extraction, dynamic weighting model and multi-head self-attention mechanism.
It improves the accuracy and robustness of driving status recognition, enhances the real-time and computing efficiency of the model, can accurately identify the driver's driving status, and supports real-time driver behavior monitoring and traffic accident prevention.
Smart Images

Figure CN119416003B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation technology and relates to a driving state recognition method and system based on multi-modal fusion. Background Art
[0002] Fatigue driving and distracted driving are the two most common causes of major accidents. In the case of the coexistence of multiple modes of transportation, accurate detection of the driver's driving status can greatly improve the safety of urban road traffic and road operation efficiency. At present, mainstream research relies on visual features, EEG, EOG, EMG, etc. to identify bad driving conditions, but these data collection devices are expensive and come into contact with the body, affecting normal operation, and are easily affected by light and obstructions, making them almost unsuitable.
[0003] Fatigue driving and distracted driving are two common bad driving behaviors. Fatigue driving is usually caused by insufficient rest, long-term driving or night driving, and its manifestations include inattention, slow reaction, frequent yawning and blinking. Distracted driving is caused by drivers' habitual behaviors such as making phone calls and smoking, which also lead to distraction and slow reaction. The above phenomena are very likely to cause traffic accidents. Therefore, it is crucial to quickly and accurately detect bad driving behaviors of drivers to ensure the safety of drivers, passengers and pedestrians.
[0004] Although some driving status recognition methods have been proposed in the prior art, these methods often have poor final recognition effects due to poor feature correlation, invasive acquisition equipment, single data source, and low feature fusion accuracy.
[0005] In terms of feature correlation, the existing technology does not effectively evaluate the contribution of different features, resulting in unreasonable processing weights of different features. All features are used as inputs into the recognition model, which affects the accuracy of the model.
[0006] In terms of the invasiveness of the acquisition equipment, traditional physiological detection equipment is expensive and has low comfort, which may interfere with the driver's driving behavior and thus affect his state recognition. In addition, when using a combination of EEG equipment or physiological signal sensors to detect adverse driving conditions, many electrodes or sensors need to be attached to the driver, which may be invasive. Many actions in actual driving will produce artifacts in the measurement signal, affecting the accuracy of adverse driving state detection.
[0007] In terms of data source, the existing bad driving state recognition system mainly relies on a single signal source as input, resulting in a high false alarm rate and missed alarm rate. Although many researchers have proposed using sensors to obtain multimodal information to comprehensively detect the driver's bad driving state. However, due to many reasons such as the difficulty of data collection, the complexity of algorithm recognition, and the susceptibility of detection results to external factors, the recognition of bad driving state has brought great challenges.
[0008] In terms of feature fusion accuracy, existing methods for identifying bad driving conditions often fail to effectively fuse features from different sources, resulting in low recognition accuracy. In addition, many methods do not fully evaluate the importance of each feature, resulting in unreasonable distribution of feature processing weights and directly inputting all features evenly into the model, thus affecting the accuracy of the model. More importantly, existing technologies are insufficient in real-time performance and computational efficiency, making it difficult to meet the needs of actual applications. Summary of the invention
[0009] The purpose of the present invention is to propose a driving state recognition method based on multimodal fusion. The method is based on multimodal data feature fusion. By introducing a weighted fusion mechanism, the weight of each feature is automatically adjusted according to the contribution of each feature to the recognition of a specific driving state, which is beneficial to improving the driving state recognition accuracy while effectively improving the real-time performance and computational efficiency of the model.
[0010] In order to achieve the above object, the present invention adopts the following technical scheme:
[0011] A driving state recognition method based on multimodal fusion comprises the following steps:
[0012] Step 1. Comprehensively collect multimodal data of psychological signal data, physiological signal data, driving performance data, and wrist motion information, and preprocess the collected multimodal data;
[0013] Step 2. Further feature extraction is performed on the preprocessed multimodal data, and the features with significant differences under different driving conditions are screened out from the feature extraction results through Friedman test and Bonferroni correction;
[0014] Step 3. Build a dynamic feature weighting model to capture the dynamic changes of features and the correlation between features under driving conditions and adjust the importance weight of features in real time according to the changes in feature states;
[0015] The feature dynamic weighting model includes LSTM module, cross attention mechanism module and Transformer module;
[0016] The selected features are fed into the LSTM module, which captures the temporal dependencies between different features.
[0017] Each feature input into the LSTM module will go through the temporal modeling process in the LSTM module. The LSTM learns the dynamic changes of the driver under different behavior states from the temporal pattern of each feature.
[0018] The LSTM module outputs the hidden state vector of each feature at different time steps; the hidden state vector output by the LSTM module is divided into four categories, corresponding to physiological signal data, psychological signal data, driving performance data and wrist movement information;
[0019] Among them, each type of hidden state vector contains the time-dependent information of each feature in the corresponding modality;
[0020] The four types of hidden state vectors are input into the cross-attention mechanism module, which calculates the attention weights between multiple modalities, dynamically adjusts the contribution of each modality feature, and integrates the information of different modalities;
[0021] The output features of the cross-attention mechanism module are input into the Transformer module, and the multi-head self-attention mechanism of the Transformer module is further used to achieve deep interaction and feature enhancement of different modal data to obtain fusion features, namely weighted physiological signal data, psychological signal data, driving performance data, and wrist motion information features;
[0022] Step 4. Build a driving state recognition model based on XGBoost and perform model training. The trained driving state recognition model is based on the fusion features calculated in step 3 to predict the driving state recognition results.
[0023] In addition, based on the above-mentioned driving state recognition method based on multimodal fusion, the present invention also proposes a corresponding driving state recognition system based on multimodal fusion, which adopts the following technical solutions:
[0024] A driving state recognition system based on multimodal fusion, comprising a sensing device and a computer device; wherein the sensing device comprises a non-invasive physiological bracelet, a pedal sensor and an inertial navigation sensor;
[0025] The non-invasive physiological bracelet is used to collect physiological and psychological signal data and wrist movement information, wherein the physiological signal and psychological signal data are collected by the built-in sensors in the bracelet; the wrist movement information is collected by the accelerometer and gyroscope in the bracelet;
[0026] Driving performance data is collected by the vehicle’s pedal sensors and inertial navigation sensors;
[0027] Each sensor device is connected to a computer device and is used to transmit the collected data to the computer device;
[0028] The computer device includes a memory and one or more processors; the memory stores executable codes, and when the processor executes the executable codes, it is used to implement the driving state recognition method based on multimodal fusion as described above.
[0029] The present invention has the following advantages:
[0030] As described above, the present invention relates to a driving state recognition method based on multimodal fusion, which collects the driver's wrist motion information through a non-invasive physiological bracelet, providing an additional dimension for driving state recognition. The introduction of wrist motion information can more accurately capture the driver's micro-motion changes during dynamic driving, enhance the monitoring of the driver's state, and combine wrist motion information with traditional physiological signals (including HRV, EDA) and driving performance data to comprehensively improve the accuracy and robustness of driving state recognition. In addition, the present invention also proposes a feature dynamic weighted model (abbreviated as LCA-Transformer) based on LSTM module, cross-attention mechanism module Cross-Attention and Transformer module. In the LCA-Transformer model, the LSTM module is used to capture the long-term dependency in the time series, and the Cross-Attention module weights the features between different modes through the self-attention mechanism, so that the dynamic importance of each feature can be adjusted according to different driving states. This mechanism significantly improves the adaptability and performance of the model in complex driving scenarios, and can accurately identify the driver's driving state. Based on the traditional LSTM network, the present invention combines the multi-head self-attention mechanism of Transformer to achieve deep interaction and feature enhancement of different modal data. In addition, through the Cross-Attention mechanism, a stronger correlation is established between the features of each modality. The LCA-Transformer model can make full use of multimodal information and enhance the understanding of nonlinear relationships between features, thereby facilitating the improvement of the classification accuracy of driving status. Through the comprehensive analysis of multimodal features, the present invention uses the XGBoost model for classification. The model can accurately distinguish different driving states based on wrist motion information, driving performance data, and psychological and physiological signal characteristics. The present invention can accurately perform three-classification tasks, that is, identify the three driving states of the driver: normal driving, fatigue driving, and distracted driving. This three-classification capability provides important support for real-time driver behavior monitoring and can effectively prevent the occurrence of traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1is a processing flow chart of a driving state recognition method based on multimodal fusion in an embodiment of the present invention;
[0032] Figure 2 It is a network structure diagram of the feature dynamic weighting model LCA-Transformer in an embodiment of the present invention;
[0033] Figure 3 It is a processing flow chart of the feature dynamic weighting model LCA-Transformer in an embodiment of the present invention;
[0034] Figure 4 This is a processing flow chart of a driving state recognition model based on XGBoost in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0036] Example 1
[0037] This embodiment 1 describes a driving state recognition method based on multimodal fusion, such as Figure 1 As shown, the driving state recognition method based on multimodal fusion includes the following steps:
[0038] Step 1. Comprehensively collect multimodal data of psychological signal data, physiological signal data, driving performance data, and wrist motion information, and preprocess the collected multimodal data.
[0039] The psychological signal data and physiological signal data collected in this embodiment refer to data such as HRV (electrocardiogram), sEMG (electromyography) and EDA (skin electrical signal) collected from the driver through a non-invasive physiological bracelet.
[0040] The wrist motion information collected in this embodiment, specifically, refers to wrist X, Y, Z three-axis coordinate data, three-axis angular velocity and acceleration data obtained through the built-in accelerometer and gyroscope of the non-invasive physiological bracelet.
[0041] Traditional physiological data collection equipment such as electrocardiographs and electromyographs cause great interference to drivers and have a greater impact on driving operations. Non-invasive physiological bracelets cause less interference to drivers, are suitable for driving status research and can be worn for a long time.
[0042] The driving performance data collected in this embodiment refers to the driving speed, yaw rate, lane deviation, acceleration, steering wheel angle, acceleration, brake pedal coefficient, etc. obtained through inertial navigation and pedal sensors.
[0043] Driving performance data is collected by the vehicle's pedal sensors and inertial navigation sensors.
[0044] After obtaining the data of the above-mentioned modes, the data is first preprocessed as follows:
[0045] First, data cleaning is performed to remove missing values and outliers to ensure the integrity and accuracy of the data. The upper and lower quartiles of the data and the boundaries of outliers are identified through box plots, and outliers are deleted.
[0046] The Kalman filter algorithm is used to denoise the data and remove high-frequency noise to improve the stability and reliability of the signal, thereby laying a more robust data foundation for subsequent analysis. Features with different dimensions are standardized or normalized to put them on the same scale to facilitate effective comparison and analysis between features.
[0047] The above data preprocessing operation in this embodiment is relatively conventional and will not be further described here.
[0048] This embodiment collects a variety of data, including psychological and physiological signals (such as electrocardiogram, electrodermal electricity, electromyography), driving performance, and wrist movement information, and performs refined denoising and standardization on these data. Through multi-dimensional joint analysis, these signals provide a more accurate driver status assessment. The present invention solves the fusion problem of multiple signal sources by integrating multiple signals, and performs effective data preprocessing, thereby ensuring high-accuracy model training.
[0049] Step 2. Further feature extraction is performed on the preprocessed multimodal data, and the features with significant differences under different driving conditions are screened out from the feature extraction results through Friedman test and Bonferroni correction.
[0050] The driving performance data, wrist motion information, EDA, and HRV data were feature extracted with a time window of 30 s and a step size of 5 s. A total of 28 HRV features, 21 EDA features, 12 driving performance features, and 12 wrist features were extracted.
[0051] Tables 1 to 4 respectively show the names of the extracted HRV features, EDA features, driving performance features, and wrist motion information features, as well as the descriptions of the features (i.e., the physical meanings they represent).
[0052] Table 1 HRV characteristics
[0053]
[0054] Table 2 EDA characteristics
[0055]
[0056] Table 3 Driving performance characteristics
[0057]
[0058] Table 4 Characteristics of wrist motion information
[0059]
[0060] After different features of multiple modes are extracted, further screening is performed. In this embodiment, the features with significant differences under different driving conditions are screened out from the feature extraction results through Friedman test and Bonferroni correction.
[0061] First, the Friedman test was used to evaluate the differences in each feature under different driving conditions.
[0062] The significance level of each indicator was set to α=0.05.
[0063] When the p-value of the indicator obtained by the significance test is less than 0.05, it can be considered that the driver's state factor has an impact on the indicator at the significance level of 0.05, indicating that the indicator can be used to identify the driver's bad driving state.
[0064] For HRV characteristics, Friedman test showed that the values of NNmean, RMSSD, SDNN, SDSD, NN50, PNN50, NN20, PNN20, medianNN, rangeNN, CVNN, CVSD, SDHR, Ulf, Vlf, Lf, Lf / Hf, CVI, SD1, SD2, SampEn, and ModCSI in distracted driving were greater than those in fatigue driving and normal driving, and the values of RMSSD, SDSD, CSI, SD1, and SD2 / SD1 in fatigue driving were less than those in normal driving and distracted driving. Further pairwise comparison using Bonferroni method showed that the P values of normal and distracted, fatigue and distracted in CSI, fatigue and distracted in If / Hf, and fatigue and distracted in SD2 / SD1 were greater than 0.05, and the P values of the pairwise comparison results of other indicators were all less than 0.05, so it was considered that except for CSI, SD2 / SD1, and If / Hf indicators, other HRV indicators had statistically significant differences. Therefore, 25 distinctive HRV features were finally selected.
[0065] For EDA features, Friedman test shows that in distracted driving state, the values of meanBand and maxBand are greater than those of normal driving and fatigue driving. In fatigue driving state, the values of meanEDA, medianEDA, maxEDA, trimeanEDA, energyEDA, minBand, and Hurst are less than those of normal driving and distracted driving. Except for skewEDA, almost all EDA features have statistically significant differences in normal driving, fatigue driving, and distracted driving, which indicates that EDA features differ in bad driving states, further illustrating the effectiveness of EDA features in identifying bad driving states. Further pairwise comparison using Bonferroni method shows that except for normal and fatigue in meanEDA and medianEDA, and fatigue and distraction in minEDA, maxEDA, trimeanEDA, and energyEDA, the P values of pairwise comparison results of other indicators are less than 0.05, so it is believed that except for meanEDA, medianEDA, minEDA, maxEDA, trimeanEDA, and energyEDA, other EDA indicators have statistically significant differences. Therefore, 14 EDA features were finally selected.
[0066] For driving performance characteristics, Friedman test indicated that in distracted driving state, except for ya_mean, xa_mean, xa_std, LFD_std, SWA_mean, and vYR_mean, other driving performance characteristics showed statistically significant differences. The values of V_std and SWA_std in distracted driving state were greater than those in normal driving and fatigue driving. Further pairwise comparison using Bonferroni method showed that except for fatigue and distraction in V_std, normal and fatigue, fatigue and distraction in ya_std, normal and fatigue, fatigue and distraction in SWA_std, the P values of the pairwise comparison results of other indicators were less than 0.05, so it was considered that except for V_std, ya_std, and SWA_std indicators, other driving performance indicators had statistically significant differences. Therefore, 3 driving performance characteristics were finally selected.
[0067] For wrist motion information features, the Friedman test showed that almost all wrist motion indicators showed statistical differences. The values of GYRO_Y_mean and GYRO_Z_mean were greater than those of normal driving and fatigue driving in distracted driving, and the values of GYRO_X_std, GYRO_Z_std, ACC_X_std, ACC_Y_std, and ACC_Z_std were less than those of normal driving and distracted driving. Further pairwise comparisons using the Bonferroni method showed that, except for the normal and distracted values of GYRO_Y_mean, the P values of the pairwise comparisons of other indicators were less than 0.05, so except for GYRO_Y_mean, other wrist motion information indicators had statistically significant differences. Therefore, 11 wrist motion features were finally selected as the input of the model.
[0068] In summary, through the difference analysis of HRV features, EDA features, driving performance features and wrist motion information features, a total of 53 features were screened out, of which there were 25 HRV features, namely NNmean, RMSSD, SDNN, SDSD, NN50, PNN50, NN20, PNN20, medianNN, rangeNN, CVNN, CVSD, SDHR, Ulf, Vlf, Lf, CVI, SD1, SD2, SampEn, Mod CSI, meanHR, maxHR, minHR, Hf; there are 14 EDA features, namely meanBand, maxBand, minBand, SampEnEDA, Hurst, stdEDA, rangeEDA, IQREDA, trimeanBand, kurtBand, skewBand, medianBand, IQRBand, midhingeBand; there are 3 driving performance features, namely V_mean, LFD_mean, vYR_std; there are 11 wrist motion information features, namely ACC_X_mean, ACC_X_std, ACC_Y_mean, ACC_Y_std, ACC_Z_mean, ACC_Z_std, GYRO_X_mean, GYRO_X_std, GYRO_Y_std, GYRO_Z_mean, GYRO_Z_std.
[0069] The present invention extracts multiple features (such as 28 HRV features, 21 EDA features, etc.) from the collected data, and screens out features with significant differences under different driving conditions through Friedman test and Bonferroni correction. This statistical method ensures the efficiency and accuracy of feature selection and effectively improves the discrimination of the model.
[0070] Step 3. Build a feature dynamic weighting model (LSTM-Cross-Attention-Transformer, LCA-Transformer for short), such as Figure 2 As shown, it is used to capture the dynamic changes of features and the correlation between features under driving conditions (such as normal, fatigue, distraction) and adjust the importance weights of features in real time according to the changes in feature states.
[0071] The feature dynamic weighted model includes an LSTM module, a cross-attention mechanism module, and a Transformer module.
[0072] The LSTM module is used to capture long-term dependencies in time series.
[0073] The cross-attention mechanism module weights the features between different modalities through the self-attention mechanism, so that the dynamic importance of each feature can be adjusted according to different driving conditions.
[0074] The multi-head self-attention mechanism of the Transformer module is used to achieve deep interaction and feature enhancement of data of different modalities.
[0075] The overall processing flow of the LCA-Transformer model is as follows:
[0076] First, the filtered features are fed into the LSTM module, which captures the temporal dependencies between different features.
[0077] Each feature input into the LSTM module will undergo a time series modeling process in the LSTM module. LSTM learns the dynamic changes of the driver under different behavior states from the time series pattern of each feature.
[0078] The LSTM module outputs the hidden state vector of each feature at different time steps. The hidden state vectors output by the LSTM module are divided into four categories, corresponding to physiological signal data, psychological signal data, driving performance data and wrist movement information. Each type of hidden state vector contains the time dependency information of each feature in the corresponding modality.
[0079] The four types of hidden state vectors are input into the cross-attention mechanism module, which calculates the attention weights between multiple modalities, dynamically adjusts the contribution of each modality feature, and fuses information from different modalities.
[0080] The output features of the cross-attention mechanism module are input into the Transformer module, and the multi-head self-attention mechanism of the Transformer module is further used to achieve deep interaction and feature enhancement of data of different modalities to obtain fused features, namely weighted physiological signal data, psychological signal data, driving performance data and wrist motion information features.
[0081] The LCA-Transformer model in this embodiment dynamically assigns feature weights to features. The dynamic weighting method can significantly enhance the role of key features, reduce the interference of irrelevant features, improve the accuracy and robustness of the classification model, and adapt to data distribution fluctuations in complex driving scenarios, providing more accurate support for the identification of adverse driving conditions.
[0082] like Figure 3 As shown, the various sub-modules that make up the LCA-Transformer model are described in detail below.
[0083] First, the filtered features are input into the LSTM module, which captures the temporal dependencies between different features.
[0084] Each input feature The LSTM module undergoes a temporal modeling process, where LSTM learns the dynamic changes of the driver under different behavioral states from the temporal pattern of each feature.
[0085] The output of LSTM is the hidden state at each time step , which preserves the dynamic changes in the time series:
[0086] .
[0087] in, represents the LSTM hidden state matrix, Preserve time-dependent properties.
[0088] In the LSTM module, HRV features (such as RMSSD) capture the changing trend during fatigue driving, especially the decrease of RMSSD, through the recursive structure of the network; EDA features reflect the driver's skin electrical response, which may fluctuate due to anxiety or tension. LSTM can capture the increase in EDA fluctuations under distracted and fatigued driving conditions; driving performance features can reflect the driver's control situation, and LSTM learns the time series pattern of unstable vehicle speed during fatigue driving; at the same time, LSTM can also capture the timing fluctuations of the driver's frequently changing hand movements, especially in poor driving conditions.
[0089] LSTM captures these changes and trends in time series data through a recursive structure and outputs the hidden state vector of each feature at different time steps, which are h HRV 、h EDA 、h Perf 、h Wrist , which respectively contain the time-dependent information of the electrocardiogram signal features, skin electrical signal features, driving performance features and wrist movement information features corresponding to each other.
[0090] Among them, h HRV Corresponding to the 25 selected HRV features, h EDA Corresponding to the 14 EDA features selected, h Perf Corresponding to the three selected driving performance characteristics, h Wrist It corresponds to the 11 selected wrist motion information features.
[0091] The four types of hidden state vectors output by the LSTM module, that is, the four types of features, are input into the cross-attention mechanism module and processed in the Cross-Attention module.
[0092] The Cross-Attention module first interactively processes the four types of features. Specifically, it establishes the relationship between the features of different modes by calculating the query, key, and value matrices.
[0093] Step I. In order to calculate the interaction between different modalities, firstly, from each type of feature h HRV 、h EDA 、h Perf 、h Wrist Three matrices are extracted from the query matrix Q, the key matrix K, and the value matrix V.
[0094] These matrices are used to compute the attention weights and weighted sums.
[0095] Step II. Select any query matrix Q from the four calculated query matrices Q, and select a key matrix K from the four calculated key matrices K whose feature source is different from that of the query matrix Q.
[0096] The criss-cross attention mechanism uses the similarity between the selected query matrix Q and the key matrix K to calculate the attention weights.
[0097] The dot product between the query matrix Q and the key matrix K represents the similarity between them, and softmax is applied to obtain the normalized attention distribution so that each element becomes a probability, indicating the dependency between features.
[0098] Taking the leftmost branch in the Cross-Attention module as an example, the query matrix Q comes from the HRV feature, the key matrix K comes from the EDA feature, and the value matrix V comes from the driving performance feature. The formula for generating the matrix is as follows:
[0099] , , .
[0100] in , , is the trainable weight matrix, , , They respectively indicate that the query matrix Q comes from the HRV features, the key matrix K comes from the EDA features, and the value matrix V comes from the driving performance features.
[0101] The Cross-Attention module calculates the attention weights by calculating the similarity between the query matrix and the key matrix.
[0102] The dot product between the query matrix Q and the key matrix K represents the similarity between them. Softmax is applied to obtain the normalized attention distribution so that each element becomes a probability, indicating the dependency between features.
[0103] .
[0104] in, is the dimension of the key, used as a scaling factor to prevent the dot product value from being too large.
[0105] It shows the relationship between HRV features and EDA features, reflecting their importance. Similarly, the attention weight matrix between other modalities, such as HRV and Perf, HRV and Wrist, Perf and EDA, etc., is calculated.
[0106] Step III. Further use the calculated attention weight matrix to perform weighted summation on the value matrix to obtain a new feature representation. The weighted feature representation contains the interaction information between different features.
[0107] The feature sources of the value matrix V are different from those of the query matrix Q and the key matrix K mentioned above.
[0108] Taking the leftmost branch in the Cross-Attention module as an example, the value matrix V selected by this branch comes from the driving performance features (as mentioned above, it is different from the query matrix , key matrix from different sources).
[0109] Using the calculated attention matrix, the value matrix Perform weighted summation to obtain a new feature representation:
[0110] .
[0111] From the formula, we can see that the weighted feature representation , which contains the interactive information between different features. Similarly, in this way, the weighted feature representation of other modal combinations can be obtained , , wait.
[0112] Step IV. Repeat steps II to III above until all combinations of query matrix Q, key matrix K, and value matrix V are traversed to obtain weighted feature representations of multiple different modal combinations.
[0113] The feature sources of the query matrix Q, key matrix K, and value matrix V used in each calculation process are different.
[0114] Step V. Connect the multiple weighted feature representations obtained in step IV to obtain the output feature , where the fusion feature The expression is: .
[0115] The Cross-Attention module can effectively calculate the attention weights between multiple modalities, dynamically adjust the contribution of each modal feature, and fuse information from different modalities, thereby enhancing the expressiveness of the model and improving the effect of multimodal learning.
[0116] Feature representation after Cross-Attention module , which has combined the interactive information between different modalities. These features will serve as the input of the Transformer module , recorded as: .
[0117] The core of Transformer is the self-attention mechanism, which requires the construction of query (Query, Q), key (Key, K), and value (Value, V) matrices.
[0118] First, build the Transformer query, key, and value matrices.
[0119] The characteristics Input to the Transformer module, enhance the global dependency of features, and build , , matrix:
[0120] .
[0121] in, , , are the weights of the query matrix, key matrix, and value matrix respectively; , , Representing query, key, and value respectively, is accomplished by Multiply , , Got it.
[0122] Calculate the self-attention matrix , the formula is as follows:
[0123] .
[0124] in is the dot product of the query and the key, indicating the similarity between features, is the normalized scaling factor, and finally the attention distribution is obtained through the softmax function , making it a probability distribution that represents the importance of each feature.
[0125] The self-attention matrix With value matrix Weighted summation to obtain the fused global feature representation :
[0126] .
[0127] in is the attention matrix, which contains the importance weights between each feature. It is a value matrix, which contains the information of the feature itself. The fused global feature representation is obtained through weighted summation operation. , including the interdependencies between features.
[0128] At the output of Transformer In the final feature representation, feature weighting is performed to calculate feature weights. Input a fully connected layer MLP to perform feature weighting coefficients, and calculate them through the softmax function:
[0129] .
[0130] in, Represents the dynamic weight coefficient of each feature, , is the dynamic weighting coefficient corresponding to each screened feature, T is the number of features, and the sum of all dynamic weighting coefficients after normalization is 1.
[0131] Use the weight coefficient to perform weighted fusion on the original feature X to obtain the weighted feature express:
[0132] .
[0133] in, , represents the original eigenvalue, represents the weighted eigenvalue, , represents the weighted eigenvalue.
[0134] Through the LCA-Transformer model, the system can dynamically adjust the importance weight of each feature according to changes in the driver's state. This dynamic feature weighting mechanism can effectively enhance the role of key features while reducing the interference of irrelevant features, thereby improving the accuracy, stability and robustness of the model and adapting to complex and changing driving scenarios.
[0135] Step 4. Build a driving state recognition model based on XGBoost and perform model training. The trained driving state recognition model is based on the fusion features calculated in step 3 to predict the driving state recognition results.
[0136] In the previous step, the accuracy of the driving state recognition model is improved by dynamically adjusting the feature weighting coefficient input through the LCA-Transformer model, thereby better utilizing the recognition ability of each feature for a specific driving state.
[0137] When you start training an XGBoost model, you first go through a simple initialization step to establish initial predictions. The initial model is usually a constant that reflects the global statistics of the training data.
[0138] Assuming that most of the samples in the current training set belong to "normal driving", the initial prediction value may tend to predict all samples as "normal driving". The purpose of this step is to provide a starting point for the construction of subsequent decision trees. Although the initial model has limited accuracy for complex classification tasks, it lays the foundation for the subsequent residual fitting through gradient boosting.
[0139] The XGBoost model achieves accurate classification of driving status (normal driving, fatigue driving, distracted driving) by comprehensively analyzing weighted multimodal features (such as wrist movement information, psychological signal features, physiological signal features, and driving performance). During the training process, the model minimizes the objective function through the gradient boosting algorithm, gradually optimizes the importance of features, and adjusts the split nodes based on the feature contribution value to distinguish different driving states.
[0140] The processing process of the driving state recognition model based on XGBoost is as follows:
[0141] I. Gradient boosting. XGBoost gradually optimizes model performance through the gradient boosting algorithm. The core idea is to fit the residual of the current model prediction error each iteration and gradually correct the model's prediction ability.
[0142] First, an initial model is generated using the mean of the input feature samples. At this time, the model may have a poor prediction of the target variable. The initial model is regarded as the first decision tree. When constructing new leaves in each round, XGBoost calculates the residual of the existing model, that is, the error between the true value and the predicted value.
[0143] These residuals measure the inadequacy of the current model and are defined as follows:
[0144] .
[0145] in, represents the residual, Represents the true value, that is, the weighted eigenvalue. Represents the current prediction value calculated by the model based on the input features. The existing model refers to the cumulative model of XGBoost in each iteration, that is, the set of all decision trees that have been trained at the current moment. These residuals are the targets that the next tree needs to fit.
[0146] In each iteration, XGBoost calculates the gradient of the current prediction value and the second-order gradient .
[0147] .
[0148] When the driver's state is "drowsy driving" but is predicted to be "normal driving", the residual and gradient The new tree will adjust the model predictions based on this information.
[0149] XGBoost uses the gradient boosting algorithm to gradually optimize the model parameters during training. Each round of tree training will be adjusted according to the residual of the previous round to learn the contribution of different features. During the training process, XGBoost will evaluate the impact of each feature on the classification results, and more important features will be given greater weight at the splitting nodes of the tree.
[0150] II. Decision tree construction.
[0151] XGBoost takes feature matrix is the input; feature matrix Each column represents a different weighted feature (such as wrist movement feature, driving performance feature, physiological signal feature, psychological signal feature).
[0152] The construction of the decision tree starts from the root node, selects features and split points in turn, calculates the contribution of each feature to the current classification through the gain function, and selects the feature with the largest gain and its split point for data division.
[0153] Gain function The calculation formula is:
[0154] .
[0155] Among them, g represents the target gradient of each sample, which is the first-order derivative and is used to update the predicted value. Each sample refers to a feature vector, which is an input data for model training and is one of all feature subsets. h represents the target second-order derivative of each sample, which is used to adjust the learning rate. λ represents the regularization parameter, which is used to control the complexity of the model. γ represents the splitting cost, which is used to limit the node splitting from being too complex. Represents the first-order derivative of a sample of a child node after splitting, Represents the second-order derivative of a sample of a child node after splitting.
[0156] The gain indicates the degree of optimization of the loss function after splitting. The model will select the feature with the largest gain and its splitting point.
[0157] Traverse all input features, select different splitting thresholds, and select the features and splitting points with the largest gain as the basis for splitting the current node; high-value wrist movement features may help distinguish normal driving from other states, so they may appear at the root node of the decision tree. Driving performance features are very important in identifying fatigue and distracted driving. Higher LFD_mean values usually indicate fatigue driving states, while larger vYR_std values may indicate distracted driving. Physiological signal features reveal signs of fatigue or distracted driving through the changes in heart rate they reflect. New decision trees are constructed to fit these residuals, and each newly added tree will better compensate for the errors of the existing model by splitting nodes; the tree construction process is similar to the first round of tree construction, that is, splitting nodes are selected based on the splitting gain of the features.
[0158] The output of the new tree will be weighted and added to the current prediction result. The update formula is as follows:
[0159] ;
[0160] in, represents the predicted value of the tth time, represents the predicted value of the previous round, Represents the learning rate, which is used to control the impact of each new tree on the overall model. Represents the output of the tth decision tree.
[0161] If the current model under-predicts distracted driving, the new tree may correct the prediction by paying more attention to vYR_std (because this parameter has a positive effect on distracted driving identification, so it increases the attention to this feature).
[0162] The above steps will be repeated to build new decision trees in rounds, and each round will fit a new tree based on the current residual. In this way, the model gradually approaches the optimal solution.
[0163] III. Classification.
[0164] XGBoost converts the score of each sample into a probability distribution through the Softmax function:
[0165] .
[0166] in, is the probability that the sample belongs to a category, K=3, corresponding to normal driving, fatigue driving, and distracted driving respectively; represents the score of the sample belonging to category j, Represents the score of the sample in category k; through the Softmax function, XGBoost normalizes the scores of different categories into probability distribution and selects the category with the maximum probability as the final classification result.
[0167] Define the loss function. XGBoost uses the objective function to measure the difference between the predicted value and the true value, and adds a regularization term to prevent the model from overfitting. The objective function is defined as:
[0168] .
[0169] in, is the prediction error, is a regularization term used to limit model complexity to prevent overfitting.
[0170] Among the input features, high values of GYRO_Z_mean will significantly affect the classification of normal driving categories, while changes in NNmean may affect the distinction of fatigue driving. The optimization of the objective function can capture the actual contribution of these features to the classification task.
[0171] In order to verify the model performance, XGBoost uses cross-validation to evaluate classification indicators, such as accuracy, recall, F1-score, etc. In driving state recognition, a higher GYRO_Z_mean value supports normal driving classification, a larger vYR_std value supports distracted driving classification, and the change trend of NNmean helps to distinguish fatigue driving status.
[0172] Through comprehensive analysis of these features, the XGBoost classification model achieves efficient classification of driving status.
[0173] Through comprehensive analysis of multimodal features and optimization of the XGBoost classification model, the present invention can efficiently and accurately identify the driver's status. The XGBoost classification model can accurately distinguish three driving states, namely normal driving, fatigue driving and distracted driving, thereby providing more accurate monitoring and early warning for road safety and timely preventing dangerous driving behaviors.
[0174] Based on the detected driving status, the system can improve driving safety through real-time intervention.
[0175] Under normal driving conditions, the system provides voice reminders or driving habit optimization suggestions to help the driver maintain a good state.
[0176] If fatigue driving is detected, the system can remind the driver through voice alarms, seat vibrations or adjustment of the interior environment, and suggest stopping to rest; if fatigue continues, the system can limit the vehicle speed or switch to assisted driving mode.
[0177] For distracted driving, the system reduces risks by blocking mobile phone notifications, limiting in-car entertainment functions or triggering assisted driving functions, and issues real-time reminders to correct the driver's behavior.
[0178] In addition, the system can dynamically adjust the intervention intensity according to the level of danger, gradually upgrading from mild prompts to mandatory intervention, and generate behavior reports based on driving history data to provide long-term safety management support.
[0179] This comprehensive intervention strategy can effectively reduce driving risks and ensure road traffic safety.
[0180] Example 2
[0181] This embodiment 2 describes a driving state recognition system based on multimodal fusion, including a sensing device and a computer device; wherein the sensing device includes a non-invasive physiological bracelet, a pedal sensor, and an inertial navigation sensor.
[0182] The non-invasive physiological bracelet is used to collect physiological and psychological signal data and wrist movement information, where the physiological signal and psychological signal data are collected by the built-in sensors in the bracelet; the wrist movement information is collected by the accelerometer and gyroscope in the bracelet.
[0183] Driving performance data is collected by the vehicle's pedal sensors and inertial navigation sensors.
[0184] Each sensor device is connected to the computer device and is used to transmit the collected data to the computer device.
[0185] The computer device includes a memory and one or more processors; the memory stores executable codes, and when the processor executes the executable codes, it is used to implement the driving state recognition method based on multimodal fusion as described above.
[0186] In this embodiment, the computer device is any device or apparatus with data processing capability, which will not be described in detail here.
[0187] The driving state recognition method and system based on multimodal fusion described in the present invention can be applied to driver state recognition, intelligent assisted driving system, vehicle monitoring system and driving behavior evaluation in the field of traffic safety and driving behavior analysis. Through the dynamic feature fusion method proposed in the present invention, the present invention can be widely used to improve driving safety, optimize driving experience and support high-level automatic driving and driving behavior, thereby providing important technical support for traffic safety management.
[0188] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with the field under the guidance of this specification fall within the essential scope of this specification and should be protected by the present invention.
Claims
1. A driving state recognition method based on multimodal fusion, characterized in that: The steps include: Step 1. Comprehensively collect multimodal data information of psychological signal data, physiological signal data, driving performance data and wrist movement information, and pre-process the collected multimodal data; Driving performance data is collected by the vehicle's pedal sensors and inertial navigation sensors; Step 2. Further feature extraction is performed on the preprocessed multimodal data, and the features with significant differences under different driving conditions are screened out from the feature extraction results through Friedman test and Bonferroni correction; Step 3. Build a dynamic feature weighting model to capture the dynamic changes of features and the correlation between features under driving conditions and adjust the importance weight of features in real time according to the changes in feature states; The feature dynamic weighting model includes LSTM module, cross attention mechanism module and Transformer module; The selected features are fed into the LSTM module, which captures the temporal dependencies between different features. Each feature input into the LSTM module will go through the temporal modeling process in the LSTM module. The LSTM learns the dynamic changes of the driver under different behavior states from the temporal pattern of each feature. The LSTM module outputs the hidden state vector of each feature at different time steps; the hidden state vector output by the LSTM module is divided into four categories, corresponding to physiological signal data, psychological signal data, driving performance data and wrist movement information; Among them, each type of hidden state vector contains the time-dependent information of each feature in the corresponding modality; The four types of hidden state vectors are input into the cross-attention mechanism module, which calculates the attention weights between multiple modalities, dynamically adjusts the contribution of each modality feature, and integrates the information of different modalities; The output features of the cross-attention mechanism module are input into the Transformer module, and the multi-head self-attention mechanism of the Transformer module is further used to achieve deep interaction and feature enhancement of different modal data. final On the top, perform feature weighting to calculate feature weights, and convert Z final Input a fully connected layer MLP to perform feature weighting coefficients, and calculate the dynamic weighting coefficient of each feature through the softmax function. The sum of all dynamic weighting coefficients after normalization is 1. Use the dynamic weighting coefficient of each feature to perform weighted fusion on the original features to obtain the weighted feature X′, that is, the weighted physiological signal data, psychological signal data, driving performance data, and wrist motion information features; X′=X⊙W; Where X = [x1, x2, ..., x T ],x1,x2,...,x T represents the original eigenvalue, X′ represents the weighted eigenvalue, X′=[x1w1,x2w2,…,x T w T ]=[x′1,x′2,...,x′ T ],x′1,x′2,...,x′ T represents the weighted eigenvalue; Step 4. Build a driving state recognition model based on XGBoost and perform model training. The trained driving state recognition model is based on the weighted feature X′ calculated in step 3 to predict the driving state recognition result.
2. The driving state recognition method based on multimodal fusion according to claim 1 is characterized in that: The features filtered out in step 2 are input into the LSTM module, which captures the temporal dependencies between different features. Where each input feature X = [x1, x2, ..., x T ]In the LSTM module, the temporal modeling process is carried out. LSTM learns the dynamic changes of the driver under different behavior states from the temporal pattern of each feature; where x1,x2,...,x T Represents the features selected in step 2; The output of LSTM is the hidden state h at each time step LSTM , which retains the dynamic change information in the time series: h LSTM =LSTM(X); Among them, h LSTM represents the LSTM hidden state matrix, h LSTM Preserve time-dependent properties.
3. The driving state recognition method based on multimodal fusion according to claim 1 is characterized in that: The four hidden state vectors that define the output of the LSTM module are h HRV 、h EDA 、h Perf 、h Wrist , corresponding to ECG signal features, skin electrical signal features, driving performance features, and wrist motion information features respectively; The above four types of hidden state vectors, i.e. four types of features, are all input into the cross attention mechanism module and processed as follows: Step I. First, from each class of features h HRV 、h EDA 、h Perf 、h Wrist Three matrices are extracted from the query matrix Q, the key matrix K, and the sum matrix V, which are used to calculate the attention weights and weighted sums; Step II. randomly select a query matrix Q from the four query matrices Q obtained by calculation, and select a key matrix K from the four key matrices K obtained by calculation, the key matrix K having a feature source different from that of the query matrix Q; The cross attention mechanism uses the similarity between the selected query matrix Q and the key matrix K to calculate the attention weight; The dot product between the query matrix Q and the key matrix K represents the similarity between them, and softmax is applied to obtain the normalized attention distribution so that each element becomes a probability, indicating the dependency between features; Step III. Further use the calculated attention weight matrix to perform weighted summation on the value matrix to obtain a new feature representation. The weighted feature representation contains the interaction information between different features. Among them, the feature sources of the value matrix V are different from those of the query matrix Q and the key matrix K mentioned above; Step IV. Repeat the above steps II to III until all combinations of query matrix Q, key matrix K, and value matrix V are traversed to obtain weighted feature representations of multiple different modal combinations; The feature sources of the query matrix Q, key matrix K, and value matrix V used in each calculation process are different; Step V. Connect the multiple weighted feature representations obtained in step IV to obtain output features.
4. The driving state recognition method based on multimodal fusion according to claim 1, characterized in that: The output feature Z of the cross attention mechanism module fusion Input to the Transformer module and serve as the input feature X of the Transformer module input , that is, X input =Z fusion , and construct Q′, K′, V′ matrices; Q′=X input W Q ′,K′=X input W K ′,V′=X input W V ′; Among them, W Q ′、W K ′、W V ′ are the weights of query matrix, key matrix and value matrix respectively; Q′, K′, V′ represent query, key and value respectively, which are obtained by input Multiply W Q ′、W K ′、W V 'obtained; Calculate the self-attention matrix A′, the formula is as follows: Among them, Q′K′ T is the dot product of the query and the key, indicating the similarity between features, is the scaling factor for normalization, and finally the attention distribution A′ is obtained through the softmax function, making it a probability distribution that indicates the importance of each feature; The weighted sum of the self-attention matrix A′ and the value matrix V′ is used to obtain the fused global feature representation Z final : From final =A′V′; At the Transformer output Z final In the final feature representation, feature weighting is performed to calculate feature weights, and Z final Input a fully connected layer MLP to perform feature weighting coefficients, and calculate them through the softmax function: W=softmax(MLP(Z final )); Where W represents the dynamic weighting coefficient of each feature, W = [w1, w2, ..., w T ],w1,w2,...,w T is the dynamic weighting coefficient corresponding to each screened feature, T is the number of features, and the sum of all dynamic weighting coefficients after normalization is 1.
5. The driving state recognition method based on multimodal fusion according to claim 4 is characterized in that: In step 4, the processing process of the driving state recognition model based on XGBoost is as follows: First, an initial model is generated using the mean of the input feature samples. The initial model is regarded as the first decision tree. When constructing new leaves in each round, XGBoost calculates the residual of the existing model, that is, the error between the true value and the predicted value, which is defined as follows: Among them, r i Residual, y i represents the true value, Represents the current prediction value calculated by the model based on the input features; in each iteration, XGBoost calculates the gradient g of the current prediction value i and the second-order gradient h i ; XGBoost takes the feature matrix X′ as input, and each column of the feature matrix X′ represents a different weighted feature; The construction of the decision tree starts from the root node, selects features and split points in turn, calculates the contribution of each feature to the current classification through the gain function, and selects the feature with the largest gain and its split point for data division; The calculation formula of the gain function Gain is: Among them, g represents the target gradient of each sample, that is, each feature vector, which is the first-order derivative and is used to update the prediction value; h represents the target second-order derivative of each sample, which is used to adjust the learning rate; λ represents the regularization parameter, which is used to control the complexity of the model; γ represents the splitting cost, which is used to limit the complexity of node splitting; g′ represents the first-order derivative of the sample of a child node after splitting, and h′ represents the second-order derivative of the sample of a child node after splitting; The model will select the feature with the largest gain and its split point; it will traverse all input features, select different split thresholds, and select the feature with the largest gain and split point as the basis for splitting the current node; New decision trees are constructed to fit the residuals. Each new tree will better compensate for the errors of the existing model by splitting nodes. The output of the new tree will be weighted and added to the current prediction results. The update formula is as follows: in, represents the predicted value of round t, represents the predicted value of the previous round, η represents the learning rate, which is used to control the impact of each new tree on the overall model, and f(X i ′) represents the output of the t-th decision tree; Repeat the above steps to build new decision trees round by round. In each round, a new tree is fitted according to the current residual, and the model gradually approaches the optimal solution. XGBoost converts the score of each sample into a probability distribution through the Softmax function: Among them, P(c k |x) is the probability that the sample belongs to a category, K = 3, corresponding to normal driving, fatigue driving, and distracted driving respectively; y j Indicates the score of the sample belonging to category j, y k Represents the score of the sample in category k; Through the Softmax function, the scores of different categories are normalized into probability distribution; The category with the maximum probability is selected as the final classification result of the driving state recognition model based on XGBoost.
6. The driving state recognition method based on multimodal fusion according to claim 5 is characterized in that: In step 4, the loss function is defined. XGBoost uses the objective function to measure the difference between the predicted value and the true value, and adds a regularization term to prevent the model from overfitting. The objective function Defined as: in, is the prediction error, Ω(f t ) is a regularization term used to limit model complexity to prevent overfitting.
7. The driving state recognition method based on multimodal fusion according to claim 1 is characterized in that: In step 1, the wrist movement information collected includes wrist X, Y, Z three-axis coordinate data, three-axis angular velocity and acceleration data obtained by the built-in accelerometer and gyroscope of the physiological bracelet.
8. The driving state recognition method based on multimodal fusion according to claim 1, characterized in that: In step 1, the process of preprocessing the collected multimodal data is as follows: First, data cleaning is performed to remove missing values and outliers to ensure the integrity and accuracy of the data. The upper and lower quartiles of the data and the boundaries of outliers are identified through box plots, and outliers are deleted. The Kalman filter algorithm is used to denoise the data and remove high-frequency noise to improve the stability and reliability of the signal; Standardize or normalize features with different dimensions so that they are on the same scale.
9. The driving state recognition method based on multimodal fusion according to claim 1, characterized in that: In step 2, 11 wrist motion information features are screened out; The characteristics of each wrist motion information are the average value of X-axis acceleration, the standard deviation of X-axis acceleration, the average value of Y-axis acceleration, the standard deviation of Y-axis acceleration, the average value of Z-axis acceleration, the standard deviation of Z-axis acceleration, the average value of X-axis angular velocity, the standard deviation of X-axis angular velocity, the standard deviation of Y-axis angular velocity, the average value of Z-axis angular velocity, and the standard deviation of Z-axis angular velocity.
10. A driving state recognition system based on multimodal fusion, comprising a sensor device and a computer device; wherein: Sensing devices include non-invasive physiological bracelets, as well as pedal sensors and inertial navigation sensors; The non-invasive physiological bracelet is used to collect physiological and psychological signal data and wrist movement information; the physiological signal and psychological signal data are collected by the built-in sensors in the bracelet; Wrist movement information is collected through the accelerometer and gyroscope in the bracelet; Driving performance data is collected by the vehicle’s pedal sensors and inertial navigation sensors; Each sensor device is connected to the computer device and is used to transmit the collected data to the computer device; The computer device comprises a memory and one or more processors; characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the executable codes are used to implement the driving state recognition method based on multimodal fusion as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Pedestrian intention prediction method based on Transform multi-modal fusion strategy
CN118781575A
Vehicle networking driving state monitoring method and system
CN119218227A