Driver distraction degree real-time evaluation and intervention method based on multi-modal data

Through the methods of multimodal data fusion and dynamic modal weight adjustment, the driver's attention state is evaluated in real time, which solves the problem of inaccurate evaluation in existing technologies, achieves more accurate attention assessment and effective intervention, and improves driving safety.

CN120753656APending Publication Date: 2025-10-10HEFEI UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510849753.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing driver attention monitoring systems find it difficult to accurately assess the driver's attention state in real time in complex driving environments, and existing multimodal data fusion models fail to effectively consider the dynamic changes between different modalities, resulting in inaccurate assessments.

Method used

A multimodal data fusion method is adopted. By introducing the Transformer encoding module and the dynamic modal fusion module, the weight of each modal data is adjusted in real time. In combination with driving situations and physiological reactions, physiological data, eye movement data and vehicle status data are used to assess the driver's attention, and customized intervention is carried out based on the assessment results.

Benefits of technology

It improves the accuracy and comprehensiveness of driver attention assessment, enhances the system's responsiveness to changing driving environments, provides customized intervention plans, reduces negative emotions caused by excessive intervention, and ensures that drivers regain their attention in a timely manner to avoid potential driving risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120753656A_ABST
    Figure CN120753656A_ABST
Patent Text Reader

Abstract

The invention discloses a driver distraction degree real-time evaluation and intervention method based on multi-modal data, and relates to the technical field of driving safety, and the method comprises the steps: collecting the data of three modals, i.e., the physiological data, eye movement data and vehicle state of a driver, and extracting depth features through a Transform coding module; inputting the feature scalar quantities of all n time steps into a dynamic modal fusion module to calculate a driver attention degree score DAI under each time step; calculating the average value of the DAIs of the drivers in the n time steps by using a weight method, determining the distraction degree of the current driver and performing intervention; by introducing a dynamic modal fusion module and an attention mechanism, the weight of each modal data is adjusted in real time according to different driving situations and physiological reactions of a driver, and it is ensured that in the driving process, the various modal data can generate proper influences on a final attention evaluation result according to the importance and real-time changes of the modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of driving safety technology, and in particular to a real-time evaluation and intervention method for driver distraction based on multimodal data. Background Art

[0002] With the continuous development of intelligent driving technology, driver attention monitoring has become a key link in ensuring driving safety. Driver distraction is not only one of the important causes of traffic accidents, but its dynamic characteristics require real-time and accurate monitoring and intervention. Traditional driver attention monitoring methods mostly rely on simple driving behavior or physiological indicators, which are difficult to accurately reflect the driver's attention state in real time. Existing attention monitoring systems mostly use single sensor data (such as vehicle status or eye movement data) for judgment, but these methods cannot effectively cope with complex driving environments and the diversity of driver attention states.

[0003] To address this issue, in recent years, both academia and industry have conducted research on multimodal data fusion, attempting to more accurately assess the driver's attention state by integrating multiple data sources, such as eye movements, heart rate, and brain waves. However, current fusion models are relatively simple in their approach to modal fusion and weighting, and often fail to take into account the dynamic changes between different modalities. For example, some systems set the weights of all modalities to fixed values, ignoring the changes in the importance of each modality in different scenarios. This fixed weighting approach limits the potential of multimodal data, making it difficult to fully tap the potential information of each modality, resulting in poor performance of the system in complex driving environments.

[0004] Based on this, the present invention aims to provide a real-time evaluation and intervention method for driver distraction based on multimodal data to solve the problem of difficulty in fully mining the potential information of each modality. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technical problems, the purpose of the present invention is to provide a real-time assessment and intervention method for the degree of driver distraction based on multimodal data. By introducing a dynamic modal fusion module and an attention mechanism, the weight of each modal data is adjusted in real time according to different driving scenarios and the driver's physiological reactions, ensuring that during the driving process, various modal data can have an appropriate impact on the final attention assessment result based on their importance and real-time changes.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A method for real-time assessment and intervention of driver distraction based on multimodal data includes the following steps:

[0008] (1) Collecting data from three modalities: driver physiological data, eye movement data, and vehicle status, and converting the three modal data into a multimodal feature vector;

[0009] (2) Input the multimodal feature vector into the Transformer encoding module, extract the deep features at each time step, and perform dimensionality reduction on the deep features of all time steps to become scalars;

[0010] (3) The characteristic scalars of all n time steps in step (2) are input into the dynamic modal fusion module. The time gating mechanism and modal confidence evaluation are introduced into the dynamic modal fusion module. The weight of each modality at different time steps is dynamically adjusted based on the historical state, and then the driver's attention level score DAI at each time step is calculated;

[0011] (4) Calculate the driver’s attention level score DAI for the past n time steps according to step (3), and assign different weights σ to the DAI of each time step. i , weight σ i As time approaches the current time step, the average of the driver's DAI is calculated, and the average is compared with the threshold of the driver's attention level under each distraction level to determine the current driver's distraction level;

[0012] (5) Take appropriate interventions based on the driver’s level of distraction.

[0013] In the present invention, the multimodal feature vector x(t):

[0014] x(t)=[x phys (t),x eye (t),x motion (t)]

[0015] Each subvector is:

[0016] x phys (t)=[HR(t),GSR(t),RRI(t)]

[0017] x eye (t)=[Pupil(t),GazeDev(t),BlinkDur(t)]

[0018] x motion (t)=[LatAcc(t),LongAcc(t),Steering(t)]

[0019] Where HR is heart rate, GSR is galvanic skin response estimation, RRI is RR interval, Pupil is pupil diameter, GazeDev is gaze deviation, BlinkDur is blink duration, LatAcc is lateral acceleration, LongAcc is longitudinal acceleration, and Steering is steering wheel angle.

[0020] In the present invention, the scalar h in step (2) i The process of obtaining (t) is as follows:

[0021] For each mode’s eigenvector x i (t) is input into the Transformer network and encoded through multiple layers of self-attention mechanism to obtain the deep feature H of each time step i (t), expressed as:

[0022] H i (t) = Transformer(x i (t))

[0023] Among them, x i (t) is the feature of the i-th mode at time step t, i = 1, 2, 3, respectively refers to x phys (t), x eye (t), x motion (t), H i (t) is the depth feature of each time step;

[0024] Finally, the features of all time steps are reduced to a scalar h using principal component analysis i (t), expressed as:

[0025] h i (t) = PCA(H i (t))

[0026] Here, PCA(·) represents the principal component analysis dimensionality reduction process.

[0027] In the present invention, the time gating mechanism dynamically adjusts the importance of each mode at each time step according to the historical state in the time series to adapt to the volatility of the driver's state change; the change of weight depends not only on the modal performance of the current time step, but also on the trend of the previous time steps for smooth update, which can be expressed as:

[0028]

[0029] Among them, ω i (t) is the dynamic weight of the i-th mode at time step t; k i (t) is the feature vector of the i-th mode at time step t; k is the time step, and W and b are both trainable parameters.

[0030] In the present invention, the modal confidence factor is used to evaluate the signal-to-noise ratio or stability of each modal sensor in real time, and its participation is reduced when the modal quality decreases. The system performs confidence correction on the dynamic weight, and the fusion weight is updated as follows:

[0031]

[0032] Among them, ω i '(t) is the modified fusion weight; M is the total number of modalities; R i (t) is the confidence level of the ith mode at time step t, which is calculated based on the signal fluctuation amplitude:

[0033]

[0034] In the present invention, the driver's attention level score DAI is calculated as follows:

[0035]

[0036] Among them, ω i (0) is the basic attention weight of the i-th modality at the initial time step, ω i '(t) is the dynamic attention weight of the i-th modality in the subsequent time step, h i (t) is the dimensionality reduction feature table of the i-th mode at time step t.

[0037] In the present invention, the mean of the driver's attention level scores at the current time step t is calculated as follows:

[0038] At the current time step t, calculate the DAI value of the past n time steps and assign different weights σ to the DAI of each time step i ;

[0039]

[0040] DAI(k) is the distraction state index at the kth time step;

[0041] σ k is the weight of the kth time step;

[0042] n is the size of the sliding window;

[0043] DAI avg (t) is the average value of DAI within the weighted sliding window;

[0044] Weight σ k Calculated using the decaying exponential function, expressed as:

[0045] σ k =α n-k

[0046] Among them, α is the decay factor that controls the decay speed; k is the index of the corresponding time step, the current time step is t, t-1 is the previous time step, and so on.

[0047] In the present invention, the distraction levels include concentration, mild distraction, moderate distraction and severe distraction. Different reminders are given according to the distraction status through voice reminder, vibration feedback and instrument screen flashing.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] (1) Improve the accuracy and comprehensiveness of attention assessment

[0050] This method integrates multimodal data (including physiological data, eye movement data, and vehicle status information) and uses the Transformer model and self-attention mechanism to perform deep feature extraction to comprehensively assess the driver's attention state. Multimodal data fusion can provide richer driver status information, reduce the possibility of failure or noise from a single data source, and thus improve the accuracy of the assessment.

[0051] (2) The present invention introduces a dynamic modal fusion module and an attention mechanism to adjust the weights of various modal data in real time according to different driving scenarios and the driver's physiological reactions, ensuring that during the driving process, various modal data can enable the system to make adaptive adjustments according to real-time conditions, thereby improving the system's responsiveness to changing driving environments and ensuring a more accurate assessment of the driver's attention status.

[0052] (3) The present invention provides customized intervention solutions based on the driver's level of distraction (e.g., focused, mildly distracted, moderately distracted, severely distracted). Specifically, the system will precisely adjust intervention parameters (e.g., volume, vibration frequency, flashing frequency, etc.) according to different levels of distraction through a combination of voice reminders, vibration feedback, instrument panel flashing, etc., to reduce the driver's negative emotions caused by excessive intervention, while improving the effectiveness of intervention, ensuring that the driver can regain attention in a timely manner and avoid potential driving risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0055] The present invention discloses a real-time driver distraction assessment and intervention method based on multimodal data, specifically as follows:

[0056] (1) Use on-board sensors, cameras, heart rate monitors and other equipment to collect real-time data on the driver's physiological and eye movements, as well as vehicle motion.

[0057] Physiological data: using heart rate monitors, galvanic skin response sensors, electroencephalogram (EEG) headsets, etc.

[0058] Eye movement data: The driver’s eye movement trajectory and gaze point are monitored through on-board cameras and eye trackers.

[0059] Vehicle data: Obtain vehicle status information such as speed, direction, acceleration, etc. through on-board sensors.

[0060] These data are pre-processed by the on-board central control unit and converted into multimodal feature vectors:

[0061] x(t)=[x phys (t),x eye (t),x motion (t)]

[0062] Each subvector is:

[0063] x phys (t)=[HR(t),GSR(t),RRI(t)]

[0064] x eye (t)=[Pupil(t),GazeDev(t),BlinkDur(t)]

[0065] x motion (t)=[LatAcc(t),LongAcc(t),Steering(t)]

[0066] Among them, HR is Heart Rate, which indicates heart rate; GSR is Galvanic Skin Response, which indicates skin electrical response; RRI is RR-Interval, which indicates RR interval; Pupil is pupil diameter; GazeDev is Gaze Deviation, which indicates gaze deviation; BlinkDur is BlinkDuration, which indicates blink duration; LatAcc is LateralAcceleration, which indicates lateral acceleration; LongAcc is LongitudinalAcceleration, which indicates longitudinal acceleration; Steering is the steering wheel angle.

[0067] (2) Through the embedded computing unit, the Transformer model is used to perform deep encoding on the modal data collected in real time to extract deep features, and the deep features of all time steps are reduced to scalars.

[0068] This module performs real-time calculations based on embedded computing units. On hardware, the Transformer model uses the self-attention mechanism to encode the data of each modality and accelerates the model inference process through parallel computing.

[0069] The eigenvector x of each mode i Input into the Transformer network, after multi-layer self-attention mechanism encoding, the deep feature H of each time step is obtained i (t), expressed as:

[0070] H i (t) = Transformer(x i (t))

[0071] Among them, x i (t) is the feature of the i-th mode at time step t, i = 1, 2, 3, respectively refers to x phys (t), x eye (t), x motion (t), after Transformer encoding, the output H i (t) is the depth feature of each time step. In this invention, the time step t = 0.5s is taken. Finally, the features of all time steps are reduced to a scalar h using the principal component analysis method. i (t), expressed as:

[0072] h i (t) = PCA(H i (t))

[0073] Here, PCA(·) represents the principal component analysis dimensionality reduction process.

[0074] (3) The scalar inputs of all time steps are fed into the dynamic modal fusion module, and the weights of the modalities are dynamically adjusted using the attention mechanism. Based on the weight of each modality at each time step, the driver’s attention level score DAI at each time step is calculated.

[0075] From the deep feature representation of each modality, the score representing the driver’s attention level (DriverAttentionIndex, DAI) is calculated. The weighted fusion model is based on the basic weight ω of each modality. i To calculate the weighted attention score. The specific formula is:

[0076]

[0077] Among them, ω i (0) is the basic attention weight of the i-th modality at the initial time step, ω i '(t) is the dynamic attention weight of the i-th modality in the subsequent time step, h i (t) is the dimensionality reduction feature representation of the i-th mode at time step t.

[0078] To further enhance the dynamic adaptability of modal fusion, the system introduces a temporal attention gating mechanism and modal confidence assessment in the weight calculation process to achieve dynamic adjustment of weights based on historical status.

[0079] The time gating mechanism dynamically adjusts the importance of each mode at each time step based on the historical state in the time series to adapt to the volatility of the driver's state changes. The change in weight depends not only on the modal performance of the current time step, but also on the trend of the previous time steps for smooth update, which is expressed as:

[0080]

[0081] Among them, ω i (t) is the dynamic weight of the i-th mode at time step t; k i (t) is the eigenvector of the i-th mode at time step t; t is the time step, t = 0.5s; W and b are both trainable parameters.

[0082] The system introduces modal confidence to achieve real-time evaluation of the signal-to-noise ratio or stability of each modal sensor, and reduces its participation when the modal quality deteriorates. The system makes confidence corrections to the dynamic weights, and the fusion weights are updated as follows:

[0083]

[0084] Among them, ω i '(t) is the modified fusion weight; M is the total number of modalities; R i (t) is the confidence level of the ith mode at time step t, which is calculated based on the signal fluctuation amplitude:

[0085]

[0086] (4) Score the driver's attention level DAI in the past n time steps and assign different weights σ to the DAI of each time step i , calculate the mean of the driver's attention level scores, compare the mean with the threshold of the driver's attention level scores under each distraction level, and determine the current driver's distraction level.

[0087] In order to more accurately evaluate the driver's distracted state, a weighted sliding window method is adopted to give different weights to the DAI of the last n time steps, so as to achieve higher attention to the distracted state in the recent time steps.

[0088] First, a fixed time window size n is set, which means that the DAI data of the past n time steps are considered. Unlike the ordinary sliding window method, the DAI values ​​of the current time step and the past time step will be weighted according to the weight.

[0089] Then, at the current time step t, calculate the DAI values ​​of the past n time steps (including the current time step) and assign different weights σ to the DAI of each time step i The closer the time step is to the current one, the greater the weight is, and the farther the time step is from the current one, the smaller the weight is.

[0090]

[0091] in,

[0092] DAI(k) is the distraction state index at the kth time step;

[0093] σ i is the weight of the i-th time step;

[0094] n is the size of the sliding window;

[0095] DAI avg (t) is the average value of DAI within the weighted sliding window.

[0096] Weight σ i Calculated using the decaying exponential function, expressed as:

[0097] σ k =α n-k

[0098] Where: α is the decay factor, which controls the decay speed. Generally, a smaller value will make the weight decay faster. In this invention, α=5.

[0099] k is the index of the corresponding time step, with the current time step being t, t-1 being the previous time step, and so on.

[0100] Finally, according to the current DAI avg (t) Determine the current distraction state. The criteria are given in Table 1.

[0101] State DAI interval Focus <![CDATA[DAI avg >0.75]]> Mild distraction <![CDATA[0.5<DAI avg ≤0.75]]> Moderate distraction <![CDATA[0.25<DAI avg ≤0.5]]> Severe distraction <![CDATA[DAI avg ≤0.25]]>

[0102] (5) Intervene based on the driver’s current level of distraction to help them return to a state of focus.

[0103] Intervention methods include voice reminders, vibration feedback and instrument screen flashing. Different intervention methods of different degrees are used according to the different levels of distraction of the driver to help the driver return to a state of concentration.

[0104] Table 2 shows the reminder contents under different distraction states.

[0105]

[0106] The above is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A real-time driver distraction assessment and intervention method based on multimodal data, characterized in that: The following steps are involved: (1) Collecting data from three modalities: driver physiological data, eye movement data, and vehicle status, and converting the three modal data into a multimodal feature vector; (2) Input the multimodal feature vector into the Transformer encoding module, extract the deep features at each time step, and perform dimensionality reduction on the deep features of all time steps to become scalars; (3) The characteristic scalars of all n time steps in step (2) are input into the dynamic modal fusion module. The time gating mechanism and modal confidence evaluation are introduced into the dynamic modal fusion module. The weight of each modality at different time steps is dynamically adjusted based on the historical state, and then the driver's attention level score DAI at each time step is calculated; (4) Calculate the driver’s attention level score DAI for the past n time steps according to step (3), and assign different weights σ to the DAI of each time step. i , weight σ i As time approaches the current time step, the average value of the driver's DAI is calculated, and the average value is compared with the threshold value of the driver's attention level at each distraction level to determine the current driver's distraction level; (5) Take appropriate interventions based on the driver’s level of distraction.

2. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 1, characterized in that: The multimodal feature vector x(t): x(t)=[x phys (t),x eye (t),x motion (t)] Each subvector is: x phys (t)=[HR(t),GSR(t),RRI(t)] x eye (t)=[Pupil(t),GazeDev(t),BlinkDur(t)] x motion (t)=[LatAcc(t),LongAcc(t),Steering(t)] Where HR is heart rate, GSR is galvanic skin response, RRI is RR interval, Pupil is pupil diameter, GazeDev is gaze deviation, BlinkDur is blink duration, LatAcc is lateral acceleration, LongAcc is longitudinal acceleration, and Steering is steering wheel angle.

3. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 2, characterized in that: The characteristic scalar h in step (2) i The process of obtaining (t) is as follows: For each mode’s eigenvector x i (t) is input into the Transformer network and encoded through multiple layers of self-attention mechanism to obtain the deep feature H of each time step i (t), expressed as: H i (t)=Transformer(x i (t)) Among them, x i (t) is the feature of the i-th mode at time step t, i = 1, 2, 3, respectively refers to x phys (t), x eye (t), x motion (t), H i (t) is the depth feature of each time step; Finally, the features of all time steps are reduced to a scalar h using principal component analysis i (t), expressed as: h i (t)=PCA(H i (t)) Here, PCA(·) represents the principal component analysis dimensionality reduction process.

4. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 1, characterized in that: The time gating mechanism dynamically adjusts the importance of each mode at each time step based on the historical state in the time series to adapt to the volatility of the driver's state change. The change in weight depends not only on the modal performance of the current time step, but also on the trend of the previous time steps for smooth update, which can be expressed as: Among them, ω i (t) is the dynamic weight of the i-th mode at time step t; k i (t) is the feature vector of the i-th mode at time step t; k is the time step, and W and b are both trainable parameters.

5. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 1, characterized in that: The modal confidence factor is used to evaluate the signal-to-noise ratio or stability of each modal sensor in real time, and its participation is reduced when the modal quality decreases. The system makes confidence corrections to the dynamic weights, and the fusion weights are updated as follows: Among them, ω i '(t) is the modified fusion weight; M is the total number of modalities; R i (t) is the confidence level of the ith mode at time step t, which is calculated based on the signal fluctuation amplitude:

6. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 1, characterized in that: The calculation of the driver's attention level score DAI is as follows: Among them, ω i (0) is the basic attention weight of the i-th modality at the initial time step, ω i '(t) is the dynamic attention weight of the i-th modality in the subsequent time step, h i (t) is the dimension-reduced feature scalar of the i-th mode at time step t.

7. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 1, characterized in that: The mean of the driver's attention score at the current time step t is calculated as follows: At the current time step t, calculate the DAI value of the past n time steps and assign different weights σ to the DAI of each time step i ; DAI(k) is the distraction state index at the kth time step; σ k is the weight of the kth time step; n is the size of the sliding window; DAI avg (t) is the average value of DAI within the weighted sliding window; Weight σ k Calculated using the decaying exponential function, expressed as: s k =a n-k Among them, α is the decay factor that controls the decay speed; k is the index of the corresponding time step, the current time step is t, t-1 is the previous time step, and so on.

8. The method for real-time assessment and intervention of driver distraction based on multimodal data according to claim 1, characterized in that: The distraction levels include focus, mild distraction, moderate distraction and severe distraction. Different reminders are given according to the distraction status through voice reminders, vibration feedback and flashing instrument screen.