Driver state perception and cooperative driving regulation and control method based on multi-modal fusion

By using multimodal fusion driver state perception technology and a deep fusion model of EEG and eye movement signals, the driver state can be accurately decoded and personalized collaborative driving control can be achieved. This solves the problems of insufficient perception accuracy and lagging control strategies in existing technologies, and improves driving safety and comfort.

CN121650679APending Publication Date: 2026-03-13CHINA AUTOMOTIVE ENG RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing driver state perception technologies suffer from insufficient single-modal perception accuracy, high false alarm rates, and crude and lagging control response strategies. They lack deep integration and personalized collaborative driving control, resulting in low confidence in state recognition and poor robustness, making it impossible to achieve a balance between safety and comfort in human-machine collaboration.

Method used

A multimodal fusion driver state perception method is adopted. By simultaneously collecting EEG signals and eye movement signals, a deep fusion model based on attention mechanism is used to perform feature-level fusion to generate a fusion feature vector representing the driver's cognitive state. Combined with the current vehicle state and environmental information, personalized collaborative driving control strategies are dynamically selected, including vehicle assistance system control and human-machine interface adaptation.

Benefits of technology

It achieves accurate decoding and high-precision perception of driver status, provides intelligent and personalized forward-looking collaborative driving control, improves the robustness of driver status recognition and system acceptance, and ensures driving safety and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121650679A_ABST
    Figure CN121650679A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent driving, and discloses a multi-modal fusion-based driver state perception and cooperative driving regulation and control method, which comprises the following steps of S1, synchronously acquiring an electroencephalogram signal and an eye movement signal of a driver; s2, performing preprocessing and feature extraction to obtain an electroencephalogram feature sequence and an eye movement feature sequence; s3, performing feature-level fusion by using a multi-modal fusion model based on an attention mechanism to generate a fusion feature vector; s4, inputting the fusion feature vector into a state recognition model, and recognizing the cognitive state of the driver in real time; s5, based on the recognized cognitive state, the current vehicle state and the environment perception information, the matching degree of the driving task demand and the driver state is evaluated; and S6, according to the matching degree, dynamically selecting and executing corresponding regulation and control actions. According to the method, multi-modal physiological and behavior data can be deeply fused, a complex cognitive state is accurately decoded, and intelligent, personalized and prospective collaborative driving regulation and control are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, specifically to a driver state perception and collaborative driving control method based on multimodal fusion. Background Technology

[0002] Existing driver state perception technologies, as a key component in improving road traffic safety, have undergone years of development and formed several mainstream technical approaches, including steering wheel-based manipulation behavior analysis, vision-based facial expressions and eye tracking, and physiological signal monitoring based on wearable devices. Most of these technologies rely on single-modal data sources, and their theoretical basis lies in indirectly inferring a driver's internal state by monitoring changes in specific external behaviors or physiological indicators. For example, judging fatigue by the frequency of steering wheel micro-manipulation, or identifying distraction by the degree of gaze dispersion.

[0003] While these methods have demonstrated some effectiveness in controlled laboratory environments or simple scenarios, they have significant limitations in real-world, complex, dynamic driving environments. Single-modal perception accuracy is insufficient, and the false alarm rate is high. Single-sensor modalities are susceptible to various external interference factors. For example, vision-based methods heavily rely on lighting conditions, and drivers wearing glasses, sunglasses, or changes in makeup can significantly reduce recognition rates. Furthermore, normal driving observation behaviors (such as glancing at the dashboard or rearview mirror) are easily misinterpreted as inattention. On the other hand, methods based on physiological signals such as electroencephalography (EEG) can more directly reflect neural activity, but they are highly susceptible to head motion artifacts and complex electromagnetic noise in the non-static cockpit, making it difficult to guarantee the signal-to-noise ratio. Moreover, they lack sufficient specificity for cognitive states related to specific sensory channels, such as "visual distraction."

[0004] Furthermore, even when combined with multimodal approaches, without an effective fusion mechanism that is deep, adaptive, and temporally aligned, information is often simply spliced ​​together or voted at the decision level. This "information overlay" rather than "deep fusion" approach fails to fundamentally solve the problem of the inherent correlation between multi-source heterogeneous data. It cannot fully explore the deep complementary correlation and temporal dynamic coupling between EEG (reflecting neurocognitive activity) and eye movement (reflecting visual attention allocation), resulting in low confidence and poor robustness in state recognition.

[0005] Besides the accuracy issues in the perception stage, the control and response strategies of existing driver condition monitoring systems are also crude and lagging, making it difficult to achieve a balance between safety and comfort in human-machine collaboration. Most systems currently still use alarm mechanisms based on fixed thresholds, such as triggering audible or vibration warnings when fatigue indicators exceed a preset threshold. This is passive and reactive, typically intervening only after the driver's condition has significantly deteriorated, thus losing the safety margin for proactive prevention. Secondly, the control methods are singular and abrupt; continuous alarms may actually increase driver frustration and cognitive load, reducing system acceptance. Finally, the strategies lack stratification and adaptability, failing to provide refined and tiered responses based on the severity of the abnormal condition, the specific driving task context (such as highway cruising, urban congestion, unprotected left turns), and individual driver physiological baseline differences. The existing technology system has failed to form an efficient intelligent closed loop by accurately perceiving the state, understanding the real-time scene context, and making personalized control decisions. The fundamental reason is that there is a break in the chain from "perception" to "decision" to "control". There is a lack of a complete technical architecture that can comprehensively analyze the state of "driver-vehicle-environment" and generate flexible and collaborative control commands accordingly. Summary of the Invention

[0006] The present invention aims to provide a driver state perception and collaborative driving control method based on multimodal fusion, which can deeply integrate multimodal physiological and behavioral data, accurately decode complex cognitive states, and realize intelligent, personalized, and forward-looking collaborative driving control.

[0007] The basic solution provided by this invention is: a driver state perception and cooperative driving control method based on multimodal fusion, comprising the following steps: S1, simultaneously collects the driver's EEG and eye movement signals; S2, preprocess and extract features from the EEG and eye movement signals to obtain EEG feature sequences and eye movement feature sequences; S3, using a multimodal fusion model based on attention mechanism, feature-level fusion is performed on the EEG feature sequence and eye movement feature sequence to generate a fusion feature vector representing the driver's cognitive state; S4, the fused feature vector is input into the state recognition model to identify at least one cognitive state of the driver in real time, the cognitive state including fatigue level, distraction level and cognitive load intensity; S5 assesses the match between driving task requirements and driver state based on the identified cognitive state, current vehicle state, and environmental perception information. S6. Based on the matching degree, dynamically select and execute corresponding control actions from the preset hierarchical collaborative driving control strategy library. The control actions include adjusting the control parameters of the vehicle's assisted driving system, adapting the information presentation mode of the human-machine interface, and flexibly switching the ownership of driving control.

[0008] The working principle and advantages of this invention are as follows: This invention relates to a driver state perception and collaborative driving control method based on multimodal fusion. It can deeply integrate multimodal physiological and behavioral data, accurately decode complex cognitive states, and achieve intelligent, personalized, and forward-looking collaborative driving control. The key points are: First, this solution achieves deep fusion of multimodal signals, which helps improve the accuracy, robustness, and interpretability of driver state perception.

[0009] Unlike traditional methods that simply splice multimodal signals, this approach specifically targets heterogeneous data from EEG and eye movements, two types of data with inherent physiological connections, and designs a deep fusion model based on attention mechanisms. This model can adaptively learn and dynamically weight information from different modalities at the feature level. For example, when ambient light interferes with eye movement signals, the system can automatically increase the weight of cognitive load features derived from EEG; conversely, when EEG signals are contaminated with noise, it relies more on the visual attention allocation patterns reflected in eye movements. This dynamic, data-driven fusion approach not only achieves complementarity at the signal level but also delves deeper into the temporal coupling relationship between neurocognitive activity and visual behavior, thus enabling a more nuanced distinction between states such as "fatigue," "distraction," and "high cognitive load," which may appear similar but have different underlying mechanisms.

[0010] Second, this solution enables proactive and personalized cooperative driving.

[0011] This solution doesn't simply trigger a general alarm upon detecting an anomaly. Instead, it performs a comprehensive risk assessment and driving task requirement matching analysis based on real-time driver cognitive status, combined with current vehicle status and environmental perception information. Based on this, the solution's built-in hierarchical control strategy library can generate and execute a series of gradient and personalized cooperative driving actions. For example, for initial signs of inattention, it might only provide a gentle alert by highlighting key traffic elements on the head-up display; when moderate fatigue is detected, it might automatically enhance the lane-keeping assist system's correction sensitivity and appropriately limit the cruise speed to increase the safety margin; and in cases of severe cognitive overload or high-risk scenarios, it will initiate a smooth takeover process, stabilizing the vehicle to a safe state through coordinated adjustments of longitudinal and lateral controls, rather than abrupt emergency braking. This solution seamlessly integrates control interventions into the driving task flow, providing active safety protection when necessary while minimizing unnecessary interference to the driver, thus helping to ensure driving continuity and comfort. The entire control process fully considers the context of the driving scenario and the real-time physiological feedback of the individual driver, which can form a dynamic optimization capability, thereby significantly improving the trust and acceptance of human-machine co-driving. It is applicable to a variety of scenarios, from highways to complex urban traffic, and can meet the strict requirements of L2+ and L3 level intelligent assisted driving systems for high reliability and universality. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating an embodiment of the driver state perception and cooperative driving control method based on multimodal fusion of the present invention. Detailed Implementation

[0013] The following detailed explanation illustrates the specific implementation methods: The basic implementation examples are as follows: Figure 1 As shown: A driver state perception and cooperative driving control method based on multimodal fusion includes the following steps: S1 simultaneously collects the driver's electroencephalogram (EEG) and eye movement signals.

[0014] This step is achieved through a multi-source sensing system deployed within the vehicle. This system includes an EEG acquisition module (specifically, an existing dry electrode EEG acquisition device) and an eye-tracking module (specifically, a near-infrared camera eye tracker). In this embodiment, the driver wears a lightweight dry electrode EEG acquisition device, for example, an existing 8-channel (covering FP1, FP2, Fz, Cz, Pz, O1, O2, with the reference electrode placed on the earlobe) wireless EEG headband with a sampling rate set to 256 Hz to balance data accuracy and system load. Simultaneously, a near-infrared camera eye tracker is integrated behind the vehicle's dashboard or in the rearview mirror, with a sampling rate set to 60 Hz, for non-invasive tracking of the driver's eye features.

[0015] To achieve accurate multimodal fusion, this embodiment employs a hardware-software combination for time synchronization: During system initialization, the vehicle-mounted domain controller acts as the master clock source, synchronously distributing high-precision timestamps to the EEG acquisition module and eye-tracking module via a Precise Time Protocol (PTP) network. Both modules append a unified microsecond-level timestamp from this clock source to each data packet they acquire. Furthermore, at the data link layer, periodic synchronization signal frames are used for delay calibration to ensure alignment of the EEG and eye-tracking signals on the time axis from the signal source, with alignment errors controlled within ±10 milliseconds to meet the requirements of subsequent timing analysis.

[0016] Through the above processing, the consistency of multimodal data in the time dimension can be guaranteed, laying a reliable foundation for subsequent high-precision analysis. This enables the solution to maintain stable perception performance when facing individual differences, complex environmental changes and interference from normal driver operation, providing a crucial prerequisite for achieving reliable human-machine collaboration.

[0017] S2, preprocess and extract features from the EEG and eye movement signals to obtain EEG feature sequences and eye movement feature sequences.

[0018] The preprocessing and feature extraction include: The EEG signals were filtered, artifacts were removed, and time-domain, frequency-domain, and nonlinear dynamic features were extracted.

[0019] Specifically, in this embodiment, the acquired raw EEG signal is first passed through a bandpass filter from 0.5Hz to 45Hz to remove DC drift and high-frequency noise. Subsequently, independent component analysis is used to automatically identify and remove typical artifacts caused by eye movements, blinking, ECG, and muscle activity. The preprocessed clean EEG signal is segmented into continuous time windows of 2 seconds each with an overlap of 1.5 seconds for feature extraction.

[0020] For each channel of data within each time window, calculate the following characteristics: Band power spectral density (PSD): Calculated using the Welch method, this represents the relative power (i.e., the ratio of the power in that band to the total power) of the δ band (1-4Hz), θ band (4-8Hz), α band (8-13Hz), β band (13-30Hz), and γ band (30-45Hz). This is a core indicator of fatigue and cognitive load.

[0021] Nonlinear dynamic characteristics: The sample entropy of each channel signal is calculated to describe the complexity of the EEG signal.

[0022] The above features constitute the EEG feature sequence.

[0023] The eye movement signal is calibrated and filtered, and features such as fixation point, saccade, pupil diameter change rate, and saccade path entropy are extracted.

[0024] Specifically, in this embodiment, after the eye movement signal data is smoothed by Kalman filtering, the velocity-threshold method (identifying intervals with angular velocities greater than 30° / s as saccades) is used to segment the continuous fixation point sequence into "fixation" and "saccade" events. For the same time window aligned with the EEG signal, the following eye movement features are calculated: eye movement PERCLOS, average fixation duration, saccade frequency (times / second), average peak saccade velocity (° / second), mean pupil diameter and its standard deviation, fixation point spatial distribution entropy, and eye movement fixation entropy (EME), etc.

[0025] The above features constitute an eye movement feature sequence.

[0026] A unified timestamp mechanism was used to align the extracted EEG and eye movement features in time sequence.

[0027] Because EEG and eye-tracking sampling rates differ, to ensure that each fusion unit corresponds to the same time period, this embodiment uses a time window of the EEG signal as a benchmark (a new window is generated every 0.5 seconds). For each new EEG time window, all eye-tracking data points whose timestamps fall completely within this window are selected, and the aforementioned eye-tracking features are calculated. If there are insufficient eye-tracking data points within the window (e.g., due to brief occlusion), the features of the previous valid window are used for linear interpolation to supplement them, ensuring the continuity of the feature stream.

[0028] S3. Using a multimodal fusion model based on attention mechanism, feature-level fusion is performed on the EEG feature sequence and eye movement feature sequence to generate a fusion feature vector representing the driver's cognitive state.

[0029] When performing feature-level fusion using an attention-based multimodal fusion model, the following sub-steps are included: The EEG feature sequence and eye movement feature sequence are respectively input into the encoder to obtain their respective hidden layer representations; The attentional weights of EEG features on eye movement features and the attentional weights of eye movement features on EEG features were calculated using the cross-attention mechanism. Based on the attention weights, the hidden layer representations of both parties are weighted and fused to generate the fused feature vector containing cross-modal complementary information.

[0030] S4, the fused feature vector is input into the state recognition model to identify at least one cognitive state of the driver in real time, the cognitive state including fatigue level, distraction level and cognitive load intensity.

[0031] The state recognition model is a multi-task classification or regression model that can simultaneously output quantitative indicators of fatigue level, distraction level, and cognitive load intensity.

[0032] In this embodiment, the state recognition model is a fully connected neural network with a multi-task learning structure. The state recognition model takes the fused feature vector output by S3 as input, shares several hidden layers at the bottom, and then branches out into three independent sub-networks at the top level, each used for regressing and predicting three targets: Fatigue Level (FS): Outputs a continuous value between 0 and 1, where 0 represents complete wakefulness and 1 represents extreme fatigue. Training labels are derived from expert-evaluated video clips or use validated physiological indicators (such as EEG-based fatigue indices) collected synchronously in standard fatigue-inducing experiments (such as long-duration driving simulator tasks) as supervisory signals.

[0033] Distraction Level (DL): Outputs a continuous value between 0 and 1, where 0 represents high focus and 1 represents severe distraction. Training data is obtained by designing a dual-task paradigm (e.g., primary driving task + secondary cognitive task), and the distraction label is quantified by the performance of the secondary task (e.g., reaction time, accuracy).

[0034] Cognitive Load Intensity (CLI): Outputs a continuous value between 0 and 1. Training labels are typically obtained by controlling the difficulty of the driving task (such as traffic flow density and road complexity) and combining it with the driver's subjective load rating (such as the NASA-TLX scale).

[0035] The state recognition model uses mean squared error (MSE) as the loss function for each branch, and the total loss is the weighted sum of the losses of each branch. Through end-to-end training, the model can simultaneously and efficiently output quantification metrics (FS, DL, CLI) for three states.

[0036] S5 assesses the match between driving task requirements and driver state based on the identified cognitive state, current vehicle state, and environmental perception information.

[0037] The assessment of the match between driving task requirements and driver status specifically includes: Construct a comprehensive state space that includes driver state vector, vehicle dynamic parameter vector, and environmental scene vector; Among them, the driver state vector is [FS, DL, CLI], the vehicle dynamic parameter vector is [vehicle speed, lateral acceleration, steering wheel angle, and time distance to the vehicle in front (TTC)], and the environmental scene vector is the scene label, such as highway cruise, urban traffic jam following, unprotected left turn, night driving, etc., which can be provided by the vehicle's environmental perception system.

[0038] The similarity or distance between the integrated state space and multiple preset driving scenario prototypes is calculated using a pre-trained matching evaluation network, and the matching score and potential risk level are output.

[0039] Specifically, each scenario label presets a baseline cognitive load requirement and a maximum permissible distraction level. For example, unprotected left turns have a higher cognitive load requirement and a lower maximum permissible distraction level, while highway cruising has a lower cognitive load requirement.

[0040] Based on this, the deviation between the current driver state and the scenario requirements is calculated. , .

[0041] Taking into account the effect of fatigue, a preliminary fit is defined. ; , and These are empirical weighting coefficients.

[0042] By combining the vehicle dynamic parameter vectors, a weighted calculation is performed to obtain the vehicle state risk. , For example, when TTC is below the safety threshold or lateral acceleration is excessive, the vehicle's state risk increases. Increase. The final match degree is: ;in, This is a correction factor. 1 indicates a perfect match, and 0 indicates a very poor match.

[0043] Preferably, M and Output a comprehensive risk level (low, medium, high). For example, construct a 3x3 decision matrix, with the horizontal axis representing the risk level M (low, medium, high) and the vertical axis representing the risk level M. The risk levels are categorized as low, medium, and high. Each cell in the matrix defines the overall risk level to be output under these two input combinations. The decision-making logic follows the "weakest link" principle—high risk in any dimension is sufficient to raise the overall risk level.

[0044] S6. Based on the matching degree, dynamically select and execute corresponding control actions from the preset hierarchical collaborative driving control strategy library. The control actions include adjusting the control parameters of the vehicle's assisted driving system, adapting the information presentation mode of the human-machine interface, and flexibly switching the ownership of driving control.

[0045] The adaptation of the human-computer interaction interface information presentation method includes: When the driver's distraction level is detected to be increased, non-critical information on the central control screen is simplified, and potential risk targets ahead are highlighted on the head-up display; When the system detects that the driver's cognitive load is too high, it slows down the speech rate of the voice prompts, simplifies the content, and suspends non-urgent in-vehicle entertainment information pushes.

[0046] When dynamically selecting control actions based on matching degree, fuzzy control or rule reasoning engine is used. Its input is fuzzified cognitive state indicators and vehicle environment indicators, and its output is the execution intensity or confidence degree corresponding to different control actions.

[0047] Specifically, the hierarchical cooperative driving control strategy library includes at least three levels of strategies: Level 1 suggestion strategy: When the match rate slightly decreases (e.g., When this happens, adjust the information density of the human-computer interaction interface or enhance the prompts in specific sensory channels.

[0048] For example, if the DL is high, the information on the central control screen is simplified, and only the nearest vehicle and lane lines are highlighted on the HUD; if the CLI is high, the speech rate of all voice prompts is reduced by 30%, and the content is simplified.

[0049] If FS shows an upward trend, the lateral control gain of the Lane Keeping Assist (LKA) system will be gently increased by 10%-20% to make the correction more aggressive but less noticeable.

[0050] Secondary auxiliary enhancement strategy: When the matching degree decreases moderately (e.g., When adjusting lane keeping assist, adaptive cruise control sensitivity, or following distance, adjust the sensitivity of these functions. For example, the following distance (THW) of Adaptive Cruise Control (ACC) automatically increases by one level (e.g., from 1.5 seconds to 2.0 seconds). LKA's sensitivity is increased by 50%, and a slight speed limit is imposed (e.g., not exceeding 90% of the current road speed limit). It also combines auditory (three short warning sounds) and tactile (slight vibration of the steering wheel) warnings.

[0051] Three-tier takeover strategy: When the matching degree drops significantly (e.g., ...), When the risk level is too high (i.e., the overall risk level is high), a smooth takeover process for driving control is initiated, and the vehicle's maximum speed or driving trajectory is limited.

[0052] For example, explicitly announcing "takeover in progress" via voice and initiating a minimum risk strategy (MRM). In terms of longitudinal control, using comfortable deceleration (e.g.) The vehicle will decelerate smoothly; in terms of lateral control, combined with high-precision maps and perception, the vehicle will be kept stable in the center of the current lane or the nearest safe stopping lane.

[0053] Additionally, the driving mode is downgraded from Level 3 "Conditional Automated Driving" to Level 2 "Combined Driving Assistance," and the text "Please prepare to take over" is continuously flashed on the HUD until the vehicle enters a stable low-speed state or the driver takes over actively.

[0054] S7, Feedback on Regulation Effect and Model Update Steps: After executing the control action, continuously monitor the driver's subsequent physiological signal response, vehicle dynamic behavior response, and system takeover frequency; The effectiveness of the control actions is evaluated based on monitoring data (for example, if the driver's condition indicators are improving and there is no resistance to the operation, the control actions are considered effective), and the state recognition model or control strategy selection logic is adaptively optimized online or offline using feedback data.

[0055] To more specifically illustrate the workflow and effects of the method of the present invention in actual driving scenarios, two typical application cases are provided below.

[0056] Application Case 1: Cooperative driving control in unguided left turn scenarios.

[0057] This application case simulates a common "unguided left turn" scenario in urban roads, where there are no dedicated left-turn arrow lights or road markings at the intersection, and vehicles must find a gap in oncoming traffic to complete the left turn. This scenario places high demands on the driver's attention allocation, decision-making speed, and vehicle control precision. Specifically, it includes the following operations: (1) Test initialization: The tester starts the system, wears the EEG acquisition device and completes the calibration, and the vehicle-mounted eye tracker completes the calibration simultaneously (e.g., the eye tracking error is less than 0.5°). The driving task is configured as "unguided left turn" through the host computer system, and the driver's resting physiological baseline is recorded. The vehicle is placed in a simulated or real unguided left turn intersection scenario.

[0058] (2) Data Acquisition and Status Awareness: The vehicle approaches the intersection. The EEG device and eye tracker synchronously acquire data at a period of 20ms and send it to the cockpit-driver fusion domain controller via the CAN bus. The system performs preprocessing and feature extraction in real time to obtain EEG and eye movement features.

[0059] (3) Matching evaluation and strategy matching: EEG feature vectors (such as θ / β power of each channel) and eye movement feature vectors (fixation entropy, pupil changes) are input into a multimodal fusion model based on cross-attention mechanism to generate a fusion feature vector representing the driver's cognitive state. The fusion feature vector is then input into a state recognition model to identify the driver's cognitive state in real time.

[0060] The system identifies the current scene as "unguided left turn". It queries the built-in scene benchmark table to obtain the baseline cognitive load requirement and maximum allowable distraction value for this scene. It then calculates the matching degree and combines it with... The overall risk level is determined to be "medium".

[0061] (4) Coordinated Control Execution: Based on the above decisions, the system generates and executes coordinated control instructions: Control coordination: Under the premise of ensuring safety, the system makes fine adjustments to the longitudinal control. When it is determined that there is an acceptable gap, it applies a slight acceleration trend (such as the target speed gradually increasing from 5km / h to 10km / h) to assist the driver in completing the left turn.

[0062] Human-computer interaction: The system can use the HUD to highlight suggested acceleration times and turning paths with bright color blocks.

[0063] Continuous monitoring: Throughout the process, the system continuously monitors the driver's status. If a sudden increase in distraction level is detected (such as the driver's gaze leaving the intersection for more than 1 second), the acceleration trend will be immediately canceled, and the collision warning will be enhanced.

[0064] Optionally, the passing criteria that can be referenced for coordinated control are shown in Table 1 below, corresponding to different traffic scenarios. Table 1 Collaborative Control Table

[0065] Test Completion and Evaluation: After completing the left turn, the system saves all data. The test report can analyze whether the acceleration assistance provided by the system is smooth and timely during the unguided left turn, and whether the driver's state remains stable, thereby verifying the effectiveness of the cooperative strategy.

[0066] Application Case 2: Cooperative driving control in extreme U-turn scenarios.

[0067] This application case is used to verify the degradation and takeover mechanism under long-tail conditions such as extreme U-turns in confined spaces, when the driver is under high load or the vehicle is close to its physical limits.

[0068] (1) Test initialization: The tester starts the system, wears the EEG acquisition device and completes the calibration, and the vehicle-mounted eye tracker completes the calibration simultaneously (e.g., the gaze tracking error is less than 0.5°). The driving task is configured as "extreme U-turn scenario" through the host computer system, and the driver's resting physiological baseline is recorded. The vehicle is placed in a simulated or real extreme U-turn scenario.

[0069] (2) Data Acquisition and Status Awareness: The vehicle approaches the intersection. The EEG device and eye tracker synchronously acquire data at a period of 20ms and send it to the cockpit-driver fusion domain controller via the CAN bus. The system performs preprocessing and feature extraction in real time to obtain EEG and eye movement features.

[0070] (3) Matching evaluation and strategy matching: EEG feature vectors (such as θ / β power of each channel) and eye movement feature vectors (fixation entropy, pupil changes) are input into a multimodal fusion model based on cross-attention mechanism to generate a fusion feature vector representing the driver's cognitive state. The fusion feature vector is then input into a state recognition model to identify the driver's cognitive state in real time.

[0071] The system identifies the current scene as a "limited U-turn scenario". It queries the built-in scene benchmark table to obtain the baseline cognitive load requirement and maximum allowable distraction value for this scenario. It then calculates the matching degree and combines it with... The overall risk level is determined to be "high".

[0072] (4) Coordinated Control Execution: Based on the above decisions, the system generates and executes coordinated control instructions: Control demotion: The system issues a clear warning (such as a voice prompt "Demotion to driver assistance is imminent") and downgrades the driving mode from Level 3 Conditional Automated Driving to Level 2 Driver Assistance. Simultaneously, it immediately limits drive torque output and restricts the vehicle speed to 80% of the current speed to reduce dynamic risks.

[0073] Safety Monitoring and Takeover: After downgrading, the system still provides lateral and longitudinal assistance to the vehicle, but requires the driver to assume primary control responsibility. If the system detects that the driver has not effectively taken over (e.g., the steering wheel torque sensor does not detect clear driver steering input), or the vehicle trajectory still cannot meet the requirements for a safe U-turn (e.g., there is a possibility of scraping the curb), the system will further trigger emergency measures: control the vehicle to slow down and stop, and display a strong prompt such as "Please take over immediately" on the HUD.

[0074] Scenario Adaptation: If an oncoming vehicle is encountered during a U-turn, the system will prioritize braking based on the safety rule base and wait for the risk to be eliminated.

[0075] (5) Test completion and evaluation: After the left turn is completed, the system saves all data. The test report can analyze whether the acceleration assistance provided by the system is smooth and timely during the unguided left turn, and whether the driver's state remains stable, thereby verifying the effectiveness of the cooperative strategy.

[0076] This embodiment provides a driver state perception and collaborative driving control method based on multimodal fusion, which can deeply integrate multimodal physiological and behavioral data, accurately decode complex cognitive states, and achieve intelligent, personalized, and forward-looking collaborative driving control.

[0077] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.

Claims

1. A driver state perception and cooperative driving control method based on multimodal fusion, characterized in that, Includes the following steps: S1, simultaneously collects the driver's EEG and eye movement signals; S2, preprocess and extract features from the EEG and eye movement signals to obtain EEG feature sequences and eye movement feature sequences; S3, using a multimodal fusion model based on attention mechanism, feature-level fusion is performed on the EEG feature sequence and eye movement feature sequence to generate a fusion feature vector representing the driver's cognitive state; S4, the fused feature vector is input into the state recognition model to identify at least one cognitive state of the driver in real time, the cognitive state including fatigue level, distraction level and cognitive load intensity; S5 assesses the match between driving task requirements and driver state based on the identified cognitive state, current vehicle state, and environmental perception information. S6. Based on the matching degree, dynamically select and execute corresponding control actions from the preset hierarchical collaborative driving control strategy library. The control actions include adjusting the control parameters of the vehicle's assisted driving system, adapting the information presentation mode of the human-machine interface, and flexibly switching the ownership of driving control.

2. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, The preprocessing and feature extraction include: The EEG signals were filtered, artifacts were removed, and time-domain, frequency-domain, and nonlinear dynamic features were extracted. The eye movement signal is calibrated and filtered, and features such as fixation point, saccade, pupil diameter change rate, and saccade path entropy are extracted. A unified timestamp mechanism was used to align the extracted EEG and eye movement features temporally.

3. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, In S3, when performing feature-level fusion using an attention-based multimodal fusion model, the following sub-steps are included: The EEG feature sequence and eye movement feature sequence are respectively input into the encoder to obtain their respective hidden layer representations; The attentional weights of EEG features on eye movement features and the attentional weights of eye movement features on EEG features were calculated using the cross-attention mechanism. Based on the attention weights, the hidden layer representations of both parties are weighted and fused to generate the fused feature vector containing cross-modal complementary information.

4. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, The state recognition model is a multi-task classification or regression model that can simultaneously output quantitative indicators of fatigue level, distraction level, and cognitive load intensity.

5. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, The assessment of the match between driving task requirements and driver state specifically includes: Construct a comprehensive state space that includes driver state vector, vehicle dynamic parameter vector, and environmental scene vector; The similarity or distance between the integrated state space and multiple preset driving scenario prototypes is calculated using a pre-trained matching evaluation network, and the matching score and potential risk level are output.

6. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, The hierarchical cooperative driving control strategy library includes at least three levels of strategies: Level 1 prompting strategy: When the matching degree slightly decreases, adjust the information density of the human-computer interaction interface or enhance the prompts of specific sensory channels; Level 2 Assist Enhancement Strategy: When the matching degree decreases moderately, adjust the sensitivity of lane keeping assist and adaptive cruise control or the following distance; Three-level takeover strategy: When the matching degree drops significantly or the risk level is too high, a smooth driving control takeover process is initiated, and the vehicle's maximum speed or driving trajectory is limited.

7. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, The adaptation of the human-computer interaction interface information presentation method includes: When the driver's distraction level is detected to be increased, non-critical information on the central control screen is simplified, and potential risk targets ahead are highlighted on the head-up display; When the system detects that the driver's cognitive load is too high, it slows down the speech rate of the voice prompts, simplifies the content, and suspends non-urgent in-vehicle entertainment information pushes.

8. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, It also includes steps for feedback on the regulation effect and model updates: After executing the control action, continuously monitor the driver's subsequent physiological signal response, vehicle dynamic behavior response, and system takeover frequency; The effectiveness of the control actions is evaluated based on monitoring data, and the state recognition model or control strategy selection logic is adaptively optimized online or offline using feedback data.

9. The driver state perception and cooperative driving control method based on multimodal fusion according to claim 1, characterized in that, In S6, when dynamically selecting control actions based on matching degree, fuzzy control or rule reasoning engine is used. Its input is fuzzified cognitive state indicators and vehicle environment indicators, and its output is the execution intensity or confidence level corresponding to different control actions.