Driver dangerous behavior intervention system and method based on spatiotemporal graph neural network
Through multimodal data collection and spatiotemporal graph neural network processing, a dynamic graph structure of driver-vehicle interaction is constructed, which solves the problems of insufficient multimodal fusion and insufficient physiological status monitoring in existing technologies, realizes early prediction and intelligent intervention of driver's dangerous behavior, and significantly improves recognition accuracy and response speed.
Patent Information
- Application Number
- CN202510381610.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing driver dangerous behavior monitoring systems lack multimodal information fusion capabilities, cannot monitor physiological states, lack active intervention capabilities, have limited generalization capabilities, cannot adapt to different vehicle models and cab layouts, and cannot capture micro-expressions, resulting in insufficient recognition robustness and delayed prediction time.
A multimodal data acquisition module, including millimeter-wave radar, infrared structured light sensor and steering wheel pressure sensor grid, is used to construct a dynamic graph structure of driver-vehicle interaction. Spatiotemporal graph neural network is used for multi-scale temporal convolution and recursive causal attention processing, and closed-loop optimization module is combined for personalized intervention.
It has achieved early prediction and intelligent intervention of dangerous driver behaviors, with recognition accuracy increased to 98.6%, false alarm rate reduced to 0.8%, emergency response time reduced to 0.7 seconds, and accident rate reduced by 41%.
Smart Images

Figure CN120288054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving, and specifically to a system and method for intervening in driver dangerous behaviors based on a spatiotemporal graph neural network, and more particularly to a system and method for early prediction and intelligent intervention of driver dangerous behaviors using multimodal sensing technology and a spatiotemporal graph neural network. Background Art
[0002] With the development of intelligent driving technology, driver behavior monitoring and safety intervention systems are becoming an integral part of automotive safety equipment. Prior art, such as Chinese patent CN 106650644 B, discloses a method and system for identifying dangerous driver behavior. This system uses ultrasonic technology to identify dangerous driver behavior. First, an ultrasonic field is set up inside the vehicle to identify distinct patterns of dangerous behavior on the Doppler spectrum. Principal component analysis is then used to extract key features. A support vector machine algorithm is then used to generate multiple classifiers for identifying different dangerous behaviors. A gradient model forest is then constructed based on the degree of completion and duration of the dangerous behavior. Finally, the real-time ultrasonic signal is sliced using a window-cutting algorithm. The forest gradient model then identifies the dangerous behavior and issues an alarm.
[0003] However, the above technology has the following shortcomings: First, the system relies only on a single ultrasonic signal source and lacks the ability to fuse multimodal information, resulting in insufficient recognition robustness in complex environments; second, the system is unable to monitor the driver's physiological states such as fatigue and emotions, and only focuses on external behavioral characteristics, missing the early warning opportunity for dangerous behavior; third, the system uses a simple alarm mechanism and lacks active intervention capabilities, and cannot provide personalized intervention strategies based on the individual differences of different drivers; finally, the system has difficulty adapting to different vehicle models and cab layouts, has limited generalization capabilities, and is unable to capture subtle changes such as the driver's facial micro-expressions, resulting in a time delay in the prediction of dangerous behavior. Summary of the Invention
[0004] The purpose of the present invention is to provide a driver dangerous behavior intervention system and method based on spatiotemporal graph neural network, aiming to address the shortcomings of the existing technology.
[0005] The present invention proposes a driver dangerous behavior intervention system based on spatiotemporal graph neural network, including:
[0006] Multimodal data acquisition module, used to collect driver physiological signals, driving posture and control behavior data;
[0007] The spatiotemporal graph neural network risk perception engine is connected to the multimodal data acquisition module and is used to:
[0008] Receiving multimodal data sent by the multimodal data acquisition module;
[0009] constructing a driver-vehicle interaction dynamic graph structure based on the multimodal data;
[0010] Using a multi-scale temporal convolutional network to process the temporal changes of the dynamic graph structure;
[0011] Output the driver risk status assessment results;
[0012] An intelligent intervention execution module, connected to the spatiotemporal graph neural network risk perception engine, is used to:
[0013] receiving the risk status assessment result;
[0014] selecting a tactile, visual, or vehicle control intervention strategy based on the risk status assessment result;
[0015] implement the chosen intervention strategy;
[0016] A closed-loop optimization module, connected to the intelligent intervention execution module and the spatiotemporal graph neural network risk perception engine, is used to:
[0017] Monitor changes in driver behavior after intervention;
[0018] Evaluate intervention effectiveness;
[0019] Based on the intervention effect, the driver risk model and intervention strategy are updated.
[0020] Preferably, the multimodal data acquisition module includes:
[0021] Millimeter-wave radar sensor, used to collect the driver's heart rate and breathing rate;
[0022] Infrared structured light sensor, used to build a 3D skeletal posture model of the driver;
[0023] A steering wheel pressure sensing grid to capture changes in grip force distribution;
[0024] A multi-level buffer for storing data collected by the sensor;
[0025] The timestamp calibration unit is used to synchronize the time of multi-source data.
[0026] Preferably, the spatiotemporal graph neural network risk perception engine includes:
[0027] A graph structure building module is used to uniformly model the driver's physiological state and behavioral characteristics into a dynamic graph structure;
[0028] Spatial graph convolution module, used to process the spatial relationship of nodes in the graph structure;
[0029] Multi-scale temporal convolution module, used to process temporal dimension changes by cascading temporal convolutional networks with dilation factors;
[0030] Recursive causal attention module for extracting long sequence dependencies;
[0031] The risk assessment module is used to output the driver's risk status assessment results.
[0032] Preferably, the dynamic graph structure includes:
[0033] Skeleton key point nodes, used to represent the position information of the driver's head, torso and limbs;
[0034] Physiological status nodes are used to represent physiological indicators such as heart rate, heart rate variability, and respiratory rate;
[0035] Vehicle control nodes, used to represent control interfaces such as steering wheels, pedals, and gear shifts;
[0036] Spatial correlation edges, used to connect related nodes;
[0037] Time correlation edges, used to capture the temporal change patterns of nodes;
[0038] Physiological-behavioral interaction edges are used to establish mapping relationships between physiological state nodes and behavioral nodes.
[0039] Preferably, the multi-scale temporal convolution module comprises:
[0040] The short-timescale branch is used to detect sudden dangerous actions;
[0041] The mesoscale branch is used to monitor recent trends in driving behavior;
[0042] a long-timescale branch for monitoring cumulative risk factors such as fatigue;
[0043] The time scale adaptive fusion unit is used to dynamically adjust the importance of different time scales according to the current driving scenario.
[0044] Preferably, the intelligent intervention execution module includes:
[0045] A tactile feedback channel for providing directional warnings through localized vibrations of the steering wheel;
[0046] A visual enhancement feedback channel for annotating risk sources via a heads-up display system;
[0047] Vehicle control adjustment channel, used to automatically adjust vehicle dynamic parameters when the risk level reaches a threshold;
[0048] The intervention strategy selection unit is used to select the best intervention method based on the risk level and driver characteristics.
[0049] Preferably, the closed-loop optimization module includes:
[0050] Intervention effect monitoring unit, used to track changes in driver behavior and physiological indicators after intervention;
[0051] a personalized model updating unit for optimizing the driver's personal model based on intervention response data;
[0052] A double-delayed policy gradient optimizer is used to evaluate the short-term and long-term intervention effects and update the intervention policy.
[0053] Preferably, the dual-delay policy gradient optimizer includes:
[0054] Dual-Q network, used to evaluate short-term intervention effects and long-term safety impacts respectively;
[0055] Priority experience replay buffer pool, used to store intervention interaction data;
[0056] Behavior cloning pre-training module to accelerate model convergence;
[0057] The safety-constrained optimization objective function is used to ensure that interventions do not introduce new risks.
[0058] As an option, it also includes:
[0059] Driver profiling module, used to analyze different driving habits through unsupervised learning clustering;
[0060] Personalized threshold adjustment module, used to optimize individual risk thresholds based on reinforcement learning;
[0061] Environmental perception module, used to perceive current road and traffic conditions;
[0062] Among them, the spatiotemporal graph neural network risk perception engine dynamically adjusts risk assessment parameters according to the driver portrait, personalized thresholds and environmental conditions.
[0063] The intelligent intervention method for driver dangerous behavior based on the system includes:
[0064] Collect multimodal data of driver's physiological signals, driving posture and control behavior;
[0065] constructing a driver-vehicle interaction dynamic graph structure based on the multimodal data;
[0066] Processing the dynamic graph structure using a spatiotemporal graph convolutional neural network and a multi-scale temporal convolutional network to output a driver risk status assessment result;
[0067] selecting and executing a tactile, visual, or vehicle control intervention strategy based on the risk status assessment result;
[0068] Monitor changes in driver behavior after intervention and evaluate intervention effectiveness;
[0069] Based on the intervention effect, the driver risk model and intervention strategy are updated through a double-delayed policy gradient algorithm.
[0070] The purpose of the present invention is to provide a driver dangerous behavior intervention system and method based on spatiotemporal graph neural network, aiming to solve the problems of single perception source, lack of physiological status monitoring, passive warning mechanism and insufficient adaptability in the existing technology.
[0071] This invention uses multimodal sensor fusion and a spatiotemporal graph neural network risk perception engine to achieve early prediction and intelligent intervention of driver risk behaviors, with the following beneficial effects:
[0072] 1) By integrating millimeter-wave radar vital sign monitoring with multi-channel visual analysis, a comprehensive driver status profile is constructed, increasing recognition accuracy from 95% to 98.6%;
[0073] 2) Based on physiological signal and micro-expression analysis, dangerous events can be predicted 15 to 30 seconds in advance, providing drivers with sufficient reaction time;
[0074] 3) Through an adaptive learning system and closed-loop optimization mechanism, it continuously adapts to individual driver differences, reducing the false alarm rate from 3.5% to 0.8%;
[0075] 4) Using a spatiotemporal graph neural network combined with a multi-scale time perception mechanism to simultaneously capture instantaneous behavior and long-term risk accumulation, the accident rate was reduced by 41%;
[0076] 5) Through personalized intervention strategies, emergency response time is reduced from 1.2 seconds to 0.7 seconds, significantly improving driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 This is a general architecture diagram of a driver dangerous behavior intervention system based on a spatiotemporal graph neural network in an embodiment of the present invention;
[0078] Figure 2 is a structural diagram of a multimodal data acquisition module in an embodiment of the present invention;
[0079] Figure 3 This is a schematic diagram of the structure of the spatiotemporal graph neural network risk perception engine in an embodiment of the present invention;
[0080] Figure 4 This is a schematic diagram of constructing a dynamic graph structure according to an embodiment of the present invention;
[0081] Figure 5 is a schematic structural diagram of a multi-scale temporal convolution module in an embodiment of the present invention;
[0082] Figure 6 is a schematic structural diagram of an intelligent intervention execution module in an embodiment of the present invention;
[0083] Figure 7 is a schematic structural diagram of a closed-loop optimization module in an embodiment of the present invention;
[0084] Figure 8 1 is a schematic diagram of the structure of a dual-delay policy gradient optimizer in an embodiment of the present invention;
[0085] Figure 9 This is a flow chart of a method for intelligently intervening in driver's dangerous behavior according to an embodiment of the present invention;
[0086] Figure 10 3 is a comparison chart of the recognition effects of four typical dangerous driving behaviors in an embodiment of the present invention. DETAILED DESCRIPTION
[0087] Please refer to the attached Figure 1-10 The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0088] Example 1: Overall Architecture of a Driver Dangerous Behavior Intervention System Based on Spatiotemporal Graph Neural Network
[0089] like Figure 1 As shown, the driver dangerous behavior intervention system based on spatiotemporal graph neural network provided by the present invention includes a multimodal data acquisition module 1, a spatiotemporal graph neural network risk perception engine 2, an intelligent intervention execution module 3 and a closed-loop optimization module 4.
[0090] The multimodal data acquisition module 1 is used to collect data on the driver's physiological signals, driving posture, and control behavior. This module contains multiple sensors that can comprehensively capture driver status information and provide multi-dimensional data support for subsequent analysis.
[0091] The spatiotemporal graph neural network risk perception engine 2 is connected to the multimodal data acquisition module 1. It receives the multimodal data sent by the multimodal data acquisition module 1, constructs a dynamic graph structure of the driver-vehicle interaction based on this multimodal data, and uses a multi-scale temporal convolutional network to process the temporal changes in this dynamic graph structure, ultimately outputting the driver's risk status assessment results. This engine is the core innovation of this system, using graph neural network technology to achieve accurate modeling of driver status and risk prediction.
[0092] The intelligent intervention execution module 3 is connected to the spatiotemporal graph neural network risk perception engine 2 and is used to receive risk status assessment results, select a tactile, visual, or vehicle control intervention strategy based on the results, and execute the selected intervention strategy. This module provides a gradient intervention response based on the risk level and driver characteristics, ranging from subtle prompts to active control.
[0093] Closed-loop optimization module 4 is connected to intelligent intervention execution module 3 and spatiotemporal graph neural network risk perception engine 2. It monitors changes in driver behavior after intervention, evaluates intervention effectiveness, and updates the driver risk model and intervention strategy based on the results. This module enables system self-optimization and personalized adaptation, continuously improving recognition accuracy and intervention effectiveness.
[0094] The system's workflow is as follows: First, the multimodal data acquisition module 1 collects the driver's physiological signals, driving posture, and control behavior data in real time. Second, the spatiotemporal graph neural network risk perception engine 2 receives and processes this data, constructs a dynamic graph structure, performs spatiotemporal analysis, and outputs a risk assessment result. Then, the intelligent intervention execution module 3 selects and executes an appropriate intervention strategy based on the risk assessment results. Finally, the closed-loop optimization module 4 monitors the intervention effects and continuously optimizes the system model and strategy. This entire process forms a closed-loop feedback mechanism, enabling continuous evolution and optimization of the system.
[0095] Example 2: Multimodal data acquisition module
[0096] like Figure 2 As shown, the multimodal data acquisition module 1 includes a millimeter wave radar sensor 11 , an infrared structured light sensor 12 , a steering wheel pressure sensing grid 13 , a multi-level buffer 14 and a timestamp calibration unit 15 .
[0097] Millimeter-wave radar sensor 11 operates at a 77GHz frequency, using the Doppler effect to capture the driver's minute body movements, enabling contactless heart and respiratory rate monitoring. Positioned in the instrument panel area in front of the cockpit, this sensor has a sampling rate of 50Hz and can penetrate clothing for stable physiological signal acquisition.
[0098] The infrared structured light sensor 12 projects a structured light spot and receives reflected light signals, constructing a 3D skeletal model of the driver's posture with an accuracy of ±1.2cm. This sensor uses a 30Hz sampling frequency and operates stably in a variety of lighting conditions, effectively addressing the light sensitivity of traditional vision systems.
[0099] The steering wheel pressure sensing grid 13, comprised of 16 x 16 flexible pressure sensing elements with a sampling rate of 200Hz, is used to capture changes in the driver's grip force distribution. This grid, covering the entire circumference of the steering wheel, can detect subtle changes in grip strength, position, and grip pattern, providing crucial data support for driving behavior analysis.
[0100] The multi-level buffer 14 utilizes a three-tiered buffer pool structure, comprising a raw data buffer, a pre-processed data buffer, and a feature data buffer, to store data collected by each sensor. The buffer capacities are 10 seconds, 30 seconds, and 5 minutes, respectively, enabling efficient management and access to short-term, medium-term, and long-term data.
[0101] The timestamp calibration unit 15 uses an adaptive timestamp calibration algorithm to synchronize multi-source data, resolving sampling frequency discrepancies and latency issues between sensors. This unit achieves millisecond-level time alignment, ensuring temporal consistency across multimodal data and paving the way for subsequent fusion analysis.
[0102] Preferably, the multimodal data acquisition module 1 also includes a signal enhancement and noise suppression unit (not shown in the figure), which uses adaptive bandpass filters, wavelet threshold denoising and environmental perception noise suppression technology to dynamically adjust the filtering parameters according to the vehicle vibration characteristics, thereby effectively improving the signal quality.
[0103] The data processing flow of the multimodal data acquisition module 1 is as follows: first, each sensor collects raw data according to its own sampling frequency; then, the raw data is stored in the raw data buffer after signal enhancement and noise suppression processing; then, the timestamp calibration unit 15 synchronizes the data and stores the synchronized data in the preprocessing data buffer; finally, the preprocessed data is stored in the feature data buffer after preliminary feature extraction and transmitted to the spatiotemporal graph neural network risk perception engine 2 for further analysis.
[0104] Example 3: Spatiotemporal Graph Neural Network Risk Perception Engine
[0105] like Figure 3 As shown, the spatiotemporal graph neural network risk perception engine 2 includes a graph structure construction module 21, a spatial graph convolution module 22, a multi-scale temporal convolution module 23, a recursive causal attention module 24 and a risk assessment module 25.
[0106] The graph structure construction module 21 is used to uniformly model the driver's physiological state and behavioral characteristics into a dynamic graph structure. This module receives processed feature data from the multimodal data acquisition module 1 and constructs a dynamic graph structure representing the driver-vehicle interaction state based on predefined node types and connection relationships.
[0107] The spatial graph convolution module 22 uses a graph convolutional network (GCN) structure to process the spatial relationships between nodes in the graph structure. This module effectively aggregates spatial information and can capture the mutual influence and correlation patterns between different nodes. The spatial graph convolution operation can be expressed as:
[0108]
[0109] Among them, H (l) is the node feature matrix of the lth layer, To add the adjacency matrix after self-connection, for The degree matrix, W (l) is the learnable weight matrix and σ is the activation function.
[0110] The multi-scale temporal convolution module 23 processes temporal variations by cascading a temporal convolutional network with dilation factors. This module can simultaneously capture short-term micro-behavioral changes and long-term macro-trends, enabling multi-scale temporal feature extraction.
[0111] The recursive causal attention module 24 is based on a bidirectional GRU structure and is used to extract long-term dependencies. This module introduces a causal attention mechanism that enables the current state to focus on relevant information in the historical state, enhancing the model's perception of time series.
[0112] Risk Assessment Module 25 utilizes a multi-layer perceptron architecture to integrate the features output by the aforementioned modules and generate a risk assessment. This module output includes three components: risk level (L1-L3), risk type identification, and risk development prediction, providing a comprehensive basis for subsequent intervention decisions.
[0113] The core advantage of the Spatiotemporal Graph Neural Network Risk Perception Engine 2 lies in its deep integration of physiological signals and behavioral characteristics, capturing complex spatiotemporal dependencies through a spatiotemporal graph convolutional network. Compared to traditional methods, this engine can predict potential risks earlier and more accurately, and provide fine-grained risk assessment results.
[0114] Example 4: Dynamic Graph Structure
[0115] like Figure 4 As shown, the dynamic graph structure includes skeleton key point nodes 41, physiological state nodes 42, vehicle control nodes 43, spatial correlation edges 44, temporal correlation edges 45 and physiological-behavioral interaction edges 46.
[0116] The skeleton keypoint nodes 41 include 23 keypoints, representing the positional information of the driver's head, torso, and limbs. These keypoints include the top of the head, nose tip, shoulders, elbows, wrists, chest, waist, hips, knees, and ankles, which together form the driver's posture skeleton model.
[0117] Physiological status node 42 includes seven nodes for representing physiological indicators such as heart rate, heart rate variability, and respiratory rate. These nodes include average heart rate, heart rate variability, respiratory rate, respiratory depth, stress index, fatigue index, and attention index, which comprehensively reflect the driver's physiological status.
[0118] The vehicle control node 43 includes 12 nodes, representing control interfaces such as the steering wheel, pedals, and gear shift. These nodes include the left, right, top, and bottom sections of the steering wheel; the accelerator pedal, brake pedal, clutch pedal, gear shift lever, and various function button areas, reflecting the interaction between the driver and the vehicle.
[0119] Spatial correlation edges 44 are predefined based on anatomical and behavioral knowledge and are used to connect related nodes. For example, there is a spatial correlation edge between the hands node and the steering wheel node, and between the head node and the gaze direction node, establishing spatial associations between nodes.
[0120] Temporal correlation edges 45 are used to capture the temporal change patterns of nodes. Their weights are calculated using the sliding window correlation coefficient. Temporal correlation edges connect the states of the same node at different time points, reflecting the temporal evolution and change trends of the state.
[0121] Physiological-behavioral interaction edges 46 are used to establish mappings between physiological state nodes and behavioral nodes. For example, there is an interaction edge between the heart rate variability node and the grip pattern node, and between the breathing rate node and the upper body posture node, reflecting the mutual influence between physiological state and behavioral performance.
[0122] The dynamic graph structure is maintained through a dynamic graph update mechanism, including attention-guided edge pruning, multi-time window graph sequence maintenance, and abnormal pattern enhancement. This structure can simultaneously represent the driver's physiological state, behavioral characteristics, and interaction patterns, providing a unified data representation for spatiotemporal graph convolutional networks.
[0123] Example 5: Multi-scale Temporal Convolution Module
[0124] like Figure 5 As shown, the multi-scale time convolution module 23 includes a short time scale branch 231 , a medium time scale branch 232 , a long time scale branch 233 and a time scale adaptive fusion unit 234 .
[0125] The short-timescale branch 231 uses temporal convolution layers with dilation factors of 1 and 2, with a receptive field covering 0.1 to 3 seconds. It is primarily used to detect sudden dangerous movements, such as sudden head drops and sudden turns. This branch uses 64 convolution kernels, each with a temporal length of 7, to capture transient behavioral characteristics.
[0126] The medium-scale branch 232 uses temporal convolution layers with dilation factors of 4 and 8, with a receptive field covering 3 to 30 seconds. This branch is used to monitor recent trends in driving behavior, such as gradual distraction and slow posture shifts. This branch uses 48 convolution kernels, each with a temporal length of 5, to capture medium-term behavioral changes.
[0127] The long-time-scale branch 233 uses a temporal convolution layer with a dilation factor of 16, with a receptive field covering 30 seconds to 5 minutes. This is used to monitor cumulative risk factors such as fatigue. This branch employs 32 convolution kernels, each with a time duration of 3, enabling it to capture long-term state trends.
[0128] The time-scale adaptive fusion unit 234 uses a gating mechanism to dynamically adjust the importance of different time scales according to the current driving scenario. This unit calculates the weight coefficients of each time-scale feature to achieve adaptive fusion of features. The fusion process can be expressed as:
[0129]
[0130] Among them, F i represents the characteristics of the i-th time scale branch, α i is the corresponding weight coefficient, which is calculated by the attention network. fused is the feature representation after fusion.
[0131] Each timescale branch of the multi-scale temporal convolution module 23 uses a residual connection structure, effectively alleviating the vanishing gradient problem in deep networks. Furthermore, through its multi-scale design, the module can simultaneously focus on short-term microscopic behaviors and long-term macroscopic trends, capturing comprehensive temporal information.
[0132] Preferably, the multi-scale temporal convolution module 23 also includes a temporal feature enhancement unit (not shown in the figure), which adopts a self-attention mechanism to enhance the expressiveness of temporal features and retains the original temporal information through a jump connection structure to further improve the model performance.
[0133] Example 6: Intelligent Intervention Execution Module
[0134] like Figure 6 As shown, the intelligent intervention execution module 3 includes a tactile feedback channel 31 , a visual enhancement feedback channel 32 , a vehicle control adjustment channel 33 and an intervention strategy selection unit 34 .
[0135] Haptic feedback channel 31 provides directional warnings through localized vibrations of the steering wheel. This channel utilizes 16 independently controlled linear resonators distributed around the circumference of the steering wheel, generating different patterns of tactile feedback. For example, if the driver frequently looks down, a gentle vibration alerts the upper portion of the steering wheel. If the driver uses one hand for too long, a gradually increasing vibration on the corresponding side of the steering wheel encourages two-handed control.
[0136] The visual enhancement feedback channel 32 identifies risk sources through a heads-up display system. This channel projects risk assessment results onto the vehicle's windshield, visually presenting risk information using augmented reality technology. The display includes a risk level indicator, the location of the risk source, and recommended actions, helping drivers quickly identify and respond to potential risks.
[0137] Vehicle control adjustment channel 33 automatically adjusts vehicle dynamic parameters when the risk level reaches a threshold. This channel communicates with the vehicle control system via the vehicle's CAN bus and can adjust adaptive cruise control system parameters, limit maximum speed, increase lane departure warning sensitivity, and other measures, enabling proactive vehicle-level intervention.
[0138] The intervention strategy selection unit 34 selects the optimal intervention method based on the risk level and driver characteristics. This unit implements the decision-making logic for the intervention strategy, dynamically selecting a single or combined intervention method based on the severity of the risk assessment results, the risk type, and the driver's personal characteristics, and determining the intervention intensity and duration.
[0139] Intelligent Intervention Execution Module 3 employs a gradient intervention strategy with three levels of intervention: Level 1 (mild reminder) primarily through tactile feedback; Level 2 (clear warning) combines tactile and visual feedback; and Level 3 (emergency intervention) initiates vehicle control adjustments. This gradient design avoids driver annoyance caused by excessive intervention while ensuring timely and effective intervention in high-risk situations.
[0140] Preferably, the intelligent intervention execution module 3 further includes an intervention information recording unit (not shown in the figure), which records information such as the time, type, intensity and duration of each intervention to provide data support for subsequent analysis and optimization.
[0141] Example 7: Closed-loop optimization module
[0142] like Figure 7 As shown, the closed-loop optimization module 4 includes an intervention effect monitoring unit 41, a personalized model updating unit 42 and a dual-delay policy gradient optimizer 43.
[0143] Intervention effect monitoring unit 41 is used to track changes in driver behavior and physiological indicators after the intervention. This unit receives real-time data from multimodal data acquisition module 1 and compares and analyzes it with pre-intervention data to assess the intervention response. Specific monitoring indicators include response time (how quickly the driver reacts to the intervention), magnitude of behavioral change (the degree of adjustment in driving behavior), and changes in physiological indicators (such as trends in heart rate and respiratory rate).
[0144] The personalized model update unit 42 optimizes the driver's individual model based on intervention response data. This unit maintains a personalized profile for each driver, including characteristics such as risk preferences, intervention response patterns, and behavioral habits. Based on intervention effectiveness monitoring results, this unit continuously updates the individual model parameters, ensuring the system's adaptability to individual differences. The personalized model utilizes an incremental learning approach, integrating new observations while retaining existing knowledge.
[0145] The dual-delay policy gradient optimizer 43 is used to evaluate the short-term and long-term effectiveness of interventions and update the intervention policy. This optimizer uses a reinforcement learning framework to model the intervention process as a Markov decision process, continuously optimizing the intervention policy through interaction with the environment. The dual-delay design effectively mitigates the problem of Q-value overestimation and improves the stability and efficiency of policy optimization.
[0146] Closed-Loop Optimization Module 4 enables the system's self-evolutionary capabilities. Through continuous monitoring, evaluation, and optimization, the system adapts to individual drivers and maintains efficient performance in ever-changing driving environments. This module's closed-loop feedback mechanism is key to continuously improving the system's accuracy and user experience.
[0147] Preferably, the closed-loop optimization module 4 also includes a scenario knowledge base (not shown) for storing optimal intervention strategies and effect evaluation results for different driving scenarios. This knowledge base adopts a hierarchical structure, including a general knowledge layer, a scenario-specific layer, and a user-specific layer, to support efficient knowledge organization and rapid retrieval.
[0148] Example 8: Dual-delayed policy gradient optimizer
[0149] like Figure 8 As shown, the dual-delay policy gradient optimizer 43 includes a dual-Q network 431, a priority experience replay buffer pool 432, a behavior cloning pre-training module 433 and a safety constraint optimization objective function 434.
[0150] Dual Q network 431 includes an online Q network and a target Q network, which are used to evaluate short-term intervention effectiveness and long-term safety impact, respectively. The online Q network is responsible for evaluating and updating the current policy, while the target Q network parameters are updated with a lag, providing a stable learning objective. This dual network design effectively alleviates the problem of Q-value overestimation and improves the stability of policy optimization. The Q-value update formula can be expressed as:
[0151]
[0152] Among them, s t and a t Represents state and action respectively, r t is the reward, α is the learning rate, γ is the discount factor, Q i Represent two target Q networks.
[0153] The prioritized experience replay buffer 432 stores intervention interaction data and employs a prioritized sampling mechanism to improve learning efficiency for rare but important experiences. The buffer holds 10,000 experiences, each of which contains a state, action, reward, next state, and a termination flag. Priority calculation is based on the absolute value of the TD error, ensuring that high-error samples have a higher sampling probability.
[0154] The behavior cloning pre-training module 433 initializes policy network parameters through imitation learning, accelerating model convergence. This module leverages expert driving data for supervised learning, providing a good initial strategy for reinforcement learning and effectively alleviating the cold start problem. Pre-training uses a cross-entropy loss function to learn expert intervention decision-making patterns.
[0155] Safety constraint optimization objective function 434 ensures that interventions do not introduce new risks by introducing safety constraints. This function maximizes the effectiveness of interventions while considering the safety and comfort of interventions, balancing effectiveness and acceptance. The objective function can be expressed as:
[0156] J(θ)=E[R(τ)-λ1C safety (τ)-λ2C comfort (τ)],
[0157] Among them, θ is the policy parameter, R(τ) is the cumulative reward of trajectory τ, and C safety and C comfort are safety cost and comfort cost respectively, and λ1 and λ2 are trade-off coefficients.
[0158] The dual-delayed policy gradient optimizer43 continuously optimizes intervention strategies through a reinforcement learning framework, enabling the system to adapt to individual driver differences and maintain efficient performance in dynamic driving environments. This optimizer is the core engine of the system's continuous evolution, ensuring the effectiveness, safety, and personalization of intervention strategies.
[0159] Example 9: Driver Profile Module and Environmental Perception
[0160] The system according to claim 9 further includes a driver portrait module 5, a personalized threshold adjustment module 6 and an environment perception module 7.
[0161] Driver Profiling Module 5 uses unsupervised clustering learning to analyze different driving habits. This module collects and analyzes historical driver behavior data to construct a multidimensional feature vector, including control style (such as steering speed and acceleration frequency), attention allocation pattern, and risk response pattern. Using an improved K-means++ algorithm, it categorizes drivers into typical profiles, such as conservative, aggressive, steady, and distracted, providing a personalized foundation for risk assessment and intervention strategies.
[0162] Personalized Threshold Adjustment Module 6 optimizes individual risk thresholds based on reinforcement learning. This module maintains a set of personalized risk threshold parameters for each driver, including warning and intervention thresholds for different risk types. Through an interactive learning process, the system dynamically adjusts these thresholds based on driver feedback and response patterns, balancing sensitivity and specificity to avoid excessive warnings or delayed interventions.
[0163] Environmental Perception Module 7 is responsible for sensing current road and traffic conditions. This module integrates data from the vehicle's existing sensor systems, including environmental information such as speed, steering wheel angle, GPS location, road type, and weather conditions. This information provides crucial context for risk assessment, enabling the system to distinguish between normal and abnormal behavior in different scenarios.
[0164] The spatiotemporal graph neural network risk perception engine 2 dynamically adjusts risk assessment parameters based on the driver profile, personalized thresholds, and environmental conditions. Specifically, the system adjusts the risk model's sensitivity based on the driver profile; determines warning and intervention triggers based on personalized thresholds; and adjusts the contextual weighting of risk assessment based on environmental conditions. This multi-dimensional adjustment mechanism enables the system to maintain optimal performance for different drivers and scenarios.
[0165] Preferably, the system also implements a multi-driver automatic recognition function, automatically identifying the current driver through biometric features (such as sitting posture characteristics, driving habits), loading the corresponding personalized model and parameters, eliminating the need to manually switch driver profiles, thereby improving the system's usability and adaptability.
[0166] Example 10: Intelligent intervention method for driver's dangerous behavior
[0167] like Figure 9 As shown, the intelligent intervention method for driver dangerous behavior based on the above system includes the following steps:
[0168] Step S1: Collect multimodal data of the driver's physiological signals, driving posture and control behavior.
[0169] Specifically, this involves collecting heart rate and respiration signals through millimeter-wave radar, three-dimensional posture data through infrared structured light sensors, and grip force distribution data through a steering wheel pressure sensor grid. This multimodal data is preprocessed and synchronized to form a time-aligned multidimensional feature representation.
[0170] Step S2: Constructing a driver-vehicle interaction dynamic graph structure based on multimodal data.
[0171] The driver's physiological state, posture characteristics, and control behavior are uniformly represented as a graph structure, consisting of skeleton keypoint nodes, physiological state nodes, and vehicle control nodes, as well as spatial correlation edges, temporal correlation edges, and physiological-behavioral interaction edges connecting these nodes. The graph structure is dynamically updated, maintaining a multi-time window graph sequence to achieve comprehensive state representation.
[0172] Step S3: Use the spatiotemporal graph convolutional neural network and the multi-scale temporal convolutional network to process the dynamic graph structure and output the driver risk status assessment result.
[0173] A spatial graph convolutional network processes spatial relationships between nodes, a multi-scale temporal convolutional network handles temporal changes, and a recursive causal attention mechanism extracts long-term dependencies. By integrating multi-dimensional features, the system outputs comprehensive assessment results, including risk level, risk type, and risk trend.
[0174] Step S4: Based on the risk status assessment result, select and execute tactile, visual or vehicle control intervention strategies.
[0175] The intervention method and intensity are dynamically selected based on the severity of the risk assessment results, the risk type, and the driver's personal characteristics. Level 1 risks primarily use tactile feedback, level 2 risks combine tactile and visual feedback, and level 3 risks initiate vehicle control adjustments, achieving a graded intervention response.
[0176] Step S5: Monitor the driver's behavior changes after the intervention and evaluate the intervention effect.
[0177] Track the driver's response to intervention, including response time, behavioral change magnitude, and changes in physiological indicators, comprehensively evaluate the intervention effect, and provide feedback data for strategy optimization.
[0178] Step S6: Based on the intervention effect, the driver risk model and intervention strategy are updated through the double-delayed policy gradient algorithm.
[0179] A dual-Q network evaluates short-term and long-term intervention effectiveness, prioritized experience replay improves learning efficiency in rare scenarios, and safety-constrained optimization targets ensure intervention safety. Continuously updating the driver's personal model and intervention strategy enables system self-optimization and personalized adaptation.
[0180] This approach implements a complete closed-loop process, from data collection to risk prediction, and then to intelligent intervention and model optimization. Leveraging spatiotemporal graph neural network technology, the system deeply integrates multimodal information, enabling accurate modeling of driver status and early risk prediction. Furthermore, through closed-loop optimization mechanisms, the system continuously learns and adapts, providing personalized risk warning and intervention services for different drivers.
[0181] Example 11: System application effect evaluation
[0182] To validate the effectiveness of our system, we conducted large-scale testing in real-world road conditions. The tests involved 100 drivers of varying ages and experience, operating in a variety of environments, including urban, highway, and rural roads. The cumulative mileage exceeded 10,000 kilometers, and over 1,000 hours of driving data were collected.
[0183] Test results show that the system has achieved significant results in key indicators such as the accuracy of dangerous behavior identification, early warning time, and accident rate reduction:
[0184] 1) Dangerous Behavior Recognition Accuracy: The system achieved an average recognition accuracy of 98.6% for four typical dangerous driving behaviors (leaning forward to retrieve an object, eating, turning the head, and picking up an object), 3.6 percentage points higher than the system in reference document CN 106650644B. In particular, in low-light conditions, the system maintained an accuracy of 97.8%, while the conventional system's accuracy dropped to around 85%.
[0185] 2) Early Warning: Utilizing physiological signal and micro-expression analysis, the system can predict potential risks 15 to 30 seconds in advance, providing the driver with ample time to react. For example, in fatigue driving scenarios, the system, by monitoring blink rate and heart rate changes, can detect signs of fatigue an average of 18.3 seconds before traditional systems.
[0186] 3) Reduced false alarm rate: The system's false alarm rate has been reduced from 3.5% in traditional systems to 0.8%, significantly increasing user trust and acceptance. In particular, for object-picking recognition, the system's false alarm rate is only 0.6%, compared to 2.3% in traditional systems.
[0187] 4) Reduced accident rate: In a comparative test in a simulated driving environment, the accident rate of vehicles equipped with this system was 41% lower than that of the control group without the system. Among them, the accidents related to fatigue driving were reduced by 53% and the accidents related to distracted driving were reduced by 45%.
[0188] 5) User Acceptance: A user satisfaction survey showed that 92% of test drivers found the system's interventions natural and non-annoying, and 85% were willing to install the system in their vehicles. This compares to 68% and 52%, respectively, for traditional systems.
[0189] like Figure 10 The comparison charts comparing the recognition performance of four typical dangerous driving behaviors clearly demonstrate the system's stability and accuracy across diverse scenarios and environmental conditions. The system achieved F1 scores exceeding 96% for both leaning forward to retrieve and eating, and over 92% for the more complex head-turning and picking-up behaviors, demonstrating its exceptional performance.
[0190] Through actual application tests, the system of the present invention has proved its practicality and effectiveness in actual road environments, and can significantly improve driving safety and reduce the risk of traffic accidents caused by dangerous driver behavior.
[0191] The proposed system and method for intervening in driver risky behavior, based on a spatiotemporal graph neural network, has broad industrial application prospects. The system can be integrated into the ADAS system of new intelligent connected vehicles or used as a standalone product in the aftermarket, providing driving safety assurance for various vehicles.
[0192] The system's hardware cost is less than 1.3 times that of existing systems, its volume is less than 2L, its power consumption is controlled within 15W, and it can be powered directly by the vehicle's power supply, making it highly viable for mass production. The modular software architecture supports customized deployment at the functional level, adapting to the needs of different vehicle classes.
[0193] The commercial vehicle sector, particularly high-risk scenarios like long-distance freight and passenger transport, is the primary application market for this system. For new energy vehicles, this system can be deeply integrated with intelligent cockpit systems to provide a more intelligent driving experience. With the advancement of autonomous driving technology, this system can also serve as a key safety assurance system in the human-machine co-driving phase, providing intelligent assistance during driver takeover.
[0194] Through multimodal perception and intelligent intervention technology, this invention solves the problems of the existing technology such as single perception source, lack of physiological status monitoring, passive warning mechanism and insufficient adaptability, providing an innovative solution to improve driving safety and has important social value and economic benefits.
[0195] The above description is only a preferred embodiment of the present invention and does not limit the scope of patent protection of the present invention. Any equivalent structural transformation made by using the contents of the present invention description and drawings under the inventive concept of the present invention, or directly / indirectly applied in other related technical fields, shall also fall within the scope of protection of the present invention.
[0196] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A driver dangerous behavior intervention system based on spatiotemporal graph neural network, characterized by: include: Multimodal data acquisition module, used to collect driver physiological signals, driving posture and control behavior data; The spatiotemporal graph neural network risk perception engine is connected to the multimodal data acquisition module and is used to: Receiving multimodal data sent by the multimodal data acquisition module; constructing a driver-vehicle interaction dynamic graph structure based on the multimodal data; Using a multi-scale temporal convolutional network to process the temporal changes of the dynamic graph structure; Output the driver risk status assessment results; An intelligent intervention execution module, connected to the spatiotemporal graph neural network risk perception engine, is used to: receiving the risk status assessment result; selecting a tactile, visual, or vehicle control intervention strategy based on the risk status assessment result; implement the chosen intervention strategy; A closed-loop optimization module, connected to the intelligent intervention execution module and the spatiotemporal graph neural network risk perception engine, is used to: Monitor changes in driver behavior after intervention; Evaluate intervention effectiveness; updating the driver risk model and intervention strategy based on the intervention effect; Also includes: Driver profiling module, used to analyze different driving habits through unsupervised learning clustering; Personalized threshold adjustment module, used to optimize individual risk thresholds based on reinforcement learning; Environmental perception module, used to perceive current road and traffic conditions; Among them, the spatiotemporal graph neural network risk perception engine dynamically adjusts risk assessment parameters according to the driver portrait, personalized thresholds and environmental conditions.
2. The system according to claim 1, wherein: The multimodal data acquisition module includes: Millimeter-wave radar sensor, used to collect the driver's heart rate and breathing rate; Infrared structured light sensor, used to build a 3D skeletal posture model of the driver; A steering wheel pressure sensing grid to capture changes in grip force distribution; A multi-level buffer for storing data collected by the sensor; The timestamp calibration unit is used to synchronize the time of multi-source data.
3. The system according to claim 1, wherein: The spatiotemporal graph neural network risk perception engine includes: A graph structure building module is used to uniformly model the driver's physiological state and behavioral characteristics into a dynamic graph structure; Spatial graph convolution module, used to process the spatial relationship of nodes in the graph structure; Multi-scale temporal convolution module, used to process temporal dimension changes by cascading temporal convolutional networks with dilation factors; Recursive causal attention module for extracting long sequence dependencies; The risk assessment module is used to output the driver's risk status assessment results.
4. The system according to claim 3, characterized in that The dynamic graph structure includes: Skeleton key point nodes, used to represent the position information of the driver's head, torso and limbs; Physiological state nodes are used to represent physiological indicators of heart rate, heart rate variability, and respiratory rate; Vehicle control node, used to represent the steering wheel, pedals, and gear shift control interface; Spatial correlation edges, used to connect related nodes; Time correlation edges, used to capture the temporal change patterns of nodes; Physiological-behavioral interaction edges are used to establish mapping relationships between physiological state nodes and behavioral nodes.
5. The system according to claim 3, wherein: The multi-scale temporal convolution module includes: The short-timescale branch is used to detect sudden dangerous actions; The mesoscale branch is used to monitor recent trends in driving behavior; The long-time-scale branch is used to monitor fatigue cumulative risk factors; The time scale adaptive fusion unit is used to dynamically adjust the importance of different time scales according to the current driving scenario.
6. The system according to claim 1, wherein: The intelligent intervention execution module includes: A tactile feedback channel for providing directional warnings through localized vibrations of the steering wheel; A visual enhancement feedback channel for annotating risk sources via a heads-up display system; Vehicle control adjustment channel, used to automatically adjust vehicle dynamic parameters when the risk level reaches a threshold; The intervention strategy selection unit is used to select the best intervention method based on the risk level and driver characteristics.
7. The system according to claim 1, wherein: The closed-loop optimization module includes: Intervention effect monitoring unit, used to track changes in driver behavior and physiological indicators after intervention; a personalized model updating unit for optimizing the driver's personal model based on intervention response data; A double-delayed policy gradient optimizer is used to evaluate the short-term and long-term intervention effects and update the intervention policy.
8. The system according to claim 7, characterized in that The dual-delay policy gradient optimizer includes: Dual-Q network, used to evaluate short-term intervention effects and long-term safety impacts respectively; Priority experience replay buffer pool, used to store intervention interaction data; Behavior cloning pre-training module to accelerate model convergence; The safety-constrained optimization objective function is used to ensure that interventions do not introduce new risks.
9. The method for intelligently intervening in driver's dangerous behavior based on the system according to any one of claims 1 to 8 is characterized in that: include: Collect multimodal data of driver's physiological signals, driving posture and control behavior; constructing a driver-vehicle interaction dynamic graph structure based on the multimodal data; Processing the dynamic graph structure using a spatiotemporal graph convolutional neural network and a multi-scale temporal convolutional network to output a driver risk status assessment result; selecting and executing a tactile, visual, or vehicle control intervention strategy based on the risk status assessment result; Monitor changes in driver behavior after intervention and evaluate intervention effectiveness; Based on the intervention effect, the driver risk model and intervention strategy are updated through a double-delayed policy gradient algorithm.
Citation Information
Patent Citations
Driver dangerous behavior recognition method and system
CN106650644B
Driving assistance systems and methods
US20150092056A1
Driver state assessment method and apparatus, electronic device, and storage medium
WO2024087205A1