Driver behavior real-time analysis and early warning system based on end-side AI
By integrating a dedicated computing unit and a multimodal data perception module into the edge AI system, localized analysis and differentiated early warning of driver behavior are achieved, which solves the shortcomings of cloud reliance and single perception modality, and improves the accuracy and real-time performance of driver behavior recognition and early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RUIXI TECH (BEIJING) CO LTD
- Filing Date
- 2026-02-02
- Publication Date
- 2026-04-17
AI Technical Summary
Existing driver monitoring systems rely on cloud-based analysis, which is susceptible to network latency and data privacy risks. Their single-sensor modality lacks stability, making it difficult to identify complex and dangerous behaviors. Their warning methods are rigid and disconnected from vehicle status, resulting in limited recognition coverage and high false alarm and false negative rates.
The real-time driver behavior analysis and early warning system using edge AI integrates a dedicated AI computing unit, a multimodal data perception module, and a localized fusion analysis engine. The multimodal data perception module collects and processes driver and vehicle status data in real time, and combines it with vehicle bus data to perform spatiotemporal context correlation analysis to trigger differentiated and graded early warning actions.
It enables end-to-end vehicle-side processing of driver behavior analysis, ensuring the real-time nature and privacy of early warnings, improving the accuracy of identifying complex and dangerous behaviors and the effectiveness of early warnings, adapting to individual driving habits, and reducing the risk of accidents.
Smart Images

Figure CN121884522A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent driving assistance technology, specifically to a real-time driver behavior analysis and early warning system based on edge AI. Background Technology
[0002] Driver monitoring systems (DMS) are a key technology for improving active vehicle safety, aiming to reduce the risk of traffic accidents caused by human factors through real-time monitoring and early warning.
[0003] Currently, mainstream technological approaches face a series of interconnected technical bottlenecks in achieving this goal. First, while cloud-based analytics solutions can utilize powerful computing capabilities, their effectiveness heavily relies on a continuous and stable network connection. In scenarios such as tunnels and remote areas, analysis interruptions or high latency can lead to the failure of real-time warnings. Simultaneously, continuously uploading sensitive data, including facial features, to the cloud also presents risks of data leakage and privacy compliance challenges. Second, to overcome network dependence, some solutions have shifted to edge processing, but due to cost and power consumption limitations, they often employ only a single visible light camera. This single-sensor mode is susceptible to changes in lighting and occlusion interference, resulting in insufficient stability. More importantly, it struggles to accurately identify complex and dangerous behaviors that rely on multi-dimensional information, such as "operating the central control screen without steering wheel input," leading to limited system coverage and high false alarm and false negative rates. Furthermore, the warning triggering logic of existing systems is typically simple and rigid. Warning methods are often singular and fail to dynamically correlate with the vehicle's real-time operating status (such as vehicle speed and road conditions). This could cause warnings to become disconnected from the actual risk level, or cause drivers to become complacent due to frequent and unnecessary prompts, or even actively shut down the system, rendering the safety function ineffective. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a real-time driver behavior analysis and early warning system based on edge AI to solve or improve the technical problems existing in the prior art.
[0005] To address the aforementioned technical problems, this invention provides a real-time driver behavior analysis and early warning system based on edge AI, comprising: The edge-side integrated hardware platform integrates a dedicated AI computing unit for localized artificial intelligence computing, a data synchronization interface module for synchronously collecting multi-source heterogeneous data, and a local early warning output interface for directly driving vehicle-mounted actuators. The multimodal data perception module is connected to the data synchronization interface module and is used to collect and preprocess driver status data, driving behavior data and vehicle operating status data from in-vehicle vision sensors, cabin micro-motion sensors and vehicle bus in real time. The localized fusion analysis engine, which runs in the dedicated AI computing unit, is used to perform parallel feature extraction and spatiotemporal context correlation analysis on the data output by the multimodal data perception module, and generate localized identification results of the driver's dangerous behavior patterns and corresponding risk level assessments. An embedded hierarchical early warning actuator is connected to the localized fusion analysis engine and the local early warning output interface, respectively. It has a built-in configurable early warning rule matrix, which is used to make decisions locally on the vehicle end and trigger differentiated hierarchical early warning actions based on the hazard level assessment and real-time vehicle operating status.
[0006] Compared with existing technologies, the beneficial effects of this invention are as follows: By deploying an integrated hardware platform and localized fusion analysis engine on the vehicle terminal, the entire process of driver behavior analysis is processed on the vehicle side, eliminating the dependence on cloud computing power and stable networks, ensuring the real-time nature of warnings, and eliminating the privacy risks of sensitive data leakage; by adopting multimodal perception such as infrared-RGB dual-mode vision and millimeter-wave radar, and combining vehicle bus data, a robust spatiotemporal context fusion analysis capability is constructed, effectively overcoming the shortcomings of single visual modality being greatly affected by changes in lighting and occlusion interference, and significantly improving the recognition coverage and accuracy of complex and dangerous behaviors such as "operating the central control screen without steering wheel"; by using an embedded hierarchical warning actuator based on dynamic risk assessment results and real-time vehicle status, differentiated warning actions are triggered from visual prompts and auditory warnings to forced tactile intervention, solving the problems of traditional warning methods being single and rigid, and easily leading to driver desensitization, thereby improving the effectiveness of safety intervention and driver acceptance. In addition, the system has local adaptive optimization capabilities, which can continuously fine-tune the model while ensuring data privacy, so as to adapt it to individual driving habits, thereby achieving long-term self-improvement in the system's accuracy and reliability. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a diagram of the overall system architecture of the present invention; Figure 2 This is a flowchart of the multimodal data processing and early warning output of the present invention; Figure 3 This is a flowchart of the embedded hierarchical early warning decision-making process of the present invention; Figure 4 This is a flowchart of the local adaptive optimization process of the present invention. Detailed Implementation
[0009] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose.
[0010] This invention proposes a real-time driver behavior analysis and early warning system based on edge AI, deployed on an in-vehicle terminal, comprising: The edge-side integrated hardware platform integrates a dedicated AI computing unit for localized artificial intelligence computing, a data synchronization interface module for synchronously collecting multi-source heterogeneous data, and a local early warning output interface for directly driving vehicle-mounted actuators. The multimodal data perception module, connected to the data synchronization interface module, is used to collect and preprocess driver status data, driving behavior data, and vehicle operating status data from in-vehicle vision sensors, cabin micro-motion sensors, and vehicle bus in real time. The localized fusion analysis engine, which runs in the dedicated AI computing unit, is used to perform parallel feature extraction and spatiotemporal context correlation analysis on the data output by the multimodal data perception module, and generate localized identification results of the driver's dangerous behavior patterns and corresponding risk level assessments. The embedded hierarchical warning actuator is connected to the localized fusion analysis engine and the local warning output interface, respectively. It has a built-in configurable warning rule matrix, which is used to make decisions locally on the vehicle end and trigger differentiated hierarchical warning actions based on the hazard level assessment and real-time vehicle operating status.
[0011] This embodiment provides a real-time driver behavior analysis and early warning system based on edge AI, deployed on an in-vehicle terminal. The system integrates an edge-to-edge hardware platform designed to support localized artificial intelligence computing. This platform includes a general-purpose processor as a dedicated AI computing unit for executing preset algorithm models. Simultaneously, the platform can be configured with standardized data input ports as data synchronization interface modules, such as USB or Ethernet interfaces, for receiving data from external sensors. Furthermore, the platform can also be equipped with simple relays or digital output ports as local warning output interfaces to control indicator lights or buzzers inside the vehicle.
[0012] Furthermore, the system includes a multimodal data sensing module connected to the aforementioned data synchronization interface module. This sensing module can consist of multiple independent sensors, such as a standard visible light camera as an in-vehicle vision sensor, an ultrasonic sensor as a cabin micro-motion sensor, and a simple CAN bus data reader as a vehicle bus data acquisition device. This sensing module is responsible for acquiring data from these sensors in real time and performing preliminary preprocessing operations such as format conversion or noise filtering to ensure the data can be used by the subsequent analysis engine.
[0013] Building upon this foundation, a localized fusion analysis engine is embedded and runs within the aforementioned dedicated AI computing unit. This engine can be implemented using a rule-based expert system or a simple machine learning classifier. For example, the engine can receive image frames from a visual sensor and extract basic features such as brightness and color; simultaneously, it can receive distance data from a micro-motion sensor and extract its rate of change; and receive speed signals from the vehicle bus. The engine performs simple parallel processing on this data and, through preset logical judgments or threshold comparisons, generates preliminary identification results of driver behavior, such as determining whether the driver has closed their eyes or moved their hands, and provides a corresponding hazard level assessment, such as classifying it as "normal" or "abnormal."
[0014] In addition, the system includes an embedded hierarchical warning actuator, which is connected to both the localized fusion analysis engine and the local warning output interface. This actuator incorporates a simple lookup table as a configurable warning rule matrix, associating different hazard levels (e.g., "abnormal") with preset warning actions (e.g., "issue a buzzer"). When the localized fusion analysis engine outputs a hazard level assessment result, the actuator, based on this assessment result and real-time vehicle operating status obtained from the vehicle bus (e.g., whether the vehicle is in motion), queries the lookup table to make a local decision on the vehicle and trigger the corresponding warning action. For example, when "abnormal" behavior is detected, the actuator can control the local warning output interface to drive the onboard buzzer to sound an alarm.
[0015] The real-time driver behavior analysis and early warning system in this embodiment effectively avoids the network dependence and privacy leakage problems caused by traditional cloud architectures by realizing localized AI computing and multimodal data fusion at the in-vehicle terminal. This system overcomes the limitations of single-perception modality and improves the accuracy of recognizing complex driver behavior patterns. Simultaneously, its tiered early warning mechanism, combined with real-time vehicle operating status, achieves differentiated and dynamic early warning responses, thereby improving the effectiveness of early warnings and driver acceptance, and helping to reduce the risk of traffic accidents.
[0016] In the above embodiments of the present invention, a real-time driver behavior analysis and early warning system based on edge AI is proposed, which includes an integrated edge hardware platform, a multimodal data perception module, a localized fusion analysis engine, and an embedded hierarchical early warning actuator. However, in actual deployment and operation, if the core components of the hardware platform lack specific and efficient implementation methods, the system may face performance bottlenecks, data synchronization difficulties, or poor early warning effects when processing multi-source heterogeneous data, performing localized AI calculations, and executing differentiated early warnings, thereby affecting the system's real-time performance, accuracy, and user experience.
[0017] In response, this invention further proposes that in the aforementioned integrated hardware platform, the dedicated AI computing unit is a neural network processing unit (NPU) embedded in the vehicle terminal; the data synchronization interface module includes: a vision interface for accessing the driver's facial vision sensor, a radar interface for accessing the millimeter-wave radar sensor for micro-motion detection of the steering wheel and central control area, and a vehicle bus interface for directly reading vehicle CAN bus data; the local warning output interface includes: a display driver interface connected to the vehicle instrument panel, an audio control interface connected to the vehicle audio system, and a haptic driver interface connected to the haptic feedback device.
[0018] Specifically, the dedicated AI computing unit is configured as a Neural Processing Unit (NPU) embedded within the in-vehicle terminal. The NPU is a processor specifically designed to accelerate neural network computations, with its architecture optimized for core AI operations such as matrix multiplication and convolution. The NPU offers higher energy efficiency and lower latency than general-purpose CPUs or GPUs, making it particularly suitable for real-time AI inference on resource-constrained edge devices. Embedding the NPU within the in-vehicle terminal means that the computing unit is directly integrated into the vehicle's Electronic Control Unit (ECU) or infotainment system, eliminating reliance on cloud computing and ensuring real-time data processing and privacy. This can be achieved through ASIC (Application-Specific Integrated Circuit) design or flexible configuration using FPGA (Field-Programmable Gate Array) to adapt to the computational needs of different AI models.
[0019] The data synchronization interface module is configured to include a vision interface, a radar interface, and a vehicle bus interface. The vision interface connects to a facial vision sensor, receiving high-bandwidth image or video stream data. This interface typically uses standards such as the MIPICSI mobile industry processor interface, camera serial interface, or USB, and includes an image signal processor (ISP) for initial image denoising, color correction, and format conversion to ensure the quality and compatibility of the visual data. The radar interface is dedicated to connecting a millimeter-wave radar sensor, receiving radar point cloud data or target detection data. This interface can use protocols such as SPI, Ethernet, or CAN to adapt to the data output characteristics of the radar sensor and ensure real-time transmission of micro-motion detection data. The vehicle bus interface is dedicated to directly reading data from the vehicle's CAN controller local area network (CAN bus). This interface includes a CAN transceiver and a CAN controller, capable of parsing the protocol data of communication between various ECUs within the vehicle to obtain key vehicle operating status information such as vehicle speed, steering angle, and gear position, ensuring the real-time nature and accuracy of vehicle dynamic data. The integration of these interfaces enables the system to acquire driver status and vehicle environment information from multiple dimensions, providing a comprehensive data foundation for subsequent fusion analysis.
[0020] The local warning output interface is configured to include a display driver interface, an audio control interface, and a haptic driver interface. The display driver interface displays warning information in graphical or text form on the vehicle's instrument panel. This is achieved through standard interfaces such as LVDS low-voltage differential signal, HDMI, or DisplayPort, ensuring that warning icons, text prompts, or short animations are clearly and promptly displayed to the driver. The audio control interface controls the vehicle's audio system to play warning sounds or voice prompts. This interface can connect to the vehicle's audio system via I2S, SPDIF, or analog audio output to provide warnings with different volumes, tones, or voice content to attract the driver's attention. The haptic driver interface drives haptic feedback devices, such as a steering wheel vibration motor or seat vibrator, to generate haptic alarms. This is achieved through PWM pulse width modulation signals or a dedicated motor driver chip, which can generate vibrations of different intensities, frequencies, or durations according to the warning level, providing intuitive and easily noticeable warnings. The combination of these interfaces allows the system to provide multimodal and differentiated warning feedback based on the level of danger and the scenario, enhancing the effectiveness of the warnings and the driver's perception.
[0021] By concretizing the dedicated AI computing unit into an embedded neural network processing unit (NPU), this system significantly improves the efficiency and real-time performance of localized AI computing, ensuring rapid response in driver behavior analysis and hazard level assessment, and avoiding the latency and data transmission burden of cloud computing. Simultaneously, the data synchronization interface module, through explicit configuration of visual, radar, and vehicle bus interfaces, achieves accurate and synchronous acquisition of multi-source heterogeneous data, effectively solving the problems of inconsistent data formats and timing alignment issues from different sensors, providing high-quality, multi-dimensional data input for the localized fusion analysis engine. Furthermore, the local warning output interface, by integrating display driver, audio control, and tactile driver interfaces, enables the system to flexibly trigger differentiated warnings using visual, auditory, and tactile multimodals based on hazard levels and specific scenarios, greatly enhancing the intuitiveness, enforceability, and effectiveness of warnings. This allows for more timely and effective driver alerts, reducing accident risks and improving driving safety. This specific hardware platform configuration lays a solid foundation for the efficient and reliable operation of the entire system, ensuring stable performance of edge AI in complex in-vehicle environments.
[0022] In some embodiments of the present invention described above, a real-time driver behavior analysis and early warning system based on edge AI is proposed, which collects data through a multimodal data perception module. However, if the raw data from different sensors is not effectively and meticulously preprocessed and feature extracted, the subsequent localized fusion analysis engine will struggle to accurately identify dangerous driver behavior patterns, thereby affecting the reliability and real-time performance of the entire early warning system.
[0023] In response, this invention further proposes a specific implementation of the aforementioned multimodal data perception module, which is used to: acquire the driver's facial image sequence through the aforementioned visual interface, and extract visual feature data including eyelid state, gaze direction, and head posture; acquire point cloud data of the target area in the cockpit through the aforementioned radar interface, and extract the real-time position and motion trajectory data of the driver's hands relative to the steering wheel or a preset unsafe operating area; and acquire and parse vehicle dynamic data including vehicle speed, turn signal status, gear status, and accelerator pedal opening information in real time through the aforementioned vehicle bus interface.
[0024] The multimodal data perception module connects to a driver-facing facial vision sensor, such as a high-frame-rate camera, within the vehicle terminal via the visual interface. This camera continuously captures image sequences or video streams of the driver's facial area. The system then processes these image sequences in real time to extract key visual feature data. Specifically, by applying image processing algorithms and / or pre-trained deep learning models, the driver's eyelid state can be identified and quantified. For example, by calculating the eye aspect ratio (EAR) to determine blink frequency and eyelid closure duration, the degree of fatigue can be assessed. Simultaneously, by analyzing pupil position, eye movement trajectory, and facial key points, the driver's gaze direction can be accurately inferred, determining whether they are focused on the road ahead or exhibiting distracted behavior. Furthermore, through facial key point detection and pose estimation algorithms, the driver's three-dimensional head pose, including pitch, yaw, and roll angles, can be calculated to identify unsafe head postures such as looking down or looking sideways.
[0025] The multimodal data perception module is connected to a millimeter-wave radar sensor in the vehicle terminal via the radar interface. This millimeter-wave radar sensor periodically transmits and receives radar waves towards the target area within the cockpit. By measuring the time delay, Doppler shift, and angle information of the echoes, it generates point cloud data containing the three-dimensional coordinates and velocity of each point. After acquiring the point cloud data, the system performs clustering and segmentation to identify targets such as the driver's hands. Subsequently, using point cloud processing algorithms and / or machine learning models, it extracts the real-time three-dimensional position information of the driver's hands from the identified hand point cloud clusters. To further analyze driver behavior, the system continuously tracks the hand position, generating hand motion trajectory data. This hand position and motion trajectory data is compared in real-time with a pre-modeled three-dimensional spatial area model of the steering wheel and preset non-safe operating areas such as the central control screen and gear lever, thereby determining whether the driver's hands have left the steering wheel or whether there has been any unauthorized blind operation behavior involving touching non-safe operating areas.
[0026] The multimodal data sensing module is connected to the vehicle's CAN bus via the vehicle bus interface. This interface monitors and captures data frames on the CAN bus in real time. The system's built-in CAN protocol parsing module decodes the captured raw CAN data frames according to a preset DBC (DatabaseCAN) file or similar protocol specifications, thereby acquiring and parsing key vehicle dynamic data in real time, including vehicle speed, turn signal status, gear status, and accelerator pedal opening. This vehicle dynamic data provides important contextual information for subsequent driver behavior analysis. For example, vehicle speed can be used to determine the urgency of dangerous behavior, turn signal status and gear status can help determine the driver's intentions, and accelerator pedal opening reflects the vehicle's power output status.
[0027] Through the aforementioned technical solutions, the multimodal data perception module can extract high-value, structured driver state and behavioral characteristics from heterogeneous sensor data. Specifically, facial image sequence analysis of visual data can accurately quantify signs of driver fatigue and distraction; hand target recognition and trajectory tracking of radar point cloud data can monitor in real time whether the driver's hands are off the steering wheel or performing unsafe operations; simultaneously, combined with real-time vehicle operating context provided by vehicle bus data, the localized fusion analysis engine can obtain comprehensive and accurate input information. This significantly improves the system's precision and accuracy in recognizing dangerous driver behavior patterns, providing a solid data foundation for subsequent hazard level assessment and graded warnings, thereby effectively avoiding misjudgments or omissions caused by insufficient raw data processing, and ensuring the reliability and timeliness of the warning system.
[0028] In the above embodiments of the present invention, a method is proposed to collect driver status, driving behavior, and vehicle operating status data in real time through a multimodal data perception module, and process the data using a localized fusion analysis engine. However, in practical applications, data from different sensors often exhibit temporal asynchrony, and the indicative role of each modality's data in indicating dangerous driver behavior may dynamically adjust with changing circumstances. Without an efficient and accurate multi-source data fusion mechanism, the accuracy and real-time performance of behavior recognition results may be limited, making it difficult to effectively support subsequent hazard level assessment and early warning decisions.
[0029] To address this, the present invention further proposes a specific implementation method for a localized fusion analysis engine, comprising: a parallel feature extraction submodule, including a visual analysis model and a radar point cloud analysis model embedded on the NPU, used for real-time processing of the visual feature data and radar point cloud data, respectively; and a spatiotemporal fusion decision submodule, implemented in the form of hardware logic circuits, with built-in timestamp alignment circuits and rule weight matrices; wherein, the spatiotemporal fusion decision submodule is used to receive multiple feature streams from the parallel feature extraction submodule and the vehicle bus, synchronize them through the timestamp alignment circuit, and perform weighted fusion and logical judgment on cross-modal associated features based on the rule weight matrix, outputting a comprehensive behavior recognition result and a hazard level score; its fusion decision logic is based on the following feature function: ;in, Indicates at time The fusion feature output vector; The total number of feature sources participating in the fusion, corresponding to vision, radar, and vehicle status; For the first Each feature source at time... The dynamic weight coefficients are determined by the rule weight matrix based on the context state. For the first Feature extraction functions corresponding to each feature source; For a moment No. Raw or preprocessed data from each feature source; This refers to the parameter set corresponding to the fixed AI model. This feature function is used in the hardware logic to implement weighted fusion calculation of multi-source features under a unified time base.
[0030] It should be noted that there is a clear and direct mapping relationship between the multi-path feature streams and the fusion decision in the spatiotemporal fusion decision submodule, and its specific workflow is as follows: Generation of quantified state indicators: The visual analysis model and radar point cloud analysis model in the parallel feature extraction submodule do not output the original image or point cloud, but rather quantified state indicators that have undergone preliminary abstraction and can be directly used for logical judgment. For example: Visual analysis model output: Normalized gaze deviation Fatigue index wait.
[0031] Radar point cloud analysis model output: Probability of hand presence wait.
[0032] Vehicle bus directly provides: Normalized vehicle speed wait.
[0033] These indicators This constitutes the aforementioned "multi-path feature flow".
[0034] Dynamic weighting and eigenvectors Composition: The characteristic function Its core function is to assign appropriate dynamic weights to each state indicator based on the current context (defined by the rule weight matrix). And perform a weighted summation. It is itself a vector, and its dimensions correspond to the state index values after weight adjustment. In other words, It can be represented as:
[0035] Among them, weight The rule weight matrix is dynamically calculated or selected based on the real-time scene (such as lighting conditions and vehicle status).
[0036] Rule-based decision-making process: The logical judgments do not apply to abstract concepts. Instead of vectors, we use directly. The specific numerical values of the corresponding dimensions in the vector are compared with a preset threshold. For example, the specific execution process for the "high-speed distracted driving" determination rule is as follows: from Extract the dimension value corresponding to "line of sight deviation" ( (or its normalized derivative), to determine whether it continuously exceeds the threshold. .
[0037] from Extract the dimension value corresponding to "probability of hand presence" ( (or its function), to determine whether it is below the threshold (i.e., hand detachment).
[0038] from Extract the dimension value corresponding to "vehicle speed" ( ), to determine whether it is higher than the threshold.
[0039] When all the above conditions are met, the "High-Speed Distracted Driving" behavior tag is triggered. Risk Score Calculation formula variables in , , It comes directly from the composition The original state index of the vector.
[0040] Therefore, the localized fusion analysis engine implements a pipelined processing of "feature extraction - state quantization - dynamic weighting - rule determination". Feature-level fusion ( The generation of the signal provides standardized, weighted, and comparable input signals for decision-level fusion (if-then rules). The two are closely combined through a clear vector dimension mapping relationship to jointly complete the accurate identification and assessment of dangerous driver behaviors.
[0041] Specifically, the parallel feature extraction submodule is responsible for independently and in real-time extracting features from raw or preprocessed data from different sensors. It comprises a visual analysis model and a radar point cloud analysis model embedded in the neural network processing unit (NPU). The visual analysis model employs deep learning networks (e.g., convolutional neural network architectures optimized for embedded devices) to extract visual feature data such as eyelid state, gaze direction, and head posture from driver facial image sequences. The radar point cloud analysis model utilizes point cloud processing algorithms (e.g., clustering and geometric analysis-based methods) to identify and extract real-time position and trajectory data of the driver's hands from point cloud data of target areas within the cockpit. This parallel processing mechanism ensures efficient extraction of features from different modalities, providing timely and rich input for subsequent fusion decisions.
[0042] The spatiotemporal fusion decision submodule is a component for achieving deep fusion of multimodal data. Implemented as hardware logic circuits, it provides efficient and deterministic fusion decision-making capabilities. This submodule incorporates a timestamp alignment circuit and a regular weight matrix. The timestamp alignment circuit receives multiple feature streams from the parallel feature extraction submodule and the vehicle bus and performs precise time synchronization. This is achieved through hardware-level clock synchronization mechanisms, data buffering, and interpolation algorithms to ensure that all feature data participating in the fusion are processed on a unified time reference, effectively addressing potential asynchronous issues during data acquisition and processing from different sensors. The regular weight matrix stores weight coefficients used to dynamically adjust the importance of each modal feature. The system can dynamically adjust based on the current driving context, vehicle status (e.g., speed, lighting conditions), or preset risk assessment strategies. For example, at night or in low light conditions, the weight of radar data can be appropriately increased, while during the day, visual data is given more emphasis. The spatiotemporal fusion decision submodule, based on the received synchronous feature stream and rule weight matrix, outputs a comprehensive behavior recognition result and hazard level score through weighted fusion and logical judgment. Its fusion decision logic utilizes feature functions... The function performs a weighted fusion calculation of multi-source features under a unified time base in the hardware logic, where... For each feature source, there is a feature extraction function. For raw or preprocessed data, This is a set of parameters for solidifying AI models.
[0043] Through the above technical solutions, this invention can efficiently and accurately process multi-source heterogeneous data. The parallel feature extraction submodule enables simultaneous preliminary analysis of visual and radar data, significantly improving the real-time performance of feature extraction. More importantly, the spatiotemporal fusion decision submodule is implemented in the form of hardware logic circuits. The built-in timestamp alignment circuit effectively solves the problem of time asynchrony between different sensor data, ensuring that all input features are fused under a unified time reference. At the same time, the dynamic weighted fusion mechanism based on the rule weight matrix enables the system to intelligently adjust the importance of each modality feature according to the real-time context state, thereby more accurately and robustly identifying dangerous behavior patterns of drivers. This hardware-level fusion decision not only significantly improves processing speed and decision efficiency, but also significantly improves the accuracy and reliability of dangerous behavior identification through precise time synchronization and dynamic weight adjustment. It effectively avoids misjudgments or omissions caused by data inconsistency or improper fusion, providing a solid and real-time decision basis for subsequent differentiated graded early warning.
[0044] In the above embodiments of the present invention, a radar point cloud analysis model is proposed to process point cloud data of the target area within the cockpit to extract real-time position and trajectory data of the driver's hands relative to the steering wheel or a preset unsafe operating area. However, in practical applications, accurately and robustly identifying whether the driver's hands have left the steering wheel and whether there is blind operation of unsafe areas such as the central control screen and gear lever is crucial to ensuring the effectiveness of the driving safety warning system. Relying solely on general point cloud analysis makes it difficult to quantify the precise relationship between the hands and critical operating areas, which may lead to misjudgments or omissions, affecting the timeliness and accuracy of the warning.
[0045] To address this, the present invention further proposes that the radar point cloud analysis model be configured as follows: real-time matching calculation is performed based on the hand target point cloud and a preset three-dimensional spatial region model of the steering wheel to determine whether the driver's hand has left the steering wheel; the movement trajectory of the hand target point cloud is compared with preset unsafe operating areas such as the central control screen and gear lever to determine whether there is any unauthorized blind operation; and the probability of the hand being present... The following formula is used to calculate the result after point cloud clustering: ;in, This represents the number of points in the radar point cloud of the current frame that belong to the hand target cluster; For the th in this cluster Three-dimensional coordinate vectors of points; Distance on the surface of the preset steering wheel area model The coordinate vector of the nearest point; This is the standard deviation parameter for the distance metric. A value close to 1 indicates that the hands are gripping the steering wheel tightly, while a value close to 0 indicates that the hands are not gripping the steering wheel. This value is used to quantify the relative positional relationship between the hands and the steering wheel.
[0046] Specifically, the radar point cloud analysis model is a set of software modules or algorithms running on a dedicated AI computing unit (such as a neural network processing unit, NPU). It is used to identify and track the driver's hands from raw point cloud data acquired from the radar interface and analyze their spatial relationship with key operating areas inside the vehicle. To accurately determine whether the driver's hands have left the steering wheel, the model first needs to identify the "hand target point cloud" from the radar point cloud data. This is done by segmenting the raw point cloud data into different clusters using point cloud clustering algorithms (such as DBSCAN, K-means, etc.), and then using a pre-trained classifier (such as a deep learning-based point cloud segmentation network) to identify these clusters, thereby determining the set of point clouds belonging to the driver's hands. Simultaneously, the system pre-stores a "pre-defined steering wheel 3D spatial region model," which can be a precise geometric representation of the steering wheel (such as a mesh model, parametric surface, or a set of bounding boxes) obtained based on vehicle CAD data, 3D scanning, or manual calibration. During real-time matching calculations, the radar point cloud analysis model continuously calculates the spatial relationship between the hand target point cloud and the three-dimensional spatial region model of the steering wheel. For example, it calculates the shortest distance from each point in the hand point cloud to the surface of the steering wheel model, or evaluates the degree of overlap between the bounding box of the hand point cloud and the bounding box of the steering wheel model. When these distance or overlap indicators exceed a preset threshold, it can be determined that the driver's hands have left the steering wheel.
[0047] Furthermore, to identify unauthorized blind operation by the driver, the radar point cloud analysis model continuously tracks the "motion trajectory" of the hand target point cloud. The hand movement trajectory can be obtained by tracking the centroid or bounding box center of the identified hand target point cloud in consecutive frames of radar point cloud data (e.g., using algorithms such as Kalman filtering). Simultaneously, the system has pre-set 3D spatial models of "pre-defined unsafe operating areas such as the central control screen and gear shift lever," which can also be obtained through CAD data or manual calibration. The radar point cloud analysis model compares the hand movement trajectory with these unsafe operating areas in real time, for example, detecting whether the hand trajectory enters or lingers in these areas for an extended period, or whether there is a specific movement pattern of rapidly moving towards and touching these areas. Once such behavior is detected, the system can determine that unauthorized blind operation has occurred.
[0048] To provide a more refined quantification of the relative positional relationship between the hands and the steering wheel, the radar point cloud analysis model, after point cloud clustering, also calculates the "probability of hand presence". This probability is expressed by the formula. Perform the calculations. Among them, This represents the number of points in the radar point cloud of the current frame that have been identified as the hand target cluster; This indicates the first hand in the cluster. Three-dimensional coordinate vectors of points; Indicates the distance on the surface of the preset steering wheel area model The coordinate vector of the nearest point, which is usually achieved through the nearest neighbor search algorithm; It is a distance metric standard deviation parameter used to adjust for the degree to which distance affects probability decay. The formula quantifies the tightness of the hand-to-steering-wheel contact by taking an exponentially weighted average of the distances between each hand point and the nearest point on the steering wheel. A value close to 1 indicates that the hands are firmly gripping the steering wheel, while a value close to 0 indicates that the hands have left the steering wheel, providing a continuous and quantifiable indicator.
[0049] Through the above technical solutions, the radar point cloud analysis model is endowed with more refined hand behavior recognition capabilities. By performing real-time matching calculations between the hand target point cloud and a preset three-dimensional spatial model of the steering wheel, the system can accurately determine whether the driver's hands have left the steering wheel, effectively avoiding misjudgments caused by posture changes or occlusions in traditional methods. Simultaneously, comparative analysis of hand movement trajectories with preset non-safe operating areas such as the central control screen and gear lever enables the system to promptly detect and identify the driver's illegal blind operation behavior, significantly improving the detection accuracy of distracted driving. Furthermore, the probability of hand presence is introduced... The quantitative calculation provides continuous and detailed data on the relative positions of the hands and steering wheel for the localized fusion analysis engine, enabling the fusion decision submodule to more accurately assess the driver's risk level. This provides a reliable basis for the embedded graded warning actuator to trigger differentiated warning actions, thereby improving the robustness and accuracy of the driver behavior analysis and warning system as a whole.
[0050] In the above embodiments of the present invention, the localized fusion analysis engine is embedded in a dedicated AI computing unit, processing data from facial vision sensors through a visual analysis model to identify the driver's state. However, in actual driving environments, in-vehicle lighting conditions are complex and varied, such as strong light, weak light, backlight, or nighttime. This may cause a single-mode visual sensor to struggle to continuously and stably acquire high-quality facial images, thereby affecting the accuracy of the visual analysis model in extracting facial key points, and consequently reducing the reliability of judging driver fatigue and distraction.
[0051] To address this, the present invention further proposes that the facial visual sensor processed by the visual analysis model is a dual-mode infrared and RGB camera. The visual analysis model is configured to: prioritize processing RGB images under sufficient lighting conditions, and switch to processing infrared images under low light or backlight conditions to continuously and stably extract facial key points; determine fatigue driving state by analyzing the closing frequency and duration of eye contours in a continuous image sequence; determine distracted driving state by calculating the deviation angle and duration of the gaze vector from a preset safe area in front; and determine its fatigue index. Calculated by the following formula: ;in, To slide the time window The number of blinks detected internally; The total number of eyelid closure events within the same window; For the first The duration of the secondary closure event; and These are weighting coefficients, and This formula combines blink frequency and single blink duration to generate a continuous fatigue assessment value.
[0052] Specifically, the facial vision sensor is configured as a dual-mode infrared and RGB camera, integrating both visible light and infrared imaging modes. The RGB mode camera captures color images, providing rich color and texture information, suitable for facial feature recognition under normal lighting conditions. The infrared mode camera actively emits and receives infrared light, enabling it to acquire clear grayscale images even in low-light, no-light, or strong backlight conditions, effectively avoiding the impact of insufficient or overexposed lighting on image quality. This dual-mode design ensures that the system can stably acquire high-quality driver facial image data under various complex lighting conditions, providing a reliable and robust input source for subsequent visual analysis. To fully utilize the advantages of the infrared and RGB dual-mode camera, the visual analysis model is configured to intelligently select the image processing mode based on the current lighting conditions. In practice, the system can either have a built-in light sensor or perform feature analysis on the acquired images, such as brightness, contrast, and saturation, to determine the current lighting conditions inside the vehicle in real time. When sufficient lighting is detected, the system prioritizes processing the RGB image to utilize its rich color information for facial key point extraction. When adverse conditions such as low light, no light, or strong backlight are detected, the system automatically and seamlessly switches to processing infrared images. This dynamic switching mechanism ensures that no matter how the external environment changes, the visual analysis model can continuously and stably extract key feature points such as eyelid state, gaze direction, and head posture from the driver's facial images, providing a consistent and high-quality data foundation for subsequent driver state assessment.
[0053] After acquiring stable and high-quality facial keypoint data, the visual analysis model is further configured to determine driver fatigue by analyzing the driver's eye state. Specifically, the model continuously tracks changes in the driver's eye contour in a series of images, accurately identifying the opening and closing of the eyelids. By real-time monitoring and statistical analysis of the frequency of eyelid closure events (i.e., the number of blinks within a certain time window) and the duration of each closure, the system can capture typical physiological manifestations of driver fatigue. For example, when the driver's blinking frequency is abnormally high or the duration of a single closure is excessively long, these features are identified as signs of fatigue. By comparing with a preset fatigue model or threshold, the system can accurately determine whether the driver is in a state of fatigue. In addition to determining driver fatigue, the visual analysis model is also configured to determine distracted driving by analyzing the driver's gaze direction. Specifically, the model accurately estimates the driver's real-time gaze vector based on facial keypoints (such as pupil position and corners of the eyes). Simultaneously, the system presets a reference vector representing the vehicle's driving direction or safe driving area. By calculating the deviation angle between the driver's real-time gaze vector and the preset safe area directly ahead, and continuously monitoring the duration of this deviation angle, the system can identify whether the driver has been looking away from the road for an extended period. When the gaze deviation angle exceeds a preset threshold and the duration reaches a certain length, it indicates that the driver may be engaging in distracted driving behavior, such as checking the central control screen, communicating with passengers, or observing non-road areas outside the vehicle. To quantify the driver's fatigue level, this invention introduces a fatigue index. The calculation formula is as follows: This formula takes into account the sliding time window. Blink count detected internally Duration of eye closure events .in, and The weighting coefficients are preset and satisfy the following conditions: This index is used to balance the effects of blinking frequency and single eyelid closure duration on fatigue levels. By calculating this index in real time, the system can generate a continuous, numerical assessment of fatigue levels. The higher the value, the deeper the driver's fatigue level. This quantitative assessment method makes the judgment of fatigue status more precise and objective, providing a more accurate basis for subsequent early warning decisions. (Total number of eyelid closure events) And each time By introducing dual-mode infrared and RGB cameras as facial vision sensors and combining them with an intelligent adaptive switching mechanism for lighting conditions, this system effectively overcomes the inherent limitations of traditional single-mode vision sensors in extracting facial features under complex and changing in-vehicle lighting environments (such as strong light, weak light, backlight, and nighttime). This ensures continuous high-precision acquisition of key facial data of the driver (such as eyelid state and gaze direction), significantly improving the robustness of data input. Based on this, through refined analysis of the closing frequency and duration of eye contours in continuous image sequences, and accurate calculation of the deviation angle and duration of the gaze vector from the preset safe area in front, the system can more reliably and accurately identify driver fatigue and distracted driving states. In particular, through fatigue index... The quantitative calculations enable a more continuous and objective assessment of fatigue levels, providing more accurate and detailed driver status input for the localized fusion analysis engine. This not only significantly improves the accuracy and robustness of driver hazardous behavior pattern recognition, effectively avoiding misjudgments or omissions caused by changes in lighting, but also provides a more reliable decision-making basis for the embedded hierarchical warning actuator, thereby improving the timeliness and effectiveness of the entire warning system and ensuring driving safety.
[0054] In the above embodiments of the present invention, the real-time driver behavior analysis and early warning system based on edge AI can collect and preprocess driver state data, driving behavior data, and vehicle operating status data in real time through a multimodal data perception module. This data is then embedded and run in a dedicated AI computing unit by a localized fusion analysis engine, performing parallel feature extraction and spatiotemporal context correlation analysis on this data. This generates localized identification results of driver dangerous behavior patterns and corresponding risk level assessments. However, in actual driving scenarios, driver dangerous behavior is often the result of multiple intertwined factors. Accurately defining and identifying these complex high-risk or medium-risk behavior patterns, and providing a quantitative and actionable risk assessment, is crucial to ensuring the accuracy and reliability of the early warning system.
[0055] To address this, the present invention further proposes rules for logical judgment in the spatiotemporal fusion decision-making submodule, as well as a comprehensive risk score for high-risk behaviors. The method of determination.
[0056] Specifically, the spatiotemporal fusion decision submodule is configured to make logical judgments based on a preset rule matrix. For example, when the system simultaneously detects that the driver's gaze is continuously deviating from the forward safe area for more than a first threshold duration, the hands are determined to be off the steering wheel, and the vehicle speed obtained from the vehicle CAN bus data parsing is higher than a set threshold, the system will comprehensively determine it as a high-risk level "high-speed distracted driving" behavior. "Continuous deviation of the gaze from the forward safe area for more than the first threshold duration" refers to the continuous monitoring of the driver's gaze direction through a visual analysis model. When the gaze deviates from the preset safe area in front of the vehicle's direction of travel (e.g., a certain angle range on both sides of the center line of the road ahead) and continues for more than a preset first threshold time (e.g., 2 seconds or 3 seconds), it is determined to be a distracted state. This first threshold duration can be configured according to the actual application scenario and safety requirements. "Hands are determined to be off the steering wheel" refers to the analysis of the hand target point cloud through a radar point cloud analysis model. When the relative positional relationship between the hands and the steering wheel does not conform to a normal grip state, for example, the hands have a probability of being off the steering wheel... When the speed falls below a certain preset threshold, it is considered as if the driver has taken their hands off the steering wheel. This determination takes into account various situations, such as the driver taking both hands off the steering wheel or operating the steering wheel with one hand but with improper posture. "The vehicle speed parsed from the vehicle CAN bus data is higher than the set threshold" means that the system obtains vehicle speed information from the vehicle CAN bus in real time and compares it with the preset safe speed threshold. This set threshold is usually determined based on road type, traffic regulations, or vehicle safety standards, such as the minimum speed limit on highways or a certain dangerous speed limit. When the vehicle speed exceeds this threshold, any distracted behavior may lead to more serious consequences.
[0057] Furthermore, when the system identifies that the driver's eyelid closure frequency and duration conform to a fatigue model, and the vehicle is traveling at a constant speed in a straight line, the system will comprehensively determine it as a medium-risk level "potential fatigued driving" behavior. The "eyelid closure frequency and duration conforming to a fatigue model" involves continuously monitoring the driver's eye state through a visual analysis model. Based on parameters such as blinking frequency and single closure duration, combined with a preset fatigue driving discrimination model, it determines whether the driver is fatigued. This fatigue model can quantify the degree of driver fatigue. "The vehicle is traveling at a constant speed in a straight line" refers to obtaining information such as vehicle speed, steering angle, and lateral acceleration through vehicle CAN bus data. When vehicle speed fluctuations are small, steering angle is close to zero, and lateral acceleration is low, the vehicle can be determined to be in a relatively stable state of constant speed in a straight line. In this state, driver fatigue may be more obvious, and the warning interference is relatively small.
[0058] For high-risk behaviors, the comprehensive risk score Determined by the following multi-condition joint decision function: In this formula, This indicates the angle of visual deviation; the larger the value, the more serious the driver's visual deviation from the main road direction. The safety line-of-sight angle threshold is used to define the critical point of line-of-sight deviation. This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. It is used to convert discrete condition judgments into numerical values for calculation. This represents the probability of the hand being off the steering wheel; the closer the value is to 0, the greater the likelihood that the hand will leave the steering wheel. Real-time vehicle speed, reflecting the vehicle's motion status; The maximum reference speed preset by the system is used to normalize the real-time vehicle speed. These are the normalized weight coefficients of each risk factor, satisfying... These weighting coefficients can be adjusted based on actual driving scenarios, safety strategies, or expert experience to reflect the relative importance of different risk factors in the overall risk assessment. This formula achieves a quantitative assessment of high-risk behaviors by mapping detection results from different modalities to a unified risk score.
[0059] Through the above technical solution, this invention overcomes the limitations of traditional systems in recognizing dangerous behavior patterns in complex driving scenarios, such as insufficient accuracy and lack of quantitative standards for risk assessment. The spatiotemporal fusion decision submodule, by introducing explicit logical judgment rules, effectively integrates the discrete features output by the multimodal data perception module and the parallel feature extraction submodule into practically meaningful dangerous behavior patterns, such as "distracted driving at high speed" and "potential fatigue driving." This mechanism based on multi-condition joint judgment significantly improves the accuracy and robustness of dangerous behavior recognition, avoiding false alarms or missed alarms that may result from single-feature judgments. Simultaneously, a comprehensive risk score is introduced for high-risk behaviors. The formula deflects the line of sight at an angle The probability of hand presence and real-time vehicle speed By weighting and fusing key risk factors, a continuous and quantitative risk assessment value is generated. This enables the system not only to identify dangerous behaviors but also to accurately assess their risk levels, thus providing more refined decision-making basis for embedded hierarchical early warning actuators and achieving differentiated and intelligent early warning responses. This quantitative assessment mechanism allows the system to more objectively and comprehensively reflect the driver's real-time risk status, thereby effectively improving the overall performance and practical value of the driver behavior analysis and early warning system.
[0060] In the above embodiments of the present invention, a real-time driver behavior analysis and early warning system based on edge AI is proposed. This system can locally identify dangerous behavior patterns of drivers and assess their hazard levels, and issue warnings through an embedded hierarchical early warning actuator. However, if the early warning method lacks specificity and fails to provide corresponding warning intensity and modality based on differences in hazard levels, it may lead to unclear perception of different risks by drivers, making it impossible to effectively distinguish the severity and urgency of risks. This would affect the timeliness and effectiveness of the early warning, and may even result in false alarms due to a single warning method, reducing drivers' trust in the early warning system.
[0061] In response, this invention further proposes that the embedded graded warning actuator performs graded warning actions including: primary warning: when the behavioral hazard level is assessed as low risk, the warning icon on the dashboard flashes via the display driver interface; intermediate warning: when the behavioral hazard level is assessed as medium risk, a short warning sound is triggered via the audio control interface, while the warning icon is maintained; advanced warning: when the behavioral hazard level is assessed as high risk, a strong tactile alarm is generated by activating the vibration motor in the steering wheel or seat via the tactile driver interface, and the volume of the in-vehicle entertainment system is simultaneously reduced via the audio control interface.
[0062] Specifically, the embedded graded warning actuator is configured to perform corresponding graded warning actions based on the hazard level assessment results output by the localized fusion analysis engine. When the system assesses the driver's behavior as low-risk, a primary warning is triggered. The primary warning controls the flashing of a warning icon on the vehicle's instrument panel via a display driver interface. This display driver interface is connected to the instrument panel and can send specific control signals, such as via CAN bus protocol or video signals like LVDS / HDMI, to activate preset warning lights on the instrument panel or render flashing icons on a digital display screen. The flashing frequency and color of the warning icon can be pre-configured to subtly alert the driver to potential risks without excessively distracting them.
[0063] When the system assesses the driver's behavior as being in a medium-risk state, a medium-level warning will be triggered. While maintaining the warning icon display, the medium-level warning triggers a short warning sound via the audio control interface. This audio control interface connects to the vehicle's audio system and can send audio playback commands, such as via I2S, SPDIF, or analog audio signals, to play a preset warning sound effect. This warning sound is typically designed to be short, clear, and have a certain degree of penetration to effectively attract the driver's attention. The combined visual and auditory warning modalities enhance the warning effect and encourage the driver to increase alertness.
[0064] When the system assesses that the driver's behavior is at high risk, an advanced warning will be triggered. The advanced warning activates a vibration motor in the steering wheel or seat via a haptic drive interface, generating a strong haptic alarm, and simultaneously lowers the volume of the in-vehicle entertainment system via an audio control interface. The haptic drive interface connects to the vibration motor inside the steering wheel or seat, driving it to produce strong, continuous vibrations by sending PWM signals or digital control commands, directly impacting the driver's body and forcing them to perceive the danger. Simultaneously, the audio control interface sends volume adjustment commands to the in-vehicle entertainment system, such as via a CAN bus or dedicated control protocol, to lower or mute the entertainment system's volume to ensure the warning sound is clearly audible and to avoid interference. This multi-sensory, mandatory warning method aims to maximize driver awareness and prompt immediate corrective action.
[0065] Through the aforementioned technical solution, this system can intelligently select and execute differentiated, tiered warning actions based on the risk level assessment results of driver behavior. For low-risk situations, gentle visual cues are used to avoid excessive interference with driving; for medium-risk situations, visual and auditory warnings are combined to enhance driver alertness; and for high-risk situations, multimodal, mandatory warning methods such as tactile, auditory, and ambient sound levels are used to ensure that the driver can quickly and clearly perceive the urgency of the danger and take immediate intervention measures. This refined, tiered warning mechanism effectively solves the problems of poor warning effectiveness and unclear driver perception that may result from a single warning method, improves the user experience of the warning system and the driver's trust in the system, thereby significantly improving driving safety and the effectiveness of warnings.
[0066] While edge AI-based real-time driver behavior analysis and early warning systems can identify and warn drivers of dangerous behaviors in real time, their built-in AI models are typically fixed after deployment. This means the system may struggle to adapt to individual driver habits, long-term vehicle wear and tear, or subtle environmental features, potentially leading to decreased accuracy of warnings over time, false alarms, or missed warnings, thus impacting user experience and the system's long-term effectiveness.
[0067] To address this, the present invention further proposes a local adaptive optimization module. This module enables the system to better adapt to the habits of specific drivers and the vehicle's operating environment through continuous learning and adjustment. Its function is to achieve localized and personalized optimization of the AI model, thereby improving the long-term accuracy and reliability of the warning system.
[0068] This local adaptive optimization module includes a secure storage area and offline fine-tuning routines. The secure storage area is designed to securely store characteristic data of typical scenarios that trigger warnings locally within the vehicle. To protect driver privacy, all stored data is anonymized and de-identified, ensuring it cannot be traced back to any specific individual. This data forms the basis for subsequent system learning and optimization, providing valuable samples of real-world driving behavior and warning events.
[0069] The offline fine-tuning routine is an intelligent learning mechanism that operates under conditions set to detect when the vehicle is off and the system is idle, in order to avoid affecting normal vehicle operation and driving safety. When these conditions are met, the routine automatically starts, utilizing anonymized data accumulated in a secure storage area to perform supervised incremental learning and parameter fine-tuning on at least one AI model in the localized fusion analytics engine. This fine-tuning process aims to enable the AI model to better understand and predict the specific driving habits of the vehicle's driver, thereby optimizing recognition accuracy.
[0070] The fine-tuning process minimizes the following local loss function. To update model parameters Θ : ; in, This refers to the number of samples used for this round of fine-tuning in the secure storage area; For the first Feature data of each sample and its corresponding warning label; This is the model's current prediction output; The cross-entropy loss function; These are the initial parameters for the model; This is the regularization strength coefficient. The function consists of two parts: the first part is the cross-entropy loss function. Used to measure the current predicted output of the model. With real warning labels The differences between them drive the model to learn in a more accurate direction; the second part is the regularization term. The function of this regularization term is to restrict the model parameters. Deviation from initial parameters during fine-tuning This ensures that the model can adapt to individual habits to a certain extent, thereby effectively preventing catastrophic forgetting when adapting to individual habits. In other words, it prevents the model from forgetting the general knowledge and performance it has already mastered when learning new knowledge, and ensures that the model maintains its original reliability and generalization ability while adapting to individual needs.
[0071] By introducing a local adaptive optimization module, the system of this invention can overcome the problem of insufficient adaptability that may occur in traditional fixed AI models during long-term use. Specifically, the secure storage area anonymizes and desensitizes the characteristic data of typical scenarios that trigger warnings, providing a secure and compliant data foundation for subsequent learning. When the vehicle is off and the system is idle, the offline fine-tuning routine uses this data to perform supervised incremental learning and parameter fine-tuning of the AI model in the localized fusion analysis engine. This process minimizes the local loss function. This loss function not only drives the model to learn new driving habit features, but also effectively prevents catastrophic forgetting when adapting to individual habits through regularization terms, ensuring that the model retains its original general recognition ability after personalized optimization. Therefore, the system of this invention can continuously optimize the recognition accuracy for the vehicle's driving habits, significantly reduce false alarms and false negatives, and improve the accuracy and timeliness of warnings, thereby providing drivers with more personalized, reliable, and long-term effective driving safety protection, greatly enhancing the system's practicality and user acceptance.
[0072] The present invention also provides an embodiment of a practical application of the system of the present invention: I. Application Scenarios: 1.1 Test vehicle and driver information: Vehicle model: A mid-size SUV of a certain brand (2023 model); In-vehicle terminal: Intelligent cockpit system with integrated NPU (computing power: 4TOPS); Test section: G42 Shanghai-Chengdu Expressway, Nanjing to Hefei section (night, clear weather, moderate traffic). Driver: Male, 35 years old, 10 years of driving experience; Test period: After 3 hours of continuous driving (20:00-23:00).
[0073] 1.2 System Configuration Parameters:
[0074] II. Specific Implementation and Data Recording at Each Stage: S1 multimodal data real-time acquisition and preprocessing: Table 1: Monitoring Dimensions and Historical Benchmark Data (Statistics for the Past Month)
[0075] Table 2: Real-time monitoring data (early warning trigger period: t1-t4, a total of 60 seconds)
[0076] S2 localized fusion analytics engine performs real-time calculations: 2.1 Visual Feature Extraction and Fatigue Index Calculation: Sliding time window: seconds ; Weighting coefficients: ; Fatigue index calculation (time t3):
[0077] (where, seconds) ).
[0078] 2.2 Radar point cloud analysis and probability of hand presence: Steering wheel area model: Preset 3D mesh (CAD import); Standard deviation parameter: ; Calculation of the probability of hand presence (time t4):
[0079] (Number of cluster points) ) 2.3 Calculation of Comprehensive Risk Score: Risk factor weights: ; Safe line-of-sight threshold: ; Maximum reference speed: ; Overall risk score (time t4):
[0080] S3 Embedded Hierarchical Early Warning Actuator Response:
[0081] S4: Local Adaptive Optimization Module (Learned After the Fact): 4.1 Secure storage area data logging: Storage content: 60 seconds of multimodal feature data that triggered the alert (de-identified); Data size: Approximately 2MB (encrypted storage); Tags: High-risk driving (distraction + fatigue); 4.2 Offline fine-tuning process (after the vehicle is turned off); Fine-tuning the model: Visual analysis model (fatigue detection branch); Sample size: ; Regularization strength: ; Loss function decreases: from initial Down to ; Fine-tuning results: The model's recognition accuracy improved by approximately 12% for the driver's nighttime fatigue characteristics.
[0082] III. Effect Verification and Comparative Analysis: Table 3: Comparison with traditional early warning systems
[0083] Table 4: Summary of the Effect of this Early Warning Event
[0084] IV. Conclusion: This practical application verification shows that: Multimodal data fusion significantly improves recognition accuracy, especially in nighttime environments, where infrared-RGB dual-mode cameras ensure stable extraction of visual features; a tiered warning mechanism effectively balances safety and driving experience, with progressive warnings from prompts to mandatory intervention being more acceptable; edge computing ensures real-time performance and privacy, with a response latency of less than 1 second throughout the process; and local adaptive optimization capabilities allow the system to learn autonomously after the vehicle is turned off, gradually adapting to the driver's individual habits.
[0085] Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art and related fields based on the embodiments of the present invention without inventive effort should fall within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described and explained in the present invention, unless otherwise specified or limited, shall be implemented according to conventional means in the art.
Claims
1. A real-time driver behavior analysis and early warning system based on edge AI, deployed on an in-vehicle terminal, characterized in that, include: The edge-side integrated hardware platform integrates a dedicated AI computing unit for localized artificial intelligence computing, a data synchronization interface module for synchronously collecting multi-source heterogeneous data, and a local early warning output interface for directly driving vehicle-mounted actuators. The multimodal data perception module is connected to the data synchronization interface module and is used to collect and preprocess driver status data, driving behavior data and vehicle operating status data from in-vehicle vision sensors, cabin micro-motion sensors and vehicle bus in real time. The localized fusion analysis engine, which runs in the dedicated AI computing unit, is used to perform parallel feature extraction and spatiotemporal context correlation analysis on the data output by the multimodal data perception module, and generate localized identification results of the driver's dangerous behavior patterns and corresponding risk level assessments. An embedded hierarchical early warning actuator is connected to the localized fusion analysis engine and the local early warning output interface, respectively. It has a built-in configurable early warning rule matrix, which is used to make decisions locally on the vehicle end and trigger differentiated hierarchical early warning actions based on the hazard level assessment and real-time vehicle operating status.
2. The real-time driver behavior analysis and early warning system based on edge AI according to claim 1, characterized in that, In the aforementioned integrated edge hardware platform: The dedicated AI computing unit is a neural network processing unit (NPU) embedded in the vehicle terminal. The data synchronization interface module includes: a vision interface for accessing a driver-facing facial vision sensor, a radar interface for accessing a millimeter-wave radar sensor for detecting micro-motions in the steering wheel and center console area, and a vehicle bus interface for directly reading vehicle CAN bus data. The local warning output interface includes: a display driver interface connected to the vehicle instrument panel, an audio control interface connected to the vehicle audio system, and a haptic driver interface connected to the haptic feedback device.
3. The real-time driver behavior analysis and early warning system based on edge AI according to claim 2, characterized in that, The multimodal data sensing module is specifically used for: The driver's facial image sequence is obtained through the visual interface, and visual feature data including eyelid state, gaze direction and head posture are extracted. The radar interface is used to acquire point cloud data of the target area in the cockpit and extract the real-time position and movement trajectory data of the driver's hands relative to the steering wheel or a preset unsafe operating area. The vehicle bus interface is used to acquire and parse dynamic vehicle data in real time, including information on vehicle speed, turn signal status, gear status, and accelerator pedal opening.
4. The real-time driver behavior analysis and early warning system based on edge AI according to claim 3, characterized in that, The localized fusion analysis engine includes: The parallel feature extraction submodule includes a visual analysis model and a radar point cloud analysis model that are embedded and run on the NPU, respectively used to process the visual feature data and radar point cloud data in real time. The spatiotemporal fusion decision submodule is implemented in the form of hardware logic circuits and has built-in timestamp alignment circuits and rule weight matrices. The spatiotemporal fusion decision submodule receives multiple feature streams from the parallel feature extraction submodule and the vehicle bus, synchronizes them through the timestamp alignment circuit, and performs weighted fusion and logical judgment on the cross-modal associated features based on the rule weight matrix, outputting a comprehensive behavior recognition result and a hazard level score; its fusion decision logic is based on the following feature function: ; in, Indicates at time The fusion feature output vector; The total number of feature sources participating in the fusion, corresponding to vision, radar, and vehicle status; For the first Each feature source at time... The dynamic weight coefficients are determined by the rule weight matrix based on the context state. For the first Feature extraction functions corresponding to each feature source; For a moment No. Raw or preprocessed data from each feature source; This refers to the parameter set corresponding to the fixed AI model.
5. The real-time driver behavior analysis and early warning system based on edge AI according to claim 4, characterized in that, The radar point cloud analysis model is configured as follows: Real-time matching calculations are performed based on the hand target point cloud and the preset three-dimensional spatial area model of the steering wheel to determine whether the driver's hands have left the steering wheel; The motion trajectory of the hand target point cloud is compared with the preset non-safe operation areas such as the central control screen and gear lever to determine whether there is any unauthorized blind operation behavior. The probability of its hand being present The following formula is used to calculate the result after point cloud clustering: ; in, This represents the number of points in the radar point cloud of the current frame that belong to the hand target cluster; For the th in this cluster Three-dimensional coordinate vectors of points; Distance on the surface of the preset steering wheel area model The coordinate vector of the nearest point; This is the standard deviation parameter for the distance metric. A value close to 1 indicates that the hands are gripping the steering wheel tightly, while a value close to 0 indicates that the hands are not gripping the steering wheel. This value is used to quantify the relative positional relationship between the hands and the steering wheel.
6. The real-time driver behavior analysis and early warning system based on edge AI according to claim 4, characterized in that, The facial visual sensor processed by the visual analysis model is an infrared and RGB dual-mode camera, and the visual analysis model is configured as follows: Under sufficient lighting conditions, RGB images are processed first, while under low lighting or backlight conditions, infrared images are processed to continuously and stably extract facial key points. The state of fatigued driving can be determined by analyzing the closing frequency and duration of the eye contour in a continuous image sequence. Distracted driving status is determined by calculating the deviation angle and duration between the line-of-sight vector and the preset safe zone directly in front; Its fatigue index Calculated by the following formula: ; in, To slide the time window The number of blinks detected internally; The total number of eyelid closure events within the same window; For the first The duration of the secondary closure event; and These are weighting coefficients, and .
7. The real-time driver behavior analysis and early warning system based on edge AI according to claim 4, characterized in that: The rules for logical judgment in the spatiotemporal fusion decision submodule include: When the following are simultaneously detected: "the line of sight is continuously deviated from the forward safe area for more than the first threshold duration", "the hands are judged to be off the steering wheel", and "the vehicle speed obtained from the vehicle CAN bus data is higher than the set threshold", it is comprehensively judged as a high-risk level "high-speed distracted driving" behavior. When it is identified that "the frequency and duration of eyelid closure conform to the fatigue model" and "the vehicle is in a straight line at a constant speed", it is comprehensively judged as a "potential fatigue driving" behavior of medium risk level; Among them, for high-risk behaviors, the comprehensive risk score Determined by the following multi-condition joint decision function: ; in, The angle of deviation from the line of sight; The safe line-of-sight angle threshold; This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. The probability of the hand being present; Real-time vehicle speed; The maximum reference speed preset by the system; Let be the normalized weight coefficients of each risk factor, satisfying . .
8. The real-time driver behavior analysis and early warning system based on edge AI according to claim 2, characterized in that, The hierarchical early warning actions executed by the embedded hierarchical early warning actuator include: Primary warning: When the risk level of the behavior is assessed as low, the warning icon on the dashboard will flash via the display driver interface; Intermediate warning: When the behavioral hazard level is assessed as medium risk, a short warning sound is triggered through the audio control interface, while the warning icon is displayed. Advanced warning: When the behavioral hazard level is assessed as high risk, a strong tactile alarm is generated by activating the vibration motor in the steering wheel or seat through the tactile drive interface, and the volume of the in-vehicle entertainment system is simultaneously reduced through the audio control interface.
9. The real-time driver behavior analysis and early warning system based on edge AI according to claim 1, characterized in that, It also includes a local adaptive optimization module, which includes: A secure storage area is used to anonymize and store typical scenario feature data that triggers warnings locally on the vehicle; the data has been anonymized. The offline fine-tuning routine runs automatically when the vehicle is detected to be off and the system is idle. It uses data in the secure storage area to perform supervised incremental learning and parameter fine-tuning on at least one AI model in the localized fusion analysis engine to optimize the recognition accuracy for the vehicle's driving habits. The fine-tuning process minimizes the following local loss function. To update model parameters : ; in, This refers to the number of samples used for this round of fine-tuning in the secure storage area; For the first Feature data of each sample and its corresponding warning label; This is the model's current prediction output; The cross-entropy loss function; These are the initial parameters for the model; This is the regularization intensity coefficient.