Driving safety early warning system based on video monitoring vehicle-mounted terminal

By collecting multi-source data through video surveillance vehicle terminals to construct personalized driver profiles and combining them with environmental risk situation maps for comprehensive risk assessment, the problem of poor targeted early warning in existing systems has been solved, and accurate driving safety early warning has been achieved.

CN121921903AInactive Publication Date: 2026-04-24SHANXI ZHONGHUAN SATELLITE NAVIGATION COMMUNICATIONS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANXI ZHONGHUAN SATELLITE NAVIGATION COMMUNICATIONS CO LTD
Filing Date
2026-01-22
Publication Date
2026-04-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing driving safety warning systems fail to establish a coupled correlation model between the driver's real-time status and environmental risks, resulting in poor warning targeting and a tendency for false alarms, missed alarms, or warning methods that are not suitable for the driver.

Method used

By collecting multi-source heterogeneous data through video surveillance vehicle terminals, a personalized driver profile is constructed. Combined with an environmental risk situation map, a nonlinear coupling model is used to assess the comprehensive risk level and generate an adaptive early warning strategy.

Benefits of technology

It achieves a comprehensive risk assessment that accurately quantifies driver status and environmental risks, avoiding missed and false alarms, and improving the system's early warning performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_17
    Figure SMS_17
  • Figure SMS_92
    Figure SMS_92
Patent Text Reader

Abstract

The invention provides a driving safety early warning system based on a video monitoring vehicle-mounted terminal, which relates to the technical field of driving safety and comprises a data acquisition and fusion module, a personalized driver portrait generation module, an environmental risk cognitive prediction module, a risk coupling evaluation decision module and a multi-mode self-adaptive early warning module. According to the method, the personalized driving portrait fusing the real-time state and the long-term habit is constructed, so that the early warning strategy can accurately adapt to the individual characteristics of different drivers, a coupling evaluation model of the driver state and the environmental risk is established instead of independently considering the single-dimensional risk, the accurate quantification of the comprehensive risk level is realized, and the driving safety is improved. According to the method and the system, missing report and false report caused by risk assessment splitting of an existing system are avoided, the anti-interference capability of the system is improved through complementary verification of multi-source data, and the stable early warning performance can still be kept even in complex scenes such as low illumination and target shielding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of driving safety technology, specifically relating to a driving safety early warning system based on a video surveillance vehicle terminal. Background Technology

[0002] With the increase in car ownership and the growing complexity of road traffic, driving safety warning systems have become one of the key technologies for improving driving safety. However, existing driving safety warning systems have many problems that need to be addressed: Firstly, existing technologies mostly adopt uniform warning thresholds and strategies, without taking into account the driver's real-time state, such as fatigue, distraction, and emotions, as well as long-term driving habits, such as following distance and lane-changing style. This results in poor warning targeting and is prone to false alarms, missed alarms, or warning methods that are not suitable for the driver. Secondly, existing technologies only consider environmental risks or driver status separately, without establishing a coupled correlation model between the two, and cannot accurately quantify the comprehensive risks under the interaction of driver status and environmental risks. Summary of the Invention

[0003] This invention provides a driving safety warning system based on a video surveillance vehicle terminal to solve at least one of the technical problems mentioned above.

[0004] To address the aforementioned technical problems, this invention discloses a driving safety early warning system based on a video surveillance vehicle terminal, comprising: The data acquisition and fusion module is used to collect and synchronously fuse multi-source heterogeneous data streams from the in-vehicle driver monitoring camera, the external environment perception camera and the vehicle CAN bus in real time, and generate fused perception data corresponding to the current driver. The personalized driver profile generation module is used to construct short-term state vectors and long-term habit vectors based on the fused perception data corresponding to the current driver, and to fuse the short-term state vectors and long-term habit vectors to generate a personalized driving profile of the current driver. The environmental risk perception and prediction module is used to identify dynamic targets and static traffic elements around vehicles based on fused perception data, predict the risk evolution trend in the short time domain through time-series trajectory modeling, and generate a current environmental risk situation map. The risk coupling assessment and decision-making module is used to assess the comprehensive risk level under the interaction between driver status and environmental risk based on the current driver's personalized driving profile and the current environmental risk situation map through a nonlinear coupling model, and generate the corresponding early warning and intervention decision. The multimodal adaptive warning module is used to generate the final warning command based on the current warning intervention decision, the current driver's personalized driving profile, and the driver's historical warning feedback data.

[0005] Preferably, the data acquisition and fusion module includes: The driver identification unit is used to identify the unique identity of the current driver based on the facial image obtained by the in-vehicle driver monitoring camera and to identify whether the current driver is an experienced driver or a new driver through a preset facial recognition model. When the driver is identified as a new driver, a personalized driving profile file corresponding to the identity is created. The driver data acquisition unit is used to acquire an original video stream containing the current driver's facial image, eye features, head posture, and gestures based on the in-vehicle driver monitoring camera, and to extract first temporal features based on the original video stream. The first temporal features include the eye opening sequence, gaze direction vector sequence, head three-dimensional Euler angle sequence, facial key point motion sequence, and hand key point coordinate sequence extracted from the original video stream and arranged in chronological order. The environmental data acquisition unit is used to acquire a second original video stream containing dynamic targets and static traffic elements around the vehicle based on the external environment perception camera, and to extract a second temporal feature based on the second original video stream; wherein, the second temporal feature includes a sequence of bounding boxes, a sequence of target category labels, and a sequence of target motion state vectors of each detected target extracted from the second original video stream and arranged in chronological order. The vehicle data acquisition unit is used to acquire in real time a sequence of vehicle dynamic parameters, including vehicle speed, longitudinal acceleration, steering wheel angle, yaw rate and braking status, through the vehicle CAN bus interface. The spatiotemporal synchronization fusion unit is used to perform timestamp alignment, coordinate system unification, and data normalization on the first time-series features, the second time-series features, and the vehicle dynamic parameter sequence to generate fused perception data.

[0006] Preferably, the personalized driver profile generation module includes: The short-term state analysis submodule is used to acquire, in real time, the driver's fatigue index, attention distraction index, emotional agitation index, and operational compliance index based on the fused perception data corresponding to the current driver, and to determine the driver's fatigue index. Distraction index Emotional Excitement Index and control standardization index Construct short-time state vectors; The long-term habit learning submodule, when the driver identification unit identifies the driver as an experienced driver, inputs the fused perception data of the current driver within the historical driving cycle into the trained following distance output model, lane change style evaluation model, and risk response time output model, respectively, to obtain the driver's following distance habit value, lane change style type, and risk response habit reaction time. Based on the following distance habit value, lane change style type, and risk response habit reaction time, the module constructs the driver's long-term habit vector; among which, the lane change style type includes aggressive, steady, and decisive. When the driver identification unit identifies the driver as a new driver, the default long-term habit vector is loaded as the initial value. The driving profile generation submodule is used to fuse the short-term state vector obtained by the short-term state analysis submodule and the long-term habit vector obtained by the long-term habit learning submodule to generate a personalized driving profile of the current driver.

[0007] Preferably, the short-time state analysis submodule includes: The eye state monitoring unit is used to calculate the driver's fatigue index within the current preset time window based on the current driver's eye opening and closing sequence, head three-dimensional Euler angle sequence, and facial key point motion sequence from the fused perception data. ; The gaze direction analysis unit is used to calculate the driver's attention distraction index within the current preset time window based on the current driver gaze direction vector sequence and hand key point coordinate sequence from the fused perception data. ; The facial emotion recognition unit is used to calculate the driver's emotional arousal index within a current preset time window based on the motion sequence of key facial points in the fused perception data. ; The driving behavior evaluation unit is used to calculate the driver's driving compliance index within the current preset time window based on the sequence of vehicle dynamic parameters obtained from the vehicle bus. .

[0008] Preferably, the facial emotion recognition unit includes: The emotion-related key point grouping and parsing subunit is used to divide key point groups into two core emotions, driving stress and anger, based on the facial key point motion sequence in the fused perception data and extract the motion features of the corresponding key points. The key point groups include stress emotion association group and anger emotion association group. The stress-related emotion group includes four key points on the inner and outer sides of both eyebrows, and six key points on the upper and lower edges of both eyelids and the corners of the eyes; the anger-related emotion group includes four key points on the inner and outer sides of both eyebrows, and four key points on the upper and lower lips; the corresponding facial movements for the stress-related emotion group include drooping eyebrows, tight eyelids, abnormal blinking frequency, and stiff eye muscles; the corresponding facial movements for the anger-related emotion group include furrowed eyebrows, retracted corners of the mouth, pursed lips, and tight lips; the motion characteristics include the temporal variation of three-dimensional coordinates, the rate of change of three-dimensional coordinates, the relative distance variation of key points in the same group, and the regional contour deformation coefficient; The action unit intensity dynamic quantization subunit is used to detect facial movements associated with stress and anger within the current preset time window, and record the duration and intensity dynamic feature vector of continuous activation of each emotion-associated facial movement. The intensity dynamic feature vector corresponding to the k-th emotion-related facial action is represented as follows: ;in, For the first The average intensity of facial movements associated with various emotions; For the first Peak intensity of facial movements associated with certain emotions; For the first The average slope of the rising edge of the intensity of facial movements associated with certain emotions; The multi-dimensional emotion contribution calculation subunit is used to calculate the contribution of each emotion-related facial action to the overall emotional state based on the dynamic feature vector of the duration and intensity of continuous activation of each emotion-related facial action, and the emotion association weight of each emotion-related facial action. The emotional arousal index generation subunit is used to accumulate the contribution of all activated facial movements within the current preset time window to generate the driver's emotional arousal index. .

[0009] Preferably, the multi-dimensional emotion contribution calculation subunit calculates the contribution of each emotion-related facial movement to the overall emotional state, including: ; in, For the first The contribution of facial movements associated with emotions to the overall emotional state; For the first Preset weighting coefficients for facial movements associated with emotions; For the first Emotional association weights for facial movements related to emotions; For the first The duration of continuous activation of facial movements associated with certain emotions; , , These are the feature weight coefficients corresponding to the average intensity, peak intensity, and average slope of intensity rise, respectively. , , These are the feature activation thresholds corresponding to average intensity, peak intensity, and average slope of intensity increase, respectively. It is a piecewise nonlinear saturation function used to limit the excessive contribution of a single feature.

[0010] Preferably, the environmental risk perception and prediction module includes: The target detection and tracking unit is used to detect and identify dynamic targets and static traffic elements in the vehicle's surrounding environment in real time based on the second temporal features in the fused perception data and through a preset target detection algorithm, and to perform cross-frame trajectory association and continuous tracking of each dynamic target. The risk situation modeling unit is used to calculate the instantaneous collision risk value of each dynamic target based on the identified static traffic elements and the historical trajectory, current motion state and relative position relationship of each dynamic target with respect to the vehicle, and to construct a risk distribution map of the current moment centered on the vehicle. The short-term prediction unit is used to predict the position and status of each dynamic target within a preset time period in the future, based on the current risk distribution map and the motion model of each dynamic target, and to generate a current environmental risk situation map.

[0011] Preferably, the risk-coupled assessment and decision-making module includes: The coupling risk calculation unit is used to calculate the coupling risk value of the corresponding potential risk event in the current driver state based on each potential risk event identified in the current environmental risk situation map, combined with all parameters of the current driver's personalized driving profile, through a nonlinear coupling model. The comprehensive risk assessment unit is used to conduct a comprehensive risk assessment of all potential risk events based on the coupled risk values ​​of all potential risk events and in combination with the temporal and spatial urgency of the potential risk events. The early warning decision generation unit is used to generate the corresponding early warning intervention decision based on the comprehensive risk level and type of all potential risk events, combined with the current driver's personalized driving profile parameters.

[0012] Preferably, the multimodal adaptive early warning module includes: The warning strategy library is used to store at least one preset warning strategy template. Each preset warning strategy template is associated with a specific combination of comprehensive risk level, potential risk event type and current driver personalized driving profile parameters. The strategy selection and optimization unit is used to match the baseline warning strategy from the warning strategy library based on the warning intervention decision, and to perform personalized optimization of the baseline warning strategy by combining all parameters of the current driver's personalized driving profile and the driver's historical warning feedback data, and generate the final warning instruction.

[0013] Compared with the prior art, the present invention has the following beneficial effects: This invention ensures the comprehensiveness of risk assessment and the integrity of data support by synchronously fusing multi-source heterogeneous data on driver status, external environment, and vehicle dynamics. It innovatively constructs a personalized driving profile that integrates real-time status and long-term habits, enabling early warning strategies to accurately adapt to the individual characteristics of different drivers. It establishes a coupled assessment model of driver status and environmental risk, rather than considering a single dimension of risk, thus achieving accurate quantification of comprehensive risk levels. This avoids the missed and false alarms caused by the fragmented risk assessment in existing systems. The complementary verification of multi-source data enhances the system's anti-interference capability, maintaining stable early warning performance even in complex scenarios such as low light and target occlusion. Attached Figure Description

[0014] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the vehicle safety warning system of the present invention. Detailed Implementation

[0015] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0016] Furthermore, in this invention, the use of terms such as "first" and "second" is for descriptive purposes only and does not specifically refer to any order or sequence, nor is it intended to limit the invention. They are merely used to distinguish components or operations described using the same technical terms and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions and features of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If a combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0017] The present invention provides the following embodiments. Example 1 This invention provides a driving safety warning system based on a video surveillance vehicle terminal, such as... Figure 1 As shown, it includes: The data acquisition and fusion module is used to collect and synchronously fuse multi-source heterogeneous data streams from the in-vehicle driver monitoring camera, the external environment perception camera and the vehicle CAN bus in real time, and generate fused perception data corresponding to the current driver. The personalized driver profile generation module is used to construct short-term state vectors and long-term habit vectors based on the fused perception data corresponding to the current driver, and to fuse the short-term state vectors and long-term habit vectors to generate a personalized driving profile of the current driver. The environmental risk perception and prediction module is used to identify dynamic targets and static traffic elements around vehicles based on fused perception data, predict the risk evolution trend in the short time domain through time-series trajectory modeling, and generate a current environmental risk situation map. The risk coupling assessment and decision-making module is used to assess the comprehensive risk level under the interaction between driver status and environmental risk based on the current driver's personalized driving profile and the current environmental risk situation map through a nonlinear coupling model, and generate the corresponding early warning and intervention decision. The multimodal adaptive warning module is used to generate the final warning command based on the current warning intervention decision, the current driver's personalized driving profile, and the driver's historical warning feedback data.

[0018] In this embodiment, the driver monitoring camera includes at least one near-infrared camera deployed in the cockpit and facing the driver's face, for stably capturing the driver's biometrics under low light and backlight conditions.

[0019] In this embodiment, the external environment perception camera includes at least four wide-angle cameras deployed around the vehicle to form a surround view.

[0020] In this embodiment, the short-time state vector refers to a fatigue index of the driver that is constructed in real time based on the fused perception data of the current driver. Distraction index Emotional Excitement Index and control standardization index A four-dimensional vector is used to represent the driver's instantaneous state during the current driving period.

[0021] In this embodiment, the long-term habit vector refers to a three-dimensional vector that represents the driver's long-term driving behavior preferences, based on the fused perception data of experienced drivers within their historical driving cycles and output by a trained dedicated neural network model. It includes the following distance habit value, lane change style type, and risk response habit reaction time; new drivers are loaded with default initial values.

[0022] In this embodiment, the current driver's personalized driving profile refers to a comprehensive data model that can fully and accurately describe the individual driving characteristics of a driver by fusing a short-term state vector that represents the driver's immediate state with a long-term habit vector that represents the driver's long-term driving preferences.

[0023] In this embodiment, dynamic targets around the vehicle include moving motor vehicles (cars, trucks, buses, etc.), pedestrians, non-motorized vehicles (bicycles, electric bicycles, tricycles, etc.), temporary moving obstacles (construction equipment, fallen objects, etc.), cyclists, and other traffic participants with mobility attributes.

[0024] In this embodiment, the static traffic elements around the vehicle include: lane lines (solid lines, dashed lines, double yellow lines, etc.), traffic signs (speed limit signs, prohibition signs, warning signs, etc.), traffic markings (stop lines, zebra crossings, directional arrows, etc.), road guardrails, curbs, traffic lights, bridges, tunnels, construction barriers, and fixed road infrastructure.

[0025] In this embodiment, the environmental risk situation map is a visualized data model centered on the vehicle, integrating the location and motion state of dynamic targets around the vehicle, the spatial distribution of static traffic elements, the instantaneous collision risk value of each dynamic target, the risk evolution trend within a preset time period, and potential collision points.

[0026] In this embodiment, the driver's historical warning feedback data includes the response speed of historical warning events, i.e., the time interval from the triggering of the warning to the driver taking action; the type of response behavior, such as deceleration, steering, ignoring, turning off the warning, and no action; the fluctuation value of the control normativity index; the acceptance of each warning mode; the pattern of repeated warning feedback in similar risk scenarios; and the risk avoidance effect after warning intervention.

[0027] The beneficial effects of the above technical solution are as follows: This invention ensures the comprehensiveness of risk assessment and the integrity of data support by synchronously integrating multi-source heterogeneous data on driver status, external environment, and vehicle dynamics. It innovatively constructs a personalized driving profile that integrates real-time status and long-term habits, enabling the early warning strategy to accurately adapt to the individual characteristics of different drivers. It establishes a coupled assessment model of driver status and environmental risk, rather than considering a single dimension of risk, thus achieving accurate quantification of comprehensive risk level. It avoids the missed and false alarms caused by the fragmented risk assessment in existing systems. The complementary verification of multi-source data improves the system's anti-interference capability, and it can maintain stable early warning performance even in complex scenarios such as low light and target occlusion.

[0028] Example 2 Based on Example 1, the data acquisition and fusion module includes: The driver identification unit is used to identify the unique identity of the current driver based on the facial image obtained by the in-vehicle driver monitoring camera and to identify whether the current driver is an experienced driver or a new driver through a preset face recognition model. When the driver is identified as a new driver, a personalized driving profile file corresponding to the identity is created. The driver data acquisition unit is used to acquire an original video stream containing the current driver's facial image, eye features, head posture, and gestures based on the in-vehicle driver monitoring camera, and to extract first temporal features based on the original video stream. The first temporal features include the eye opening sequence, gaze direction vector sequence, head three-dimensional Euler angle sequence, facial key point motion sequence, and hand key point coordinate sequence extracted from the original video stream and arranged in chronological order. The environmental data acquisition unit is used to acquire a second original video stream containing dynamic targets and static traffic elements around the vehicle based on the external environment perception camera, and to extract a second temporal feature based on the second original video stream; wherein, the second temporal feature includes a sequence of bounding boxes, a sequence of target category labels, and a sequence of target motion state vectors of each detected target extracted from the second original video stream and arranged in chronological order. The vehicle data acquisition unit is used to acquire in real time a sequence of vehicle dynamic parameters, including vehicle speed, longitudinal acceleration, steering wheel angle, yaw rate and braking status, through the vehicle CAN bus interface. The spatiotemporal synchronization fusion unit is used to perform timestamp alignment, coordinate system unification, and data normalization on the first time-series features, the second time-series features, and the vehicle dynamic parameter sequence to generate fused perception data.

[0029] In this embodiment, the fused sensing data is a timestamp-aligned multidimensional data matrix, with its rows corresponding to various features at the same time and its columns corresponding to time series. Missing data in the matrix is ​​filled in using a Kalman filter algorithm to ensure data integrity.

[0030] The beneficial effects of the above technical solution are as follows: By collecting three core types of data—driver, environment, and vehicle—this invention ensures the relevance and comprehensiveness of data collection. The Kalman filter algorithm is used to complete missing data, avoiding deviations in subsequent profile construction and risk assessment caused by missing data, thus improving data integrity. Through spatiotemporal synchronization processing, timestamp alignment, and coordinate system unification, the effective fusion of multi-source data is achieved, ensuring the consistency of different types of data in time and space dimensions, and providing high-quality data support for subsequent modules.

[0031] Example 3 Based on Example 2, the personalized driver profile generation module includes: The short-term state analysis submodule is used to acquire, in real time, the driver's fatigue index, attention distraction index, emotional agitation index, and operational compliance index based on the fused perception data corresponding to the current driver, and to determine the driver's fatigue index. Distraction index Emotional Excitement Index and control standardization index Construct short-time state vectors; The long-term habit learning submodule, when the driver identification unit identifies the driver as an experienced driver, inputs the fused perception data of the current driver within the historical driving cycle into the trained following distance output model, lane change style evaluation model, and risk response time output model, respectively, to obtain the driver's following distance habit value, lane change style type, and risk response habit reaction time. Based on the following distance habit value, lane change style type, and risk response habit reaction time, the module constructs the driver's long-term habit vector; among which, the lane change style type includes aggressive, steady, and decisive. When the driver identification unit identifies the driver as a new driver, the default long-term habit vector is loaded as the initial value. The driving profile generation submodule is used to fuse the short-term state vector obtained by the short-term state analysis submodule and the long-term habit vector obtained by the long-term habit learning submodule to generate a personalized driving profile of the current driver.

[0032] In this embodiment, the trained following distance output model is obtained by training a neural network model with a large amount of fused perception data of historical drivers as input and the corresponding driver following distance habit value as output.

[0033] In this embodiment, the trained lane change style evaluation model is obtained by training a neural network model with a large amount of fused perception data of historical drivers as input and the corresponding driver lane change style type as output.

[0034] In this embodiment, the trained risk reaction time output model is obtained by training a neural network model with a large amount of fused perception data of historical drivers as input and the corresponding driver risk response habit reaction time as output.

[0035] The beneficial effects of the above technical solution are as follows: This invention divides the driver's state into short-term vectors of immediate state and long-term vectors of long-term habits, solving the problem that existing systems only focus on a single state dimension and the profile description is incomplete. Different long-term habit vector construction methods are adopted for experienced drivers and new drivers. Experienced drivers learn personalized habits based on historical data, while new drivers are loaded with default initial values, ensuring the flexibility and applicability of profile construction. Long-term driving habits are extracted through a dedicated neural network model, improving the accuracy of habit feature extraction, so that the profile can truly reflect the driver's driving preferences. The fused personalized driving profile can dynamically adapt to state changes during driving, preserving the driver's long-term behavioral characteristics and updating immediate state parameters in real time. This provides accurate individual feature support for subsequent coupled risk assessment and adaptive warning, enabling warning strategies to vary from person to person, improving the driver's acceptance of warnings and execution efficiency.

[0036] Example 4 Based on Example 3, the short-time state analysis submodule includes: The eye state monitoring unit is used to calculate the driver's fatigue index within the current preset time window based on the current driver's eye opening and closing sequence, head three-dimensional Euler angle sequence, and facial key point motion sequence from the fused perception data. ; The gaze direction analysis unit is used to calculate the driver's attention distraction index within the current preset time window based on the current driver gaze direction vector sequence and hand key point coordinate sequence from the fused perception data. ; The facial emotion recognition unit is used to calculate the driver's emotional arousal index within a current preset time window based on the motion sequence of key facial points in the fused perception data. ; The driving behavior evaluation unit is used to calculate the driver's driving compliance index within the current preset time window based on the sequence of vehicle dynamic parameters obtained from the vehicle bus. .

[0037] In this embodiment, the current preset time window refers to the backtracking length used at the current analysis moment. The most recent continuous period of time.

[0038] In this embodiment, the driver's fatigue index is calculated within the current preset time window as follows: Calculate the proportion of time the driver's eyes are closed within the current preset time window: ;in, This represents the proportion of time the driver's eyes are closed within the current preset time window. This refers to the cumulative time within the current preset time window during which the eyelid closure of the driver's eyes exceeds a preset threshold, based on statistics of eye opening and closing sequence. This is the value for the current preset time window; The number of times the driver yawns within the current preset time window is statistically analyzed based on facial key point motion sequences. The number of involuntary head nods was counted based on the three-dimensional Euler angle sequence of the head. ; Driver fatigue index within the current preset time window: ;in, , and The weights are respectively the proportion of time the driver's eyes are closed, the number of times the driver yawns, and the number of times the driver involuntarily nods. The calculation only uses the numerical values ​​of each parameter and does not require the physical units on both sides of the equal sign to be consistent.

[0039] In this embodiment, the driver's attention distraction index within the current preset time window is calculated as follows: Calculate the percentage of the driver's line of sight deviating from the area of ​​the road ahead within the current preset time window: ;in, This refers to the cumulative duration within the current preset time window, based on statistics of the gaze direction vector sequence, for which the driver's gaze focus falls outside the area of ​​the road ahead. Based on the hand key point coordinate sequence, the number of events in which the driver's hands take off the steering wheel within the current preset time window is statistically analyzed. ; Driver distraction index within the current preset time window: ;in, and These represent the degree of risk associated with visual deviation and hand disengagement, respectively. The calculation only uses the numerical values ​​of each parameter and does not require the physical units on both sides of the equal sign to be consistent.

[0040] In this embodiment, the driver's operational standardization index within the current preset time window is calculated as follows: Within the current preset time window, obtain the steering wheel angle sequence and the absolute value sequence of vehicle longitudinal acceleration at m sampling times based on the vehicle dynamic parameter sequence; Calculate the fluctuations in steering wheel angle and longitudinal acceleration within the current preset time window: Fluctuation in steering wheel angle within the current preset time window: ; Fluctuation of longitudinal acceleration within the current preset time window: ;in, and These represent the steering wheel angle and the vehicle longitudinal acceleration at the j-th sampling moment within the current preset time window. and These are the average steering wheel angle and the average vehicle longitudinal acceleration within the current preset time window, respectively. Driver's operational compliance index within the current preset time window: ;in, This refers to the weighting coefficient corresponding to the fluctuation in the steering wheel angle. and These are preset reference values ​​representing the maximum permissible fluctuations in steering wheel angle and longitudinal acceleration under normal driving conditions. The calculation only uses the numerical values ​​of each parameter and does not require the physical units on both sides of the equal sign to be consistent.

[0041] The beneficial effects of the above technical solution are as follows: Each index of the present invention is calculated based on multiple sets of feature data, such as FI combined with three types of features: eyes, face, and head. This avoids the errors caused by the single feature evaluation of the state in the existing system, improves the accuracy of state recognition, and adopts a dimensionless numerical calculation method, which breaks through the unit limitations of different types of parameters such as time, number, and angle, simplifies the calculation logic. Through the dynamic setting of the current preset time window, the state index is updated in real time, which can capture the driver's state changes in a timely manner. The collaborative calculation of the four indices can not only comprehensively represent the driver's real-time state, but also improve the reliability of state evaluation through the complementary verification of different indices. For example, when FI and DI rise at the same time, the system can determine that the driver is in a high-risk state, thereby strengthening the early warning intervention. This solves the problems of incomplete single-state evaluation and delayed risk judgment in the existing system, and provides accurate real-time state support for subsequent risk coupling evaluation.

[0042] Example 5 Based on Example 4, the facial emotion recognition unit includes: The emotion-related key point grouping and parsing subunit is used to divide key point groups into two core emotions, driving stress and anger, based on the facial key point motion sequence in the fused perception data and extract the motion features of the corresponding key points. The key point groups include stress emotion association group and anger emotion association group. The stress-related emotion group includes four key points on the inner and outer sides of both eyebrows, and six key points on the upper and lower edges of both eyelids and the corners of the eyes; the anger-related emotion group includes four key points on the inner and outer sides of both eyebrows, and four key points on the upper and lower lips; the corresponding facial movements for the stress-related emotion group include drooping eyebrows, tight eyelids, abnormal blinking frequency, and stiff eye muscles; the corresponding facial movements for the anger-related emotion group include furrowed eyebrows, retracted corners of the mouth, pursed lips, and tight lips; the motion characteristics include the temporal variation of three-dimensional coordinates, the rate of change of three-dimensional coordinates, the relative distance variation of key points in the same group, and the regional contour deformation coefficient; The action unit intensity dynamic quantization subunit is used to detect facial movements associated with stress and anger within the current preset time window, and record the duration and intensity dynamic feature vector of continuous activation of each emotion-associated facial movement. The intensity dynamic feature vector corresponding to the k-th emotion-related facial action is represented as follows: ;in, For the first The average intensity of facial movements associated with various emotions; For the first Peak intensity of facial movements associated with certain emotions; For the first The average slope of the rising edge of the intensity of facial movements associated with certain emotions; The multi-dimensional emotion contribution calculation subunit is used to calculate the contribution of each emotion-related facial action to the overall emotional state based on the dynamic feature vector of the duration and intensity of continuous activation of each emotion-related facial action, and the emotion association weight of each emotion-related facial action. The emotional arousal index generation subunit is used to accumulate the contribution of all activated facial movements within the current preset time window to generate the driver's emotional arousal index E. .

[0043] In this embodiment, the facial key point motion sequence refers to one of the first temporal features extracted from the original video stream captured by the driver monitoring camera. After accurately locating all target key points of the stress emotion association group and the anger emotion association group through the convolutional neural network key point detection model, the sequence of continuous frame three-dimensional coordinate data is formed in chronological order. Its core content is the frame number, key point number and the temporal matrix of three-dimensional coordinates in the horizontal, vertical and depth directions, which can completely record the position changes of each key point over time.

[0044] In this embodiment, the stress-emotion association group comprises 10 key points: There are 4 eyebrow key points, 2 on the inner and 2 on the outer sides of both eyebrows, which are recorded as M1-M4. M1 / M2 are the inner / outer sides of the left eyebrow, and M3 / M4 are the inner / outer sides of the right eyebrow. There are 6 key points on each eyelid: 3 on the upper edge, 3 on the lower edge, and 3 at the corner of the eye. These are marked as key eye points E1-E6. E1-E3 represent the upper / lower edge / corner of the left eyelid, and E4-E6 represent the upper / lower edge / corner of the right eyelid. The anger association group has 8 key points: There are 2 points on the inner and 2 on the outer sides of both eyebrows, for a total of 4 points, which are marked as eyebrow key points M1-M4. M1 / M2 are the inner / outer sides of the left eyebrow, and M3 / M4 are the inner / outer sides of the right eyebrow. There are 2 points on the upper edge of the upper lip and 2 points on the lower edge of the lower lip, for a total of 4 points, which are denoted as L1-L4. L1 / L2 are the left / right edges of the upper lip, and L3 / L4 are the left / right edges of the lower lip.

[0045] In this embodiment, the three-dimensional coordinate temporal change is the maximum displacement offset of each key point in the target frame relative to the starting frame of the preset time window, reflecting the magnitude of the motion displacement. The specific extraction steps include: taking the first frame of the preset time window as the reference frame, reading the coordinates of the target key point, then reading the coordinates of the key point in the i-th frame within the window, then calculating the three-directional coordinate offset, and finally taking the maximum value of the absolute value of the three offsets, which is the original coordinate temporal change of the key point in the frame. The temporal change rate of three-dimensional coordinates is the ratio of the offset change of the same key point between two adjacent frames to the time interval, reflecting the motion activation rate. The extraction steps include: obtaining the original temporal change of coordinates between two adjacent frames, then calculating the time interval according to the camera sampling frame rate, and then calculating the quotient of the offset difference between the two frames divided by the time interval. The relative distance change between key points in the same group is the difference between the target frame spatial distance of key point pairs within the same emotion association group and the average distance within the window, reflecting the amplitude of local muscle contraction and relaxation. The extraction steps include: determining key point pairs according to emotion-related actions, such as eyebrow drooping corresponding to M1-M2 and M3-M4; calculating the spatial distance of key point pairs in the i-th frame, such as M1-M2, using the Euclidean distance formula; statistically analyzing the distance of the key point pair in all frames within the window and taking the arithmetic mean as the average distance within the window; and calculating the difference between the target frame spatial distance of key point pairs within the emotion association group and the average distance within the window. The region contour deformation coefficient is the change value of the contour curvature of the emotion representation region (eyebrow area, eye area, lip area) fitted based on the same group of key points. It is dimensionless and normalized to [0,1]. It reflects the degree of muscle tension and relaxation. The extraction steps are as follows: use the same group of key points, such as eye area E1-E6, to fit the closed contour of the corresponding region. Take a number of contour points at equal intervals on the contour, calculate the curvature of each contour point, take the arithmetic mean of the curvature of all contour points to obtain the original average curvature of the frame, and normalize the original average curvature to obtain the region contour deformation coefficient.

[0046] In this embodiment, the emotion-related facial movements corresponding to the stress emotion association group include: Eyebrow drooping (k=1): Based on the temporal change of the three-dimensional coordinates of key eyebrow points M1-M4 and the change of the relative distance between key points in the same group (i.e. the change of the distance between the inner and outer sides of the left eyebrow M1-M2 and the inner and outer sides of the right eyebrow M3-M4), it reflects the overall downward movement of the eyebrow caused by pressure. Eyelid tightness (k=2): Based on the relative distance change between key points in the same group of eye key points E1-E6 (i.e., the distance change from the upper edge to the lower edge of the left eyelid E1-E2, and from the upper edge to the lower edge of the right eyelid E4-E5) and the regional contour deformation coefficient (change in the curvature of the eye region contour), it reflects the eyelid muscle contraction and tightness caused by pressure. Abnormal blinking frequency (k=3): Based on the temporal change rate of the three-dimensional coordinates of key eye points E1-E6 (the change rate of the distance from the lower edge of the left eyelid to the left corner of the eye E2-E3, and from the lower edge of the right eyelid to the right corner of the eye E5-E6), the number of blinks per unit time is counted (the normal range is 15-20 times / minute, and it is judged as abnormal if it is higher than 25 times / minute or lower than 10 times / minute). Eye muscle stiffness (k=4): Based on the regional contour deformation coefficient (eye area contour curvature) of eye key points E1-E6 and the standard deviation of the three-dimensional coordinate temporal change (characterizing the dispersion of the coordinate offset of each eye key point), when the eye area contour curvature ≥0.7 (high curvature, muscle tension) and the standard deviation of the three-dimensional coordinate temporal change ≤0.1 (stable coordinate change, no flexible movement), it is judged as eye muscle stiffness.

[0047] In this embodiment, the emotion-related facial actions corresponding to the anger emotion association group include: Eyebrow wrinkling (k=5): Based on the relative distance change between key points in the same group of eyebrow key points M1-M4 (the distance change from the inner side of the left eyebrow to the inner side of the right eyebrow M1-M3, and from the outer side of the left eyebrow to the outer side of the right eyebrow M2-M4) and the regional contour deformation coefficient (the change in the curvature of the eyebrow area), it reflects the movement of the eyebrows contracting and converging towards the middle caused by anger. Mouth corner retraction (k=6): Based on the three-dimensional coordinate temporal change of key lip points L1-L4 (X-axis horizontal offset of upper lip left edge L1, upper lip right edge L2, lower lip left edge L3, lower lip right edge L4) and the three-dimensional coordinate temporal change rate (rate of change of horizontal displacement), it reflects the pulling and retracting movement of the corners of the mouth to both sides and back caused by anger. Lip pursing (k=7): Based on the relative distance change between key points in the same group of lip key points L1-L4 (the distance change between the left edge of the upper lip and the left edge of the lower lip L1-L3, and between the right edge of the upper lip and the right edge of the lower lip L2-L4), and the regional contour deformation coefficient (the change in the curvature of the lip region contour), it reflects the movement of the upper and lower lips tightening and pressing together when angry. Tight lips (k=8): Based on the regional contour deformation coefficient (lip region contour curvature) of key lip points L1-L4 and the standard deviation of the three-dimensional coordinate temporal change (dispersion of the coordinate offset of each key lip point), when the lip region contour curvature ≥ 0.8 (high curvature, tight lips) and the standard deviation of the three-dimensional coordinate temporal change ≤ 0.08, it is judged as tight lips.

[0048] In this embodiment, the intensity dynamic feature vector corresponding to the k-th emotion-related facial action is represented as follows: ;in, For the first The average intensity of facial movements associated with various emotions reflects the degree of sustained stability of movement activation; the higher the value, the more stable the sustained activation intensity. For the first The peak intensity of facial movements associated with emotions reflects the intensity of movement activation; the higher the value, the higher the instantaneous intensity of the movement during the activation period. For the first The average slope of the intensity rise of facial movements associated with emotions reflects the swiftness of movement activation; the larger the value, the more rapid the process from activation to peak. in, The calculation formula is: in, For the first The number of effective sampled frames within the period of facial action activation associated with emotion is determined by the duration of continuous action activation. With camera sampling frame rate The conversion yields, i.e. ; For the first The first type of facial movement associated with emotion The instantaneous intensity of a frame is obtained by weighted fusion of normalized motion features corresponding to emotion-related facial movements. The specific calculation formula is as follows: in, For the first The first type of facial movement associated with emotion The normalized value of the temporal variation of the three-dimensional coordinates of the frame, with a value range of [value range missing]. It is obtained by normalizing the original coordinate offset according to a preset range. For example, if the original offset is 3 mm, the preset range is... hour, ; For the first The first type of facial movement associated with emotion The normalized value of the temporal rate of change of the three-dimensional coordinates of the frame, with a value range of [value missing]. The result is obtained by normalizing the original rate of change according to a preset range. For example, if the original rate of change is 6 mm / s, the preset range is... hour, ; For the first The first type of facial movement associated with emotion Normalized values ​​of relative distance changes between keypoints in the same frame (value range) The distance variation is obtained by normalizing the original distance variation according to a preset range. For example, if the original variation is 1.5 mm, it is normalized according to the preset range. , ; For the first The first type of facial movement associated with emotion The region contour deformation coefficient of the frame, since it has been normalized to , directly used as input; These are weighted coefficients corresponding to the temporal change of 3D coordinates, the rate of change of 3D coordinates, the relative distance change of key points in the same group, and the regional contour deformation coefficient, respectively. The sum of these coefficients is 1. The system is adapted according to the type of facial movement associated with emotion; for example, for a drooping eyebrow movement, emphasis is placed on coordinate and distance changes. , , The action of retracting the corner of the mouth focuses on the change in coordinates and the rate of change, taking... , , ; The calculation formula is: in: Indicates the first The first frame to the second frame during the period of facial motion activation associated with emotions. The instantaneous intensity of the frame is used to filter out the strongest state of action activation by taking the maximum value; The calculation formula is: in: For the first The activation initiation frame intensity of emotion-associated facial movements, i.e., the first time the activation threshold is reached. Frame intensity at time, such as: activation threshold When the default value is 0.3, the action is activated starting from the 3rd frame, and the intensity of that frame is... ,but ; From the activation start frame to the intensity peak The actual time interval of the corresponding frame.

[0049] In this embodiment, the formula for calculating the driver's emotional agitation index is: ; in, This is the driver's emotional agitation index, which is the sum of the contributions of all emotion-related facial movements. The higher the value, the more agitated the driver is. The state is determined to be one of high emotional agitation, requiring an alert to be triggered. To identify facial movements associated with 8 emotions (4 corresponding to stress). ; Four types of anger ) contribution Summation, The calculation only uses the numerical values ​​of each parameter and does not require the physical units on both sides of the equal sign to be consistent.

[0050] The beneficial effects of the above technical solution are as follows: This invention classifies core emotions into two categories: stress and anger. It specifically divides key point groups and associated facial movements, solving the problems of ambiguous emotion recognition categories and poor targeting in existing systems. Through the extraction and dynamic quantification of intensity of multi-dimensional motion feature coordinate changes, distance changes, contour deformation, etc., it achieves accurate description of facial movements, avoiding the errors caused by the single feature recognition of facial movements in existing systems. Based on the contribution accumulation calculation of EI, it can quantify the degree of influence of different facial movements on emotions, improving the reliability of the emotional arousal index. The definition of key point groups and facial movements enables it to accurately capture subtle emotional changes during driving, such as eyelid tension caused by stress and brow furrowing caused by anger. Even when the driver does not show obvious emotional expression, such as shouting or crying, it can still accurately identify potential emotional risks, solving the problems of insufficient sensitivity of emotion recognition and easy omission of potential emotional risks in existing systems in driving scenarios, and providing accurate support for the coupled assessment of emotion-related risks.

[0051] Example 6 Based on Example 5, the multi-dimensional emotion contribution calculation subunit calculates the contribution of each emotion-related facial action to the overall emotional state, including: ; in, For the first The contribution of facial movements associated with emotions to the overall emotional state; For the first Preset weighting coefficients for facial movements associated with emotions; For the first Emotional association weights for facial movements related to emotions; For the first The duration of continuous activation of facial movements associated with certain emotions; , , These are the feature weight coefficients corresponding to the average intensity, peak intensity, and average slope of intensity rise, respectively. , , These are the feature activation thresholds corresponding to average intensity, peak intensity, and average slope of intensity increase, respectively. It is a piecewise nonlinear saturation function used to limit the excessive contribution of a single feature.

[0052] In this embodiment, stress-related emotions are associated with facial movements. Anger-related facial movements .

[0053] In this embodiment, , , Based on a large amount of driving scenario data, they were respectively calibrated as That is, the average strength needs to be ≥0.4 to make an effective contribution. That is, the peak intensity needs to be ≥0.6 to make an effective contribution. That is, the rising slope must be ≥0.2 to generate a valid contribution; if the threshold is not reached, the contribution of the corresponding feature is 0.

[0054] The beneficial effects of the above technical solution are as follows: This invention introduces emotion-related weights to distinguish the degree of emotional influence of stress and anger facial movements, making the contribution calculation more consistent with the emotional risk characteristics in driving scenarios. Setting feature activation thresholds avoids interference from low-intensity, meaningless facial movements on emotion assessment, improving the reliability of the calculation results. Using a piecewise nonlinear saturation function limits the excessive contribution of a single feature, ensuring a balanced influence of each feature on the contribution, avoiding the emotion assessment bias caused by a single feature dominance in existing systems. Through the synergistic effect of the base weight coefficient βk with emotion-related weights, duration, and intensity features, the contribution calculation can accurately reflect the emotional correlation, stability, and intensity of facial movements. For example, high-intensity, long-duration anger facial movements will receive a higher contribution, while low-intensity, short-duration stress facial movements will have a lower contribution. This differentiated calculation method allows the emotional agitation index EI to truly reflect the driver's emotional state, providing accurate emotional quantification support for subsequent risk coupling assessment.

[0055] Example 7 Based on Example 6, the environmental risk perception and prediction module includes: The target detection and tracking unit is used to detect and identify dynamic targets and static traffic elements in the vehicle's surrounding environment in real time based on the second temporal features in the fused perception data and through a preset target detection algorithm, and to perform cross-frame trajectory association and continuous tracking of each dynamic target. The risk situation modeling unit is used to calculate the instantaneous collision risk value of each dynamic target based on the identified static traffic elements and the historical trajectory, current motion state and relative position relationship of each dynamic target with respect to the vehicle, and to construct a risk distribution map of the current moment centered on the vehicle. The short-term prediction unit is used to predict the position and status of each dynamic target within a preset time period in the future, based on the current risk distribution map and the motion model of each dynamic target, and to generate a current environmental risk situation map.

[0056] In this embodiment, cross-frame trajectory association and continuous tracking of each dynamic target are achieved through the following method: a fusion scheme of Kalman filter prediction, Hungarian algorithm data association, and multi-feature matching completion is adopted. The specific steps are as follows: 1) Based on the position, velocity, and acceleration parameters of the dynamic target in the current frame, the Kalman filter algorithm is used to predict the possible position and state confidence of the target in the next frame; 2) Calculate the Euclidean distance between the predicted location and the target location detected in the next frame, and construct the cost matrix; 3) Solve for the optimal solution of the cost matrix using the Hungarian algorithm to achieve initial correlation of targets across frames; 4) For scenarios involving occlusion or loss of detection, supplement the multi-dimensional matching of target appearance features (color, outline, size) and motion features (speed consistency, direction continuity) to ensure uninterrupted tracking trajectory and tracking accuracy ≥95%.

[0057] In this embodiment, the risk distribution map at the current moment is divided into red high-risk areas, yellow medium-risk areas, and blue low-risk areas according to risk level. The red high-risk areas are areas with an expected collision time of ≤3s, the yellow medium-risk areas are areas with an expected collision time of ≤15s and the blue low-risk areas are areas with an expected collision time of >15s.

[0058] In this embodiment, the motion models of each dynamic target include uniform velocity models, uniform acceleration models, etc.

[0059] In this embodiment, based on the current risk distribution map and the motion model of each dynamic target, the prediction of the position and state of each dynamic target within a preset future time period is achieved in the following way: 1) Adapt the corresponding motion model to different types of dynamic targets, such as using a uniform acceleration model for motor vehicles, a random motion model for pedestrians, and a uniform speed + steering model for non-motorized vehicles; 2) Input the current dynamic target's position, motion state (velocity, acceleration, direction) and instantaneous collision risk value of each dynamic target into the corresponding motion model. 3) The LSTM time series prediction network is used to optimize the model output results, and historical trajectory trends and environmental constraints (such as lane line restrictions and roadside obstructions) are integrated. 4) Output the coordinate trajectory, speed change trend and relative position relationship with the vehicle of each target within the next 3s-15s (dynamically adjusted according to vehicle speed, the higher the vehicle speed, the longer the prediction time), and complete the generation of the environmental risk situation map.

[0060] The beneficial effects of the above technical solution are as follows: This invention adopts a fusion-based dynamic target tracking scheme, which solves the problem of tracking interruption in existing systems under target occlusion and detection loss scenarios, ensuring the continuity and accuracy of trajectory tracking. The risk distribution map divided by risk level intuitively presents the spatial distribution of risks inside and outside the vehicle, providing clear and visualized data support for subsequent risk assessment. Differentiated motion models are adapted for different types of dynamic targets, improving the pertinence and accuracy of future state prediction. The combination of LSTM temporal prediction network and motion model not only realizes accurate prediction of the future position and state of dynamic targets, but also predicts the risk evolution trend in advance, such as potential collision points and risk level escalation trends, enabling the system to provide early warnings. Compared with the real-time risk alarm of existing systems, more driver reaction time is reserved, improving the effectiveness of early warning and solving the pain points of delayed early warning and insufficient driver response in existing systems.

[0061] Example 8 Based on Example 3, the risk coupling assessment and decision-making module includes: The coupling risk calculation unit is used to calculate the coupling risk value of each potential risk event in the current driver state based on each potential risk event identified in the current environmental risk situation map, combined with all parameters of the current driver's personalized driving profile, through a nonlinear coupling model. The comprehensive risk assessment unit is used to conduct a comprehensive risk assessment of all potential risk events based on the coupled risk values ​​of all potential risk events and in combination with the temporal and spatial urgency of the potential risk events. The early warning decision generation unit is used to generate the corresponding early warning intervention decision based on the comprehensive risk level and type of all potential risk events, combined with the current driver's personalized driving profile parameters.

[0062] In this embodiment, a potential risk event refers to a specific scenario identified in the environmental risk situation map that may lead to a collision between the vehicle and a dynamic target, or cause a driving hazard due to the influence of static traffic elements.

[0063] In this embodiment, the coupling risk value is defined as a dimensionless numerical value representing the likelihood that a potential risk event will cause an actual danger.

[0064] In this embodiment, based on each potential risk event identified in the current environmental risk situation map, and combined with all parameters of the current driver's personalized driving profile, the coupling risk value that the potential risk event may cause under the current driver's state is calculated using a nonlinear coupling model. The specific calculation formula is as follows: in, The inherent environmental risk value is obtained from the environmental risk situation map. They are respectively , , , The weighting coefficients for the coupled risk values ​​corresponding to following distance habits, lane change style types, and risk response habit reaction times are calculated. This is the customary following distance value; For lane-changing style types, the scores are: Aggressive = 1.2, Moderate = 1.0, Decisive = 1.1; For drivers' risk response habits and reaction time; The standard reaction time is preset to 1.5 seconds; among which, The calculation only uses the numerical values ​​of each parameter and does not require the physical units on both sides of the equals sign to be consistent. The magnitude of the value directly represents the level of coupling risk. The larger the value, the higher the level of coupling risk. This value is used to quantify the relative degree of coupling risk under the interaction between driver state and environmental risk.

[0065] In this embodiment, the spatiotemporal urgency includes the distance to the vehicle and the estimated collision time.

[0066] In this embodiment, based on the coupled risk value of all potential risk events and combined with the spatiotemporal urgency of the potential risk events, a comprehensive level assessment is performed on all current potential risk events. That is, the product of the coupled risk value and the spatiotemporal urgency coefficient is used as the comprehensive risk value. Based on the magnitude of the comprehensive risk value, they are divided into four levels: No Risk Level I, Warning Level II, Attention Level III, Early Warning Level IV, and Emergency Level V. The spatiotemporal urgency coefficient is calculated as the reciprocal of the product of the distance to the vehicle and the expected collision time. The closer the distance to the vehicle and the shorter the expected collision time, the higher the urgency coefficient.

[0067] In this embodiment, the comprehensive risk level includes five levels: no risk (Level I, ...). ), Prompt Level (Level II, 0.3≤ Attention level (Level III, 0.5≤) ), Warning Level (Level IV, 0.7≤ Emergency Level (Level V) ).

[0068] In this embodiment, the types of potential risk events include sudden braking of the vehicle in front, slow driving of the vehicle in front, lane change of the vehicle in front, illegal lane change of vehicles on the side, pedestrians on the side, non-motorized vehicles crossing the road, vehicles approaching too closely from adjacent lanes, lane departure, speeding, following closely, construction barriers, road obstacles, changes in traffic lights, and risks at tunnel and bridge entrances.

[0069] In this embodiment, based on the comprehensive risk level and type of all potential risk events, and combined with the current driver's personalized driving profile parameters, the corresponding early warning intervention decision is generated. This includes determining the early warning intensity based on the comprehensive risk level, determining the early warning content based on the type of potential risk event (e.g., a pedestrian crossing corresponds to a deceleration and avoidance prompt, lane departure corresponds to a return-to-center direction guide), adjusting the early warning mode and triggering logic based on the personalized driving profile parameters (e.g., enhancing tactile warnings for high fatigue, increasing warning intensity for aggressive lane changing style, and triggering early warnings for long reaction times), and optimizing the early warning coordination method based on parameter combinations (e.g., for high distraction + sudden braking risk of the vehicle in front, a triple coordinated early warning of visual, auditory, and tactile senses is adopted).

[0070] The beneficial effects of the above technical solution are as follows: This invention breaks through the limitations of the single-dimensional risk assessment of existing systems, couples all the driver's profile parameters with environmental risks for calculation, accurately quantifies the impact of the driver's state on environmental risks, solves the problem of inaccurate assessment caused by the fragmented risk assessment of existing systems, introduces a spatiotemporal urgency coefficient, so that the comprehensive risk level can reflect the urgency of the risk, avoids the inappropriate decision-making caused by existing systems that only assess based on the size of the risk and ignore spatiotemporal characteristics, and generates targeted early warning and intervention decisions based on risk level and event type, ensuring the adaptability of the decision. The calculation of the coupled risk value not only integrates the driver's immediate state and long-term habits, but also highlights the impact of key parameters through weight allocation. For example, high fatigue FI will significantly increase the coupled risk value, so that the comprehensive risk level can accurately reflect the actual risk of a specific driver in a specific environment. The resulting early warning and intervention decisions can simultaneously adapt to the individual characteristics of the driver and the characteristics of the risk scenario, solve the problem of strong generality and poor targeting of early warning decisions in existing systems, and improve the effectiveness of early warning and intervention.

[0071] Example 9 Based on Example 1, the multimodal adaptive early warning module includes: The warning strategy library is used to store at least one preset warning strategy template. Each preset warning strategy template is associated with a specific combination of comprehensive risk level, potential risk event type and current driver personalized driving profile parameters. The strategy selection and optimization unit is used to match the baseline warning strategy from the warning strategy library based on the warning intervention decision, and to perform personalized optimization of the baseline warning strategy by combining all parameters of the current driver's personalized driving profile and the driver's historical warning feedback data, and generate the final warning instruction.

[0072] In this embodiment, the preset warning strategy template is a structured configuration scheme designed around the dual dimensions of risk adaptation and driver adaptation. Its core includes basic correlation dimensions, warning modality combination configuration, triggering conditions and timing rules, personalized design of warning content, and parameter adaptation and adjustment rules. Among them, the basic association dimension achieves strong binding between templates and comprehensive risk levels, potential risk event types, and driver profile parameter combination tags to ensure targeting. The warning modality combination configuration clearly defines the icon type, flashing frequency, display position and color scheme of the visual modality, the voice type, initial volume, tone style and buzzer frequency of the auditory modality, the initial frequency of seat vibration parts and steering wheel vibration intensity of the tactile modality, and the activation status, braking force and throttle limit ratio of the vehicle control modality. The triggering conditions and timing rules define the spatial triggering conditions, temporal triggering conditions and driver status triggering conditions of the template. The warning content is personalized and pre-set with exclusive voice prompt text, visual guidance content and risk quantification display information for different risk types. The parameter adaptation and adjustment rules reserve dynamic adjustment logic linked with driver profile parameters, and clarify the parameter correction coefficients corresponding to each profile parameter and the adaptation strategies corresponding to different long-term habit vectors.

[0073] In this embodiment, the specific method for optimizing the strategy selection is as follows: 1) Based on Optimization: For example When fatigue is high, tactile warnings are enhanced, such as increasing the frequency of seat vibration, while auditory warnings are weakened to avoid interference. When fatigue occurs, a balanced warning system using both visual and tactile senses is employed. 2) Based on Optimization: For example When highly distracted, the visual warning icon flashes more frequently, and the interval between voice prompts is shortened. When the mind is distracted, provide targeted voice guidance to bring the gaze back to the center. 3) Based on Optimization: For example When a driver is highly emotionally agitated, a low-volume, calm voice warning should be used to avoid stimulating the driver. When a person is emotionally agitated, visual warnings should primarily use cool colors. 4) Based on Optimization: For example That is, when the driving standard is low, the linkage vehicle control module limits the maximum steering wheel angle and displays smooth steering visual guidance. When the operation is in compliance with regulations, only operation suggestion warnings are provided; 5) Optimization based on long-term habit vectors: For example, when the habitual following distance is less than the safety threshold, the warning trigger distance is extended; for drivers with aggressive lane-changing style, the warning intervention intensity is increased; when the habitual reaction time for risk response is greater than the standard reaction time, the warning lead time is increased.

[0074] The beneficial effects of the above technical solution are as follows: This invention breaks through the limitations of fixed warning modes and parameters in existing systems. Based on driver profile parameters, it performs targeted optimization, enabling the warning method to adapt to the individual characteristics of different drivers. For example, a calm voice warning is used for emotionally agitated drivers. Closed-loop optimization is performed by combining historical warning feedback data from drivers, allowing the warning strategy to gradually adapt to the driver's response preferences, thereby improving the driver's acceptance of the warning. Multimodal collaborative warning output ensures the effective transmission of warning information and avoids the problem of single-modal warnings being easily ignored. The personalized optimization logic not only considers the driver's immediate state and long-term habits, but also dynamically adapts to the combination of different risk scenarios and driver characteristics. This avoids invalid warnings interfering with the driver and ensures the strong effectiveness of warnings in high-risk scenarios. It solves the problems of rigid warning methods, easy interference with drivers, or insufficient warning strength in existing systems, thereby improving the execution efficiency of warnings and driving safety.

[0075] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A driving safety early warning system based on a video surveillance vehicle terminal, characterized in that: include: The data acquisition and fusion module is used to collect and synchronously fuse multi-source heterogeneous data streams from the in-vehicle driver monitoring camera, the external environment perception camera and the vehicle CAN bus in real time, and generate fused perception data corresponding to the current driver. The personalized driver profile generation module is used to construct short-term state vectors and long-term habit vectors based on the fused perception data corresponding to the current driver, and to fuse the short-term state vectors and long-term habit vectors to generate a personalized driving profile of the current driver. The environmental risk perception and prediction module is used to identify dynamic targets and static traffic elements around vehicles based on fused perception data, predict the risk evolution trend in the short time domain through time-series trajectory modeling, and generate a current environmental risk situation map. The risk coupling assessment and decision-making module is used to assess the comprehensive risk level under the interaction between driver status and environmental risk based on the current driver's personalized driving profile and the current environmental risk situation map through a nonlinear coupling model, and generate the corresponding early warning and intervention decision. The multimodal adaptive warning module is used to generate the final warning command based on the current warning intervention decision, the current driver's personalized driving profile, and the driver's historical warning feedback data.

2. The driving safety early warning system based on a video surveillance vehicle terminal according to claim 1, characterized in that: The data acquisition and fusion module includes: The driver identification unit is used to identify the unique identity of the current driver based on the facial image obtained by the in-vehicle driver monitoring camera and to identify whether the current driver is an experienced driver or a new driver through a preset face recognition model. When the driver is identified as a new driver, a personalized driving profile file corresponding to the identity is created. The driver data acquisition unit is used to acquire an original video stream containing the current driver's facial image, eye features, head posture, and gestures based on the in-vehicle driver monitoring camera, and to extract first temporal features based on the original video stream. The first temporal features include the eye opening sequence, gaze direction vector sequence, head three-dimensional Euler angle sequence, facial key point motion sequence, and hand key point coordinate sequence extracted from the original video stream and arranged in chronological order. The environmental data acquisition unit is used to acquire a second original video stream containing dynamic targets and static traffic elements around the vehicle based on the external environment perception camera, and to extract a second temporal feature based on the second original video stream; wherein, the second temporal feature includes a sequence of bounding boxes, a sequence of target category labels, and a sequence of target motion state vectors of each detected target extracted from the second original video stream and arranged in chronological order. The vehicle data acquisition unit is used to acquire in real time a sequence of vehicle dynamic parameters, including vehicle speed, longitudinal acceleration, steering wheel angle, yaw rate and braking status, through the vehicle CAN bus interface. The spatiotemporal synchronization fusion unit is used to perform timestamp alignment, coordinate system unification, and data normalization on the first time-series features, the second time-series features, and the vehicle dynamic parameter sequence to generate fused perception data.

3. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 2, characterized in that: The personalized driver profile generation module includes: The short-term state analysis submodule is used to acquire, in real time, the driver's fatigue index, attention distraction index, emotional agitation index, and operational compliance index based on the fused perception data corresponding to the current driver, and to determine the driver's fatigue index. Distraction index Emotional Excitement Index and control standardization index Construct short-time state vectors; The long-term habit learning submodule, when the driver identification unit identifies the driver as an experienced driver, inputs the fused perception data of the current driver within the historical driving cycle into the trained following distance output model, lane change style evaluation model, and risk response time output model, respectively, to obtain the driver's following distance habit value, lane change style type, and risk response habit reaction time. Based on the following distance habit value, lane change style type, and risk response habit reaction time, the module constructs the driver's long-term habit vector; among which, the lane change style type includes aggressive, steady, and decisive. When the driver identification unit identifies the driver as a new driver, the default long-term habit vector is loaded as the initial value. The driving profile generation submodule is used to fuse the short-term state vector obtained by the short-term state analysis submodule and the long-term habit vector obtained by the long-term habit learning submodule to generate a personalized driving profile of the current driver.

4. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 3, characterized in that: The short-time state analysis submodule includes: The eye state monitoring unit is used to calculate the driver's fatigue index within a current preset time window based on the current driver's eye opening and closing sequence, head three-dimensional Euler angle sequence, and facial key point motion sequence from fused perception data. ; The gaze direction analysis unit is used to calculate the driver's attention distraction index within the current preset time window based on the current driver gaze direction vector sequence and hand key point coordinate sequence from the fused perception data. ; The facial emotion recognition unit is used to calculate the driver's emotional arousal index within a current preset time window based on the motion sequence of key facial points in the fused perception data. ; The driving behavior evaluation unit is used to calculate the driver's driving compliance index within the current preset time window based on the sequence of vehicle dynamic parameters obtained from the vehicle bus. .

5. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 4, characterized in that: The facial emotion recognition unit includes: The emotion-related key point grouping and parsing subunit is used to divide key point groups into two core emotions, driving stress and anger, based on the facial key point motion sequence in the fused perception data and extract the motion features of the corresponding key points. The key point groups include stress emotion association group and anger emotion association group. The stress-related emotion group includes four key points on the inner and outer sides of both eyebrows, and six key points on the upper and lower edges of both eyelids and the corners of the eyes; the anger-related emotion group includes four key points on the inner and outer sides of both eyebrows, and four key points on the upper and lower lips; the corresponding facial movements for the stress-related emotion group include drooping eyebrows, tight eyelids, abnormal blinking frequency, and stiff eye muscles; the corresponding facial movements for the anger-related emotion group include furrowed eyebrows, retracted corners of the mouth, pursed lips, and tight lips; the motion characteristics include the temporal variation of three-dimensional coordinates, the rate of change of three-dimensional coordinates, the relative distance variation of key points in the same group, and the regional contour deformation coefficient; The motion unit intensity dynamic quantization subunit is used to detect facial movements associated with stress and anger within the current preset time window, and record the duration and intensity dynamic feature vector of continuous activation of each emotion-associated facial movement. The intensity dynamic feature vector corresponding to the k-th emotion-related facial action is represented as follows: ;in, For the first The average intensity of facial movements associated with various emotions; For the first Peak intensity of facial movements associated with certain emotions; For the first The average slope of the rising edge of the intensity of facial movements associated with certain emotions; The multi-dimensional emotion contribution calculation subunit is used to calculate the contribution of each emotion-related facial action to the overall emotional state based on the dynamic feature vector of the duration and intensity of continuous activation of each emotion-related facial action, and the emotion association weight of each emotion-related facial action. The emotional arousal index generation subunit is used to accumulate the contribution of all activated facial movements within the current preset time window to generate the driver's emotional arousal index. .

6. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 5, characterized in that: The multi-dimensional emotion contribution calculation subunit calculates the contribution of each emotion-related facial movement to the overall emotional state, including: ; in, For the first The contribution of facial movements associated with emotions to the overall emotional state; For the first Preset weighting coefficients for facial movements associated with emotions; For the first Emotional association weights for facial movements related to emotions; For the first The duration of continuous activation of facial movements associated with certain emotions; , , These are the feature weight coefficients corresponding to the average intensity, peak intensity, and average slope of intensity rise, respectively. , , These are the feature activation thresholds corresponding to average intensity, peak intensity, and average slope of intensity increase, respectively. It is a piecewise nonlinear saturation function used to limit the excessive contribution of a single feature.

7. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 2, characterized in that: The environmental risk perception and prediction module includes: The target detection and tracking unit is used to detect and identify dynamic targets and static traffic elements in the vehicle's surrounding environment in real time based on the second temporal features in the fused perception data and through a preset target detection algorithm, and to perform cross-frame trajectory association and continuous tracking of each dynamic target. The risk situation modeling unit is used to calculate the instantaneous collision risk value of each dynamic target based on the identified static traffic elements and the historical trajectory, current motion state and relative position relationship of each dynamic target with respect to the vehicle, and to construct a risk distribution map of the current moment centered on the vehicle. The short-term prediction unit is used to predict the position and status of each dynamic target within a preset time period in the future, based on the current risk distribution map and the motion model of each dynamic target, and to generate a current environmental risk situation map.

8. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 3, characterized in that: The risk coupling assessment and decision-making module includes: The coupling risk calculation unit is used to calculate the coupling risk value of the corresponding potential risk event in the current driver state based on each potential risk event identified in the current environmental risk situation map, combined with all parameters of the current driver's personalized driving profile, through a nonlinear coupling model. The comprehensive risk assessment unit is used to conduct a comprehensive risk assessment of all potential risk events based on the coupled risk values ​​of all potential risk events and in combination with the temporal and spatial urgency of the potential risk events. The early warning decision generation unit is used to generate the corresponding early warning intervention decision based on the comprehensive risk level and type of all potential risk events, combined with the current driver's personalized driving profile parameters.

9. A driving safety early warning system based on a video surveillance vehicle terminal according to claim 1, characterized in that: The multimodal adaptive early warning module includes: The warning strategy library is used to store at least one preset warning strategy template. Each preset warning strategy template is associated with a specific combination of comprehensive risk level, potential risk event type and current driver personalized driving profile parameters. The strategy selection and optimization unit is used to match the baseline warning strategy from the warning strategy library based on the warning intervention decision, and to perform personalized optimization of the baseline warning strategy by combining all parameters of the current driver's personalized driving profile and the driver's historical warning feedback data, and generate the final warning instruction.