Driver state early warning system based on large model

By combining roadside image acquisition with multimodal large model analysis of microscopic actions and vehicle trajectory features, the limitations of single-modal detection are overcome, enabling high-precision, low-false-alarm real-time early warning of driver status.

CN121982684APending Publication Date: 2026-05-05CHIZHOU CITY PUBLIC SECURITY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHIZHOU CITY PUBLIC SECURITY BUREAU
Filing Date
2026-01-14
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing driver condition monitoring technologies mainly rely on single-modal detection, which makes it difficult to identify complex actions such as the use of prohibited substances. Furthermore, the warning signals are often delayed, making it impossible to accurately distinguish the causes of impaired driving ability.

Method used

The video stream is acquired using a roadside image acquisition module and processed in real time using an edge computing module. A multimodal large model is used to make a comprehensive judgment by combining micro-action temporal features and vehicle motion trajectory features. A multi-level continuous state verification module is used to ensure the reliability of the early warning decision.

Benefits of technology

It improves the accuracy of identifying complex dangerous behaviors such as the use of prohibited substances and the reliability of early warnings, reduces the false alarm rate, and enables real-time and accurate monitoring of the driver's status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982684A_ABST
    Figure CN121982684A_ABST
Patent Text Reader

Abstract

The invention relates to the field of traffic monitoring, and discloses a driver state early warning system based on a large model, which comprises a roadside image acquisition module, an edge calculation module, a center analysis module and an early warning module, the edge calculation module is used for processing the video stream in real time, identifying forbidden articles in the vehicle and extracting microcosmic action time sequence characteristics of a driver and vehicle movement track characteristics; the central analysis module performs space-time association and semantic understanding on the features by using a pre-trained multi-modal large model, and comprehensively judges dangerous driving behaviors and risk levels; the system further comprises a multi-stage persistent state verification module which is used for carrying out secondary verification of the track and the physiological damage state on the primary high-risk judgment and carrying out graded decision making according to the verification result. Through multi-modal fusion analysis and a multi-stage verification mechanism, the accuracy and reliability of hidden dangerous driving behavior recognition are improved, and a structured law enforcement evidence packet can be automatically generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of traffic monitoring, and specifically to a driver status early warning system based on a large model. Background Technology

[0003] Currently, driver condition monitoring technologies primarily employ single-modal detection techniques. These techniques typically rely on the analysis of data from a single type of sensor. For example: Vision-based behavior analysis: This method captures facial images of drivers using vehicle-mounted or roadside cameras and analyzes features such as eye opening and closing, gaze direction, and head posture using computer vision algorithms to determine fatigue or distraction. However, this technology struggles to identify complex hand-mouth coordination actions, such as the use of prohibited substances.

[0004] Vehicle behavior-based analysis: This method analyzes vehicle trajectories through video footage to monitor anomalies such as lane departure and steering wheel angle fluctuations. However, the warning signals from this method are delayed, typically triggered only after substantial impairment of driving ability has occurred, and it cannot distinguish whether the impairment is due to substance abuse, sudden illness, or ordinary distraction. Summary of the Invention

[0005] The purpose of this invention is to provide a driver status early warning system based on a large model, thereby solving the above-mentioned technical problems.

[0006] The objective of this invention can be achieved through the following technical solutions: A driver status early warning system based on a large model includes: The roadside image acquisition module is used to acquire video streams that include the face and hand areas of the target vehicle driver; An edge computing module, deployed on the roadside, is used to process the video stream in real time. The processing includes: identifying whether there are preset prohibited items in the vehicle based on a pre-trained target detection model, and extracting the micro-motion timing features of the driver and the motion trajectory features of the vehicle. The central analysis module is used to receive the identification results of the prohibited items, the temporal features of the micro-actions, and the trajectory features of the motion, and input them into the pre-trained multimodal large model; the multimodal large model is used to perform spatiotemporal correlation and semantic understanding of items, personnel actions, and vehicle status, so as to comprehensively judge dangerous driving behavior and risk level; The early warning module is used to automatically generate early warning instructions when a behavior is determined to be high-risk.

[0007] As a further technical solution, the analysis process of the multimodal large model specifically includes: The existence state of the prohibited items, the timing features of the micro-actions, and the trajectory features are aligned and mapped to a unified temporal embedding space. Based on the attention mechanism, cross-modal association weights between different modal features in the temporal embedding space are calculated to infer behavioral intent and the degree of impairment of driving ability; Based on the road segment type and time information of the target vehicle, a context-aware weighted assessment of the hazard level is performed.

[0008] As a further technical solution, the driver's micro-action timing features are extracted and calculated through the following steps: The judgment is based on the spatial relationship between the hand key points and the bounding box of the identified prohibited item. When the center point of the bounding box of the item is continuously within the area with the hand key points as the center and the preset pixel distance as the radius for more than the first preset number of frames, it is determined to be a handheld state. In handheld mode, calculate the Euclidean distance sequence D(t) from the center point of the object's bounding box to the standard point of the face; if D(t) decreases from greater than the first distance threshold to less than the second distance threshold within a second preset number of frames, and remains below the second distance threshold for more than a third preset number of frames, then a suspected ingestion action is determined. The pupil aspect ratio is calculated based on facial feature points, and the variance of the pupil aspect ratio within the sliding time window is statistically analyzed as an indicator of pupil instability. At the same time, the standard deviation of the head pitch angle and yaw angle within the short time window is calculated as a quantitative indicator of the decline in head control.

[0009] As a further technical solution, the vehicle's motion trajectory features are extracted and calculated from the video stream through the following steps: For each frame of the image, perspective transformation and lane line detection are performed. A coordinate system with the lane center line as the reference is established, and the lateral pixel position X(t) of the vehicle in the image is determined by the vehicle detection box. Transform X(t) to a real-world approximate coordinate system with the lane center as the zero point to obtain the lateral offset sequence f(t); calculate the standard deviation of f(t) within the evaluation window, and the cumulative time percentage of Of(t) exceeding the preset safety threshold; After applying a low-pass filter to the horizontal position sequence X(t), the mean of the absolute value sequence of the first difference of X(t) is calculated as a quantitative indicator of trajectory jitter.

[0010] As a further technical solution, the system also includes a multi-level continuous state verification module; The multi-level continuous status verification module is activated after the central analysis module initially outputs a high-risk judgment, and performs first-level continuous verification, specifically as follows: Within the first verification time window after the initial judgment, the vehicle's motion trajectory characteristics are continuously analyzed; If, among the motion trajectory features, the standard deviation of the lateral offset sequence f(t), the cumulative time percentage exceeding the preset safety threshold, or the quantitative index of trajectory jitter continuously exceeds the corresponding threshold, then it is determined that the trajectory abnormality persists.

[0011] As a further technical solution, if it is determined that the trajectory anomaly persists, the multi-level persistent state verification module performs a second-level deep verification, specifically as follows: The edge computing module is instructed to increase the sampling frequency of the driver's facial features for the target vehicle and obtain updated microscopic action temporal features; Based on the updated micro-action timing characteristics, a comprehensive physiological damage persistence risk index is calculated. The calculation formula is as follows: ; ; in: and This refers to the start and end times of the secondary depth verification time window; for The aspect ratio of the pupil at any given moment. This is the reference value for the pupil aspect ratio of the driver under normal conditions; This represents the time integral of the positive deviation of the pupil aspect ratio from the normal baseline within the secondary depth verification time window, used to quantify the cumulative degree of pupil instability. and These are the standard deviations of the head pitch angle and yaw angle within the secondary depth verification window, respectively. This means that the larger of the two absolute values ​​is taken as the representative value of the head's axial instability. and These are preset weighting coefficients used to balance the contribution weights of pupil and head instability indicators; like If the risk of physiological damage exceeds the preset threshold, it is determined that the physiological damage state is ongoing.

[0012] As a further technical solution, the multi-level continuous state verification module also includes the following process: When the Level 1 continuous verification determines that the trajectory abnormality persists and the Level 2 deep verification determines that the physiological damage state persists, the initial high-risk determination is confirmed to be valid, and the final risk confirmation level is upgraded. When only the first-level verification result is valid, the warning will be marked as requiring manual review; If both levels of verification results are invalid, the initial high-risk determination will be automatically revoked.

[0013] As a further technical solution, the duration of the first verification time window is dynamically adjusted based on the risk confidence level at the time of the initial judgment, the average vehicle speed of the current road section, and the weather visibility.

[0014] As a further technical solution, the system also includes: The evidence generation module is used to simultaneously solidify key video clips containing the entire process of the behavior when automatically generating warning instructions. These clips are then bound to vehicle identity information, behavior type tags, time, and geographical location to generate a structured law enforcement evidence package.

[0015] The beneficial effects of this invention are: (1) This invention uses the edge computing module to quickly extract the quantitative features of micro-actions and trajectories, providing a high signal-to-noise ratio input for the center analysis; the center's multimodal large model uses its powerful temporal modeling and cross-modal attention mechanism to deeply understand the causal relationship between holding objects, approaching the mouth and nose, and losing control of the vehicle, thereby improving the semantic accuracy and scene adaptability of behavior recognition, thus overcoming the shortcomings of traditional methods in accurately recognizing and understanding complex dangerous behaviors such as driving while inhaling prohibited substances.

[0016] (2) This invention achieves enhanced reliability of early warning decision-making from single judgment to dynamic verification by combining the initial judgment of the multimodal large model with the multi-level continuous state verification module, thereby overcoming the technical problem of high false alarm rate caused by occasional actions or brief interference. Specifically, after the multimodal large model makes a high-risk initial judgment, it does not immediately make a final judgment, but initiates a multi-level review process that includes trajectory continuous verification and high-sampling physiological verification, ensuring that only those cases where both behavior and damage status are continuous will be finally confirmed, greatly reducing false alarms. Attached Figure Description

[0017] The invention will now be further described with reference to the accompanying drawings.

[0018] Figure 1 This is a schematic diagram of the logical structure of the present invention; Figure 2 This is a schematic diagram of the logical structure of the multi-level persistent state verification module in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figures 1-2 As shown, the present invention is a driver status early warning system based on a large model, comprising: The roadside image acquisition module is used to acquire video streams that include the face and hand areas of the target vehicle driver; An edge computing module, deployed on the roadside, is used to process the video stream in real time. The processing includes: identifying whether there are preset prohibited items in the vehicle based on a pre-trained target detection model, and extracting the micro-motion timing features of the driver and the movement trajectory features of the vehicle; the prohibited items are laughing gas balloons or laughing gas canisters; the feature library needs to be pre-loaded with the target detection model; The central analysis module is used to receive the identification results of the prohibited items, the temporal features of the micro-actions, and the trajectory features of the motion, and input them into the pre-trained multimodal large model; the multimodal large model is used to perform spatiotemporal correlation and semantic understanding of items, personnel actions, and vehicle status, so as to comprehensively judge dangerous driving behavior and risk level; The early warning module is used to automatically generate early warning instructions when a behavior is determined to be high-risk.

[0021] The multimodal large model is a visual-text cross-modal model based on the Transformer architecture. The inputs are the prohibited item recognition results (0 / 1 binary), micro-action temporal features (multidimensional vector), and motion trajectory features (multidimensional vector). The output is the risk level, including low / medium / high / extremely high.

[0022] Risk level classification criteria: Low risk, i.e. no prohibited items, no abnormal actions, and stable trajectory; Medium risk, i.e. single-modal anomaly, such as only slight trajectory jitter; High risk, i.e. bimodal anomaly, such as detection of prohibited items or suspected drug use; Very high risk, i.e. bimodal anomaly plus passing multiple levels of verification.

[0023] Example: Roadside image acquisition: Deploy high-definition infrared cameras with a resolution of 1920×1080 and a frame rate of 30 frames per second on roadside gantries or poles, with the lenses aimed at the driver's seat of oncoming vehicles to ensure coverage of the driver's face (eyes, mouth and nose) and hand area, and collect video streams and transmit them to the edge computing module in real time.

[0024] By deploying a dedicated high-resolution camera with infrared illumination, clear and usable images of the driver's face and hands can be continuously acquired in bright daylight, at night, or in low-light environments such as tunnels. This provides a stable data source for subsequent accurate feature extraction and overcomes the technical shortcomings of traditional vision solutions that are severely limited by lighting conditions.

[0025] Edge computing module processing: Call the pre-trained YOLOv8 object detection model to detect prohibited items in each frame of the image, with a confidence threshold of ≥0.7, and output the presence / absence result; Synchronous operation is achieved through real-time detection using the MediaPipe human pose estimation model based on convolutional neural networks, extracting 21 key points of the driver's hands and 68 feature points of the face, and calculating the temporal features of micro-movements. Lane line detection and vehicle localization are performed on each frame of the image, and vehicle motion trajectory features are extracted.

[0026] The high-precision target detection and real-time pose estimation model is run in parallel on the roadside edge device, realizing the synchronous and low-latency extraction of multi-target information such as objects, hands, faces and vehicles. The complex calculation of massive video streams is moved to the edge, and only structured feature data is uploaded to the center, which greatly alleviates the network bandwidth pressure and ensures the real-time response of the system, providing high-quality and high signal-to-noise ratio input for subsequent intelligent analysis at the center.

[0027] Central analysis module processing: The three types of data received from the edge computing module are aligned with timestamps with an accuracy of ≤10ms and mapped to a unified 64-dimensional temporal embedding space. The multimodal large model calculates cross-modal weights based on the attention mechanism, and performs weighted evaluation by combining road segment type and time to output risk level; among them, highways, urban areas and rural areas are quantified as 1, 2 and 3 respectively; daytime and nighttime are quantified as 1 and 2 respectively.

[0028] Through rigorous timestamp alignment and unified embedding space mapping, heterogeneous features from different processing flows are effectively fused at both temporal and semantic levels. The multimodal large model utilizes an attention mechanism to automatically learn and assign higher weights to the causal relationship between holding a prohibited item close to the mouth and nose and subsequent abnormal vehicle trajectory. This enables behavioral risk assessment based on deep semantic understanding, rather than simple feature aggregation, thus improving the accuracy and interpretability of the judgment.

[0029] Early warning module response: If the risk level is high, an early warning instruction will be automatically generated and transmitted to the traffic police command platform via 4G / 5G.

[0030] The analysis process of the multimodal large model specifically includes: The existence state of the prohibited items, the timing features of the micro-actions, and the trajectory features are aligned and mapped to a unified temporal embedding space. The unified temporal embedding space consists of a 64-dimensional real vector space, where the state of the prohibited item occupies 1 dimension (0 = non-existent, 1 = existent); micro-action features occupy 31 dimensions (1 dimension of holding state + 1 dimension of inhalation action + 1 dimension of pupil variance + 1 dimension of head pitch angle standard deviation + 1 dimension of head yaw angle standard deviation + 31-dimensional temporal sub-features); and motion trajectory features occupy 32 dimensions (1 dimension of lateral offset standard deviation + 1 dimension of cumulative timeout percentage + 1 dimension of trajectory jitter mean + 32-dimensional temporal sub-features).

[0031] Based on the attention mechanism, cross-modal association weights between different modal features in the temporal embedding space are calculated to infer behavioral intent and the degree of impairment of driving ability; The cross-modal association weight calculation method uses scaled dot product attention, with the formula: Attention(Q,K,V)=softmax( V; where Q (query) is the action feature, K (key) is the item status + trajectory feature, and V (value) is the risk semantic vector.

[0032] Based on the road segment type and time information of the target vehicle, a context-aware weighted assessment of the hazard level is performed.

[0033] Among them, the context-weighted evaluation rule is: Road segment type weighting: Expressway, weight 1.2; Urban area, weight 1.0; Rural area, weight 0.8; Time weighting: Nighttime, weight 1.1; Daytime, weight 0.9; The final risk score is calculated as: cross-modal association score × road segment weight × time weight. A score ≥ 0.8 indicates high risk.

[0034] The training data of the multimodal large model includes paired samples consisting of roadside monitoring video clips and corresponding behavioral description texts, and the model learns the mapping from video clips to risk semantics through fine-tuning instructions.

[0035] Example: Feature alignment and embedding mapping: Based on the video frame timestamp, the prohibited item status (1 value per frame), micro-motion features (31 values ​​per frame), and motion trajectory features (32 values ​​per frame) are aligned according to the time sequence to form a feature sequence of length T (T=30 frames, i.e. 1 second); Each frame's features are mapped to a 64-dimensional embedding vector through a fully connected layer, resulting in a T×64 temporal embedding matrix.

[0036] By compressing and mapping the original features of up to 64 dimensions per frame into a unified temporal embedding vector, not only is data regularization achieved, but the most relevant semantic information is also preserved during the dimensionality reduction process. This provides a standardized input format for subsequent Transformer-based model processing, enabling the model to efficiently learn long-distance temporal dependencies.

[0037] Cross-modal association weight calculation: The encoder of the multimodal large model performs layer normalization on the temporal embedding matrix, calculates the association weights between action features and item and trajectory features through a multi-head attention mechanism (8 attention heads), and outputs cross-modal fusion features (64 dimensions).

[0038] Context-aware weighted evaluation: The current road segment type is obtained from the GIS system, and the time is obtained from the camera sensor (6:00-18:00 during the day and vice versa at night), and then quantified into corresponding weights. The fused features are mapped to an initial risk score of 0-1 by the decoder, and then multiplied by the road segment weight and time weight to obtain the final risk score and determine the risk level.

[0039] The timing features of the driver's micro-actions are extracted and calculated through the following steps: The judgment is based on the spatial relationship between the hand key points and the bounding box of the identified prohibited item. When the center point of the bounding box of the item is continuously within the area with the hand key points as the center and the preset pixel distance as the radius for more than the first preset number of frames, it is determined to be a handheld state. In handheld mode, calculate the Euclidean distance sequence D(t) from the center point of the object's bounding box to the standard point of the face; if D(t) decreases from greater than the first distance threshold to less than the second distance threshold within a second preset number of frames, and remains below the second distance threshold for more than a third preset number of frames, then a suspected ingestion action is determined. The tip of the nose, or point number 30, among the 68 facial feature points was selected as the reference point for calculating the Euclidean distance.

[0040] The pupil aspect ratio was calculated based on facial feature points, and the variance of the pupil aspect ratio within the sliding time window was statistically analyzed as an indicator of pupil instability. Simultaneously, the standard deviations of head pitch and yaw angles within a short time window were calculated as a quantitative indicator of decreased head control. The pupil region was fitted from eye feature points (points 36-47), and the aspect ratio was calculated as pupil height / pupil width. The hand key points and facial feature points are detected in real time using a human posture estimation model based on a convolutional neural network; the head pitch angle and yaw angle are calculated using a spatial perspective transformation algorithm based on the correspondence between the detected two-dimensional facial key points and a predefined three-dimensional standard head model.

[0041] For example, the first preset frame rate is 10 frames, which is approximately 0.33 seconds, and the video frame rate is 30 frames per second. The second preset frame rate is 8 frames, approximately 0.27 seconds; the third preset frame rate is 5 frames, approximately 0.17 seconds. The first distance threshold is 50 pixels, which corresponds to an actual distance of about 30cm; the second distance threshold is 20 pixels, which corresponds to an actual distance of about 12cm. Sliding time window = 2 seconds, or 60 frames; short time window = 1 second, or 30 frames.

[0042] Example: Handheld status determination: The human pose estimation model outputs the coordinates of 21 key points on the hand in real time, drawing a circle with the fingertip of the index finger (point 8) as the center and a preset radius of 30 pixels; it is compatible with camera shooting distances of 5-10 meters. If the center point of the prohibited item's bounding box falls within the circle for more than 10 frames, it is determined to be in a handheld state and marked as action tag 1; otherwise, it is marked as 0.

[0043] Detection of suspected drug use: In handheld mode, the Euclidean distance D(t) from the center point of the bounding box of the prohibited item to the 30th nose tip point is calculated in each frame; If D(t) drops from >50 pixels to <20 pixels within 8 frames and remains <20 pixels for more than 5 frames, it is determined as a suspected ingestion action, and the timestamp of the action is recorded.

[0044] By defining precise geometric relationships (circle radius) and time duration (preset frame count) criteria, subjective descriptions of handheld and inhalation behaviors are transformed into objective and quantifiable algorithmic logic. It has strong anti-interference capabilities and can effectively distinguish between handheld items and items in other locations inside the vehicle, as well as inhalation actions and similar normal behaviors such as drinking water and touching the face, thus reducing the false trigger rate from the source.

[0045] Pupil and head index calculation: Within a 60-frame sliding time window, the variance of the pupil aspect ratio is calculated. The normal range is 0.01-0.03, and anything exceeding 0.05 is considered unstable. Within a short time window of 30 frames, the standard deviation of the head pitch angle (vertical sway) and yaw angle (horizontal sway) is calculated using a spatial perspective transformation algorithm. A standard deviation greater than 5° indicates a decrease in head control.

[0046] Using the variance of pupillary aspect ratio and the standard deviation of head posture angle as physiological indicators can sensitively capture subtle physiological changes such as micro-tremors of the eyeballs, gaze instability, and decreased neck muscle control caused by the influence of neuroactive substances. These serve as indirect evidence of impaired driving ability and complement direct behavioral characteristics, providing a multi-dimensional physiological basis for risk assessment.

[0047] The vehicle's motion trajectory features are extracted and calculated from the video stream through the following steps: For each frame of the image, perspective transformation and lane line detection are performed. A coordinate system with the lane center line as the reference is established, and the lateral pixel position X(t) of the vehicle in the image is determined by the vehicle detection box. Transform X(t) to a real-world approximate coordinate system with the lane center as the zero point to obtain the lateral offset sequence f(t); calculate the standard deviation of f(t) within the evaluation window, and the cumulative time percentage of Of(t) exceeding the preset safety threshold; After applying a low-pass filter to the horizontal position sequence X(t), the mean of the absolute value sequence of the first difference of X(t) is calculated as a quantitative indicator of trajectory jitter.

[0048] Perspective transformation parameters: Based on camera intrinsic parameters, focal length f=12mm, pixel size 1.4μm, camera installation height 5 meters, the image pixel coordinates are converted to world coordinates using the Homography matrix. The conversion formula is as follows: ,in The x-coordinate of the image center point. This refers to the camera height.

[0049] Preset safety thresholds: safety thresholds for lateral offset f(t), 0.5 meters for highway sections and 0.3 meters for urban sections; cumulative time percentage threshold within the evaluation window = 30%.

[0050] Low-pass filter parameters: Butterworth low-pass filter is used, cutoff frequency = 2Hz, to filter high-frequency noise and preserve trajectory trend; evaluation window = 3 seconds, i.e. 90 frames.

[0051] Example: Coordinate system establishment: Canny edge detection and Hough line detection are performed on each frame of the image to extract lane lines. A pixel coordinate system is established with the center line of the two lanes as the y-axis and the direction perpendicular to the lane lines as the x-axis. The pixel coordinate system is converted to the world coordinate system through perspective transformation, with the center of the lane as the origin and the x-axis pointing to the right side of the road. The unit is meters.

[0052] Horizontal offset sequence calculation: The center x-coordinate X(t) (pixels) of the vehicle detection box is obtained by outputting the YOLOv8 object detection model, and then converted into the lateral offset f(t) (meters) in world coordinates. Calculate the standard deviation of f(t) within the 90-frame evaluation window. The normal range is <0.1 meters, and a range exceeding 0.2 meters is considered abnormal. Also calculate the percentage of frames in which f(t) exceeds the safety threshold. A range exceeding 30% is considered abnormal.

[0053] Track jitter index calculation: Apply a Butterworth low-pass filter to X(t) with a cutoff frequency of 2Hz to remove high-frequency noise; The first-order difference of X(t) after filtering is calculated, which is the change in the horizontal position between adjacent frames. The mean of the absolute value sequence is taken. The normal range is <5 pixels / frame, and more than 10 pixels / frame is considered as trajectory jitter.

[0054] By transforming the image coordinates to the real-world coordinate system through perspective transformation, the measurement of indicators such as lateral offset and jitter becomes physically meaningful. The evaluation results are not affected by the relative distance between the vehicle and the camera, and the evaluation criteria are more objective and uniform. Combined with the trajectory jitter index extracted by low-pass filtering, it can effectively identify high-frequency, small-amplitude unstable vibrations in steering wheel control, which are typical characteristics of fatigue or physical damage driving, thus improving the fineness of trajectory analysis.

[0055] The system also includes a multi-level continuous state verification module; The multi-level continuous status verification module is activated after the central analysis module initially outputs a high-risk judgment, and performs first-level continuous verification, specifically as follows: Within the first verification time window after the initial judgment, the vehicle's motion trajectory characteristics are continuously analyzed; If, among the motion trajectory features, the standard deviation of the lateral offset sequence f(t), the cumulative time percentage exceeding the preset safety threshold, or the quantitative index of trajectory jitter continuously exceeds the corresponding threshold, then it is determined that the trajectory abnormality persists.

[0056] The duration of the first verification time window is dynamically adjusted based on the risk confidence level at the time of the initial judgment, the average vehicle speed on the current road segment, and the weather visibility.

[0057] The introduction of a dynamically adjusted first verification time window enables the system to allocate observation resources based on the urgency of the risk (confidence level) and the complexity of the environment (vehicle speed and visibility). In high-risk, high-speed situations, the window is shortened for rapid response, while in low-visibility situations, the window is appropriately extended to collect more sufficient evidence, reflecting decision-making flexibility and optimizing the balance between resource utilization and judgment reliability.

[0058] The dynamic adjustment rules for the first verification time window are as follows: The method for quantifying risk confidence is as follows: the risk score output by the central analysis module ranges from 0 to 1, which is directly used as the C value; for example, if the score is 0.85, then C=0.85. Quantification rules for average vehicle speed and visibility on road sections: Vehicle speed V: ≤60km / h (urban area) → 0.3; 60-100km / h (suburbs) → 0.7; >100km / h (highway) → 1.0; Visibility L: ≤2km→0.5; 2-5km→0.7; >5km→1.0.

[0059] Base window duration = 5 seconds, or 150 frames; Adjustment formula: T = T0 × (1 - 0.2 × C + 0.1 × V - 0.1 × L), Where C is the initial risk confidence level, ranging from 0 to 1; V is the average vehicle speed on the road segment, in km / h, quantified as 0.1-1.0; and L is the visibility, in km, quantified as 0.5-1.0. Example: Confidence level 0.9 (high), vehicle speed 100km / h (high speed), visibility 5km, T=5×(1-0.18+0.1-0.1)=4.6 seconds.

[0060] Lateral offset standard deviation threshold = 0.25 meters; Cumulative time percentage threshold = 40%; The threshold for trajectory jitter quantification index is 12 pixels / frame.

[0061] Example: Verification Activation: When the central analysis module outputs a high-risk initial judgment, the first-level continuous verification is automatically activated.

[0062] Dynamic window determination: Obtain the initial risk confidence level C from the central database, obtain the average vehicle speed V from the roadside radar, and obtain the visibility L from the camera sensor. Substitute these values ​​into the formula to calculate the first verification time window T.

[0063] Continuous trajectory feature analysis: Within time T, the edge computing module continuously extracts vehicle motion trajectory features, including lateral offset standard deviation, cumulative timeout percentage, and trajectory jitter index; If at least one of the three indicators continuously exceeds the corresponding threshold for a duration ≥ 80% of T, it is determined that the trajectory abnormality persists; otherwise, it is determined that the trajectory abnormality has disappeared.

[0064] If the trajectory anomaly is determined to persist, the multi-level persistent state verification module performs a second-level deep verification, specifically: The edge computing module is instructed to increase the sampling frequency of the driver's facial features for the target vehicle and obtain updated microscopic action temporal features; The original sampling frequency was 30 frames per second, which was increased to 60 frames per second; the edge computing module prioritizes allocating computing power to the target vehicle. Based on the updated micro-action timing characteristics, a comprehensive physiological damage persistence risk index is calculated. The calculation formula is as follows: ; ; in: and This refers to the start and end times of the secondary depth verification time window; for The aspect ratio of the pupil at any given moment. This is the reference value for the pupil aspect ratio of the driver under normal conditions; Acquisition method: During the initial 30 seconds of normal driving, when there are no prohibited items, no abnormal movements, and the trajectory is stable, 60 frames of pupil aspect ratio are collected, and the average value is taken as the driver's pupil aspect ratio. ; This represents the time integral of the positive deviation of the pupil aspect ratio from the normal baseline within the secondary depth verification time window, used to quantify the cumulative degree of pupil instability. and These are the standard deviations of the head pitch angle and yaw angle within the secondary depth verification window, respectively. This means that the larger of the two absolute values ​​is taken as the representative value of the head's axial instability. and These are preset weighting coefficients used to balance the contribution weights of pupil and head instability indicators; after calibration with training data, the two indicators are balanced. 0.8 It is 0.5; like If the risk exceeds the preset physiological damage risk threshold, the physiological damage state is considered to be persistent. Physiological damage risk threshold = 3.0. >3.0 is considered a continuing injury.

[0065] The method used in the second-level verification Composite formula, its exponent term The design aims to amplify persistent pupillary abnormalities, i.e. This design is based on the physiological principle that persistent pupillary abnormalities are a better indicator of central nervous system inhibition than transient fluctuations. Therefore, this formula can more accurately distinguish between persistent physiological impairment caused by material damage and transient physiological noise, significantly improving the specificity of physiological damage assessment.

[0066] Example: Increased sampling frequency: The multi-level continuous state verification module sends instructions to the edge computing module to increase the sampling frequency of the driver's facial features of the target vehicle from 30 frames / second to 60 frames / second, ensuring more accurate capture of physiological features.

[0067] Microscopic motion feature update: The edge computing module collects facial feature points at 60 frames per second and calculates the pupil aspect ratio in real time. Standard deviation of head pitch angle Yaw angle standard deviation .

[0068] Persistent risk index of physiological damage Calculation: Setting the secondary verification time window [ , = 3 seconds, or 180 frames; Calculation , is a numerical integral, unit: second, dimensionless; calculate Unit: degrees; Substitute into the formula: .

[0069] Judgment result: If If the value is greater than 3.0, the physiological damage state is considered to be ongoing; otherwise, the physiological damage state is considered to have disappeared.

[0070] The multi-level persistent state verification module also includes the following processes: When the Level 1 continuous verification determines that the trajectory abnormality persists and the Level 2 deep verification determines that the physiological damage state persists, the initial high-risk determination is confirmed to be valid, and the final risk confirmation level is upgraded. Risk assessment level upgrade rules: Initial high risk → passing two levels of verification → upgraded to extremely high risk; When only the first-level verification result is valid, the warning will be marked as requiring manual review; If both levels of verification results are invalid, the initial high-risk determination will be automatically revoked.

[0071] The system also includes: The evidence generation module is used to simultaneously solidify key video clips containing the entire process of the behavior when automatically generating warning instructions. These clips are then bound to vehicle identity information, behavior type tags, time, and geographical location to generate a structured law enforcement evidence package.

[0072] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A driver status early warning system based on a large model, characterized in that, include: The roadside image acquisition module is used to acquire video streams that include the face and hand areas of the target vehicle driver; An edge computing module, deployed on the roadside, is used to process the video stream in real time. The processing includes: identifying whether there are preset prohibited items in the vehicle based on a pre-trained target detection model, and extracting the micro-motion timing features of the driver and the motion trajectory features of the vehicle. The central analysis module is used to receive the identification results of the prohibited items, the temporal features of the micro-actions, and the trajectory features of the motion, and input them into the pre-trained multimodal large model; the multimodal large model is used to perform spatiotemporal correlation and semantic understanding of items, personnel actions, and vehicle status, so as to comprehensively judge dangerous driving behavior and risk level; The early warning module is used to automatically generate early warning instructions when a behavior is determined to be high-risk.

2. The driver status early warning system based on a large model according to claim 1, characterized in that, The analysis process of the multimodal large model specifically includes: The existence state of the prohibited items, the timing features of the micro-actions, and the trajectory features are aligned and mapped to a unified temporal embedding space. Based on the attention mechanism, cross-modal association weights between different modal features in the temporal embedding space are calculated to infer behavioral intent and the degree of impairment of driving ability; Based on the road segment type and time information of the target vehicle, a context-aware weighted assessment of the hazard level is performed.

3. The driver status early warning system based on a large model according to claim 2, characterized in that, The timing features of the driver's micro-actions are extracted and calculated through the following steps: The judgment is based on the spatial relationship between the hand key points and the bounding box of the identified prohibited item. When the center point of the bounding box of the item is continuously within the area with the hand key points as the center and the preset pixel distance as the radius for more than the first preset number of frames, it is determined to be a handheld state. In handheld mode, calculate the Euclidean distance sequence D(t) from the center point of the object's bounding box to the standard point of the face; if D(t) decreases from greater than the first distance threshold to less than the second distance threshold within a second preset number of frames, and remains below the second distance threshold for more than a third preset number of frames, then a suspected ingestion action is determined. The pupil aspect ratio is calculated based on facial feature points, and the variance of the pupil aspect ratio within the sliding time window is statistically analyzed as an indicator of pupil instability. At the same time, the standard deviation of the head pitch angle and yaw angle within the short time window is calculated as a quantitative indicator of the decline in head control.

4. The driver status early warning system based on a large model according to claim 3, characterized in that, The vehicle's motion trajectory features are extracted and calculated from the video stream through the following steps: For each frame of the image, perspective transformation and lane line detection are performed. A coordinate system with the lane center line as the reference is established, and the lateral pixel position X(t) of the vehicle in the image is determined by the vehicle detection box. Transform X(t) to a real-world approximate coordinate system with the lane center as the zero point to obtain the lateral offset sequence f(t); calculate the standard deviation of f(t) within the evaluation window, and the cumulative time percentage of Of(t) exceeding the preset safety threshold; After applying a low-pass filter to the horizontal position sequence X(t), the mean of the absolute value sequence of the first difference of X(t) is calculated as a quantitative indicator of trajectory jitter.

5. The driver status early warning system based on a large model according to claim 1, characterized in that, The system also includes a multi-level continuous state verification module; The multi-level continuous status verification module is activated after the central analysis module initially outputs a high-risk judgment, and performs first-level continuous verification, specifically as follows: Within the first verification time window after the initial judgment, the vehicle's motion trajectory characteristics are continuously analyzed; If, among the motion trajectory features, the standard deviation of the lateral offset sequence f(t), the cumulative time percentage exceeding the preset safety threshold, or the quantitative index of trajectory jitter continuously exceeds the corresponding threshold, then it is determined that the trajectory abnormality persists.

6. The driver status early warning system based on a large model according to claim 5, characterized in that, If the trajectory anomaly is determined to persist, the multi-level persistent state verification module performs a second-level deep verification, specifically: The edge computing module is instructed to increase the sampling frequency of the driver's facial features for the target vehicle and obtain updated microscopic action temporal features; Based on the updated micro-action timing characteristics, a comprehensive physiological damage persistence risk index is calculated. The calculation formula is as follows: ; ; in: and This refers to the start and end times of the secondary depth verification time window; for The aspect ratio of the pupil at any given moment. This is the reference value for the pupil aspect ratio of the driver under normal conditions; This represents the time integral of the positive deviation of the pupil aspect ratio from the normal baseline within the secondary depth verification time window; and These are the standard deviations of the head pitch angle and yaw angle within the secondary depth verification window, respectively. This means that the larger of the two absolute values ​​is taken as the representative value of the head's axial instability. and These are preset weighting coefficients; like If the risk of physiological damage exceeds the preset threshold, it is determined that the physiological damage state is ongoing.

7. The driver status early warning system based on a large model according to claim 6, characterized in that, The multi-level persistent state verification module also The process includes the following: When the Level 1 continuous verification determines that the trajectory abnormality persists and the Level 2 deep verification determines that the physiological damage state persists, the initial high-risk determination is confirmed to be valid, and the final risk confirmation level is upgraded. When only the first-level verification result is valid, the warning will be marked as requiring manual review; If both levels of verification results are invalid, the initial high-risk determination will be automatically revoked.

8. The driver status early warning system based on a large model according to claim 5, characterized in that, The duration of the first verification time window is dynamically adjusted based on the risk confidence level at the time of the initial judgment, the average vehicle speed on the current road segment, and the weather visibility.

9. The driver status early warning system based on a large model according to claim 1, characterized in that, The system also includes: The evidence generation module is used to simultaneously solidify key video clips containing the entire process of the behavior when automatically generating warning instructions. These clips are then bound to vehicle identity information, behavior type tags, time, and geographical location to generate a structured law enforcement evidence package.