Method and system for analyzing abnormal driving behavior data of driver
By identifying abnormal driver behavior through multi-source data fusion and hybrid analysis models, and combining environmental data for risk assessment and intervention strategy matching, this approach addresses the limitations of existing systems in identifying abnormal behavior and the inadequacy of intervention prompts, thereby improving driving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING JIUZHOU ANHUA INFORMATION SECURITY TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing driving safety systems suffer from limitations in identifying abnormal driver behavior, risk assessments that are detached from the specific scenario, and intervention prompts that lack specificity. They are also prone to misjudgment and incompatibility, making it difficult to cope with diverse safety hazards in complex road conditions.
By acquiring multi-source data (vehicle, driver, and environmental data), feature extraction and fusion are performed. A hybrid analysis model is used to identify abnormal behaviors and probabilities. Environmental data is combined to conduct risk assessment, analyze the causes, and match intervention strategies based on risk levels and causes to generate personalized prompts.
It enables accurate identification of abnormal driver behavior, scientific risk assessment, and targeted intervention prompts, thereby improving driving safety and reducing safety hazards.
Smart Images

Figure CN121901966A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of driving safety, and in particular to a method and system for analyzing abnormal driving behavior data of drivers. Background Technology
[0002] With the rapid development of intelligent connected vehicle technology, intelligent monitoring and intervention for driving safety has become a core industry requirement. Current traditional driving safety systems often rely on a single data dimension, resulting in issues such as incomplete abnormal behavior identification, risk assessment detached from specific scenarios, and a lack of targeted intervention prompts, making it difficult to address diverse safety hazards in complex road conditions. Furthermore, due to significant differences in individual driving habits, a one-size-fits-all monitoring approach can easily lead to misjudgments of abnormal behavior and inappropriate intervention strategies, impacting both the driving experience and the effectiveness of interventions.
[0003] Against this backdrop, there is an urgent need for a method and system for analyzing abnormal driving behavior data of drivers. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a method and system for analyzing abnormal driving behavior data of drivers.
[0005] A first aspect of this application provides a method for analyzing abnormal driving behavior data of a driver, including: Acquire multi-source data during the driver's driving process, wherein the multi-source data includes vehicle data, driver data, and environmental data; Feature extraction and fusion are performed on the multi-source data to obtain a fused feature vector; The fused feature vector is input into the hybrid analysis model, which outputs the driver's abnormal behavior and probability. Based on the probability of the abnormal behavior and the environmental data, a risk assessment is performed on the abnormal behavior to obtain a risk level; Based on the risk level, the fused feature vector, and the temporal context of the multi-source data, the abnormal behavior is analyzed to obtain the cause of the abnormal behavior; Based on the risk level and the trigger, a prompting strategy is matched from a preset intervention strategy library; Based on the aforementioned prompting strategy, corresponding prompting information is generated.
[0006] A second aspect of this application provides a driver's abnormal driving behavior data analysis system, including: The multi-source data acquisition module is used to acquire multi-source data during the driver's driving process, wherein the multi-source data includes vehicle data, driver data and environmental data; The feature extraction and fusion module is used to extract and fuse features from the multi-source data to obtain a fused feature vector. An abnormal behavior identification module is used to input the fused feature vector into a hybrid analysis model and output the abnormal behavior and probability of the driver. The risk level assessment module is used to assess the risk of the abnormal behavior based on the probability of the abnormal behavior and the environmental data, and obtain the risk level. An anomaly cause analysis module is used to analyze the abnormal behavior based on the risk level, the fused feature vector, and the temporal context of the multi-source data to obtain the cause of the abnormal behavior. An intervention strategy matching module is used to match prompt strategies from a preset intervention strategy library based on the risk level and the trigger. The prompt message generation module is used to generate corresponding prompt messages based on the prompt strategy.
[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for analyzing abnormal driving behavior data of a driver.
[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the driver's abnormal driving behavior data analysis system described above.
[0009] The beneficial effects of the driver's abnormal driving behavior data analysis method and system provided in this application are as follows: This application comprehensively collects multi-source data from the vehicle, driver, and environment, and forms a complete fusion feature vector based on feature extraction and fusion technology, avoiding misjudgments caused by a single data dimension. It accurately outputs abnormal behavior and probability using a hybrid analysis model, and then conducts risk assessment in conjunction with environmental data, ensuring that the risk level determination results are both consistent with the characteristics of the behavior itself and adaptable to actual driving scenarios, thus guaranteeing the scientific nature of risk classification. Based on risk level, fusion feature vector, and time-series context, it analyzes the abnormal causes from multiple dimensions, and matches prompt strategies through a preset intervention strategy library, then generates corresponding prompt information, realizing a closed-loop management of the entire process from data collection, abnormal identification, risk assessment, cause analysis to intervention prompts. The entire process takes into account the accuracy of abnormal identification, the rationality of risk assessment, and the targeting of intervention prompts. This application improves the accuracy of judging abnormal behavior during driving while reducing safety hazards. Attached Figure Description
[0010] Figure 1A flowchart illustrating a method for analyzing abnormal driving behavior data of a driver, provided in an embodiment of this application; Figure 2 A structural block diagram of a driver's abnormal driving behavior data analysis system provided in an embodiment of this application; Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] To make the purpose, technical solution, and advantages of this application clearer, the following will be described in conjunction with the appendix. Figure 1-3 The following is an explanation using specific examples.
[0013] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a method for analyzing abnormal driving behavior data of a driver according to an embodiment of this application. The method includes: S101: Acquire multi-source data during the driver's driving process, including vehicle data, driver data, and environmental data.
[0014] In this embodiment, vehicle data is collected through the on-board diagnostic system interface, on-board sensor network, and vehicle bus. The vehicle data includes real-time vehicle speed, acceleration, steering angle, brake pedal travel, engine speed, transmission gear, tire pressure data, light status, and airbag trigger signal. The data sampling frequency is set according to the importance of the parameters. For example, if it is necessary to capture instantaneous changes such as rapid acceleration, this parameter is a primary parameter, so the sampling frequency for real-time vehicle speed and acceleration is set to 80Hz. Since transmission gear and light status are switched less frequently, this parameter is a secondary parameter, so its sampling frequency is set to 10Hz. The normal tire pressure is 10Hz. When abnormal data is collected, the parameter corresponding to the abnormal data is automatically raised to 80Hz for recording.
[0015] Driver data is collected via biosensors on the steering wheel to capture physiological signals such as heart rate; cameras and computer vision algorithms built into the cockpit capture eyelid opening and closing, gaze direction, head posture, and facial expressions; and sensors acquire and record sound during the driving process.
[0016] Environmental data is acquired by fusing vehicle-mounted sensing devices with external data sources. For example, millimeter-wave radar and lidar collect distance and speed information of surrounding vehicles and pedestrians, forward-facing cameras identify traffic lights, lane lines and road signs, and vehicle-mounted weather sensors record real-time rainfall, visibility and ambient temperature. At the same time, high-precision maps and real-time traffic cloud platform data are accessed to obtain information such as road type, speed limit information, real-time traffic conditions and distribution of surrounding infrastructure, forming a spatiotemporally correlated environmental feature dataset.
[0017] S102: Perform feature extraction and fusion on multi-source data to obtain a fused feature vector.
[0018] In this embodiment, time-domain and frequency-domain feature extraction of multi-source data is first carried out: For vehicle data, time-domain statistics such as mean vehicle speed, variance, rate of change of acceleration, and distribution of brake pedal operation time intervals are calculated, and frequency-domain features such as engine speed main frequency components and transmission gear shift spectrum energy distribution are extracted through Fourier transform; For driver data, time-domain features such as heart rate fluctuation amplitude and time-series distribution of gaze deviation duration are extracted, as well as facial micro-expression frequency-domain features such as blink frequency spectrum analysis are extracted; For environmental data, time-domain features such as road curvature time-domain change trend and traffic flow time-segment distribution are extracted, and frequency-domain energy features of rainfall intensity are obtained through wavelet transform to form a multi-dimensional original feature vector.
[0019] Secondly, a three-dimensional weight mapping table is constructed to determine the basic weights. The weights are set according to the importance of data parameters, the stability of feature vectors, and the contribution of historical data. For example, the weights of core safety-related features such as vehicle speed and brake pedal travel are set to 0.8-0.9, and the weights of secondary features such as light status are set to 0.2-0.3. The weights are also updated regularly through a feature importance evaluation algorithm.
[0020] Finally, the driving scenario adjustment weights are combined. Specifically, based on the output of the scenario classification model, and according to the scenario-weight adjustment coefficient table, the adjustment coefficients for lane line recognition features in highway cruising scenarios and the adjustment coefficients for following distance attention in urban congestion scenarios are matched. The basic weights are multiplied by the adjustment coefficients to obtain the scenario-adaptive adjustment weights. After weighted calculation and fusion of features from various dimensions, principal component analysis is used to reduce dimensionality and retain more than 95% of the feature information, ultimately forming a fusion feature vector with unified dimensions that fits the current scenario.
[0021] S103: Input the fused feature vector into the hybrid analysis model and output the abnormal behavior and probability of the driver.
[0022] In this embodiment, before inputting the fused feature vectors into the hybrid analysis model, data standardization is required. The Z-Score standardization formula is used to eliminate the dimensional differences between features of different dimensions. For example, features such as vehicle speed (km / h) and eye-closing frequency (times / minute) are converted into values of the same order of magnitude to avoid affecting the model's calculation accuracy due to differences in feature scale. At the same time, missing values in the vectors are supplemented using interpolation, and outliers are smoothed using the 3σ principle. Finally, the processed 512-2048 dimension fused feature vectors are input into the analysis model in batches. The data flow pipeline of the TensorFlow or PyTorch framework is used to achieve efficient data loading to ensure the input stability during real-time model computation.
[0023] In this embodiment, the hybrid analysis model includes a base layer and an advanced layer. Specifically, the base layer first performs preliminary screening on the input fused feature vectors. Through feature importance evaluation using the random forest algorithm, it extracts features strongly correlated with single abnormal behaviors, such as the speed-to-speed limit difference being correlated with speeding behavior, and the standard deviation of steering wheel angle being correlated with sharp turning behavior. Subsequently, it uses an SVM classifier to construct multiple binary classification models, where each model corresponds to a simple abnormal behavior. By calculating the distance between the feature vector and the classification hyperplane, it preliminarily determines whether there are clearly defined abnormal behaviors such as speeding and sudden braking, and outputs a preliminary anomaly probability.
[0024] The advanced layer receives the feature filtering results from the base layer and the original fused feature vector. It captures the temporal dependencies of the feature vectors through the gating mechanism of the LSTM network. For example, it analyzes the changing trends of the frequency of eye closure and the steering wheel operation amplitude within 10 consecutive seconds. Secondly, based on the multi-head attention mechanism of the Transformer, it assigns higher attention weights to key features. Through the fully connected layer, it fuses the temporal features with the key features to build a hybrid classification model to identify complex abnormal behaviors such as fatigued driving and distracted driving, and outputs the probability value of the corresponding abnormal behavior.
[0025] After the hybrid classification model is completed, the outputs of the base layer and the advanced layer are first fused. A weighted voting method is used to comprehensively determine the type of abnormal behavior. For example, if the base layer determines the probability of speeding to be 0.6 and the advanced layer determines the probability of fatigued driving to be 0.8, then fatigued driving is finally prioritized. Subsequently, the output probabilities are calibrated using the Platt scaling method, converting the logarithmic probabilities of the original model output into confidence levels between 0 and 1 to ensure that the probability values are consistent with the credibility of the actual scenario. For example, the probability of distracted driving was 0.82 before calibration and was adjusted to 0.78 after processing. Finally, the results are output according to the abnormal behavior type-probability format, such as making or receiving a phone call -0.75, sudden acceleration and sudden braking -0.62. At the same time, the computation time and feature importance ranking of the hybrid classification model are recorded to provide data support for subsequent abnormal behavior cause analysis.
[0026] S104: Based on the probability of abnormal behavior and environmental data, conduct a risk assessment of the abnormal behavior to obtain a risk level.
[0027] In this embodiment, before risk assessment, it is necessary to calculate the environmental hazard coefficients of multiple factors. In addition to basic road type and weather conditions, road conditions are added: wet and slippery road surface, icy road surface, and dry road surface; lighting conditions: no lighting at night, strong light during the day, and normal lighting; special scene markers: construction area, tunnel entrance, and intersection. The factor values are updated by road surface humidity sensor and light sensor, rather than relying on fixed values.
[0028] Using the probability of abnormal behavior output by a hybrid analysis model for assessment is prone to bias and needs to be corrected by combining the duration of the behavior with historical data. For example, if an abnormal behavior lasts for more than 3 seconds, the original probability is multiplied by a duration coefficient of 1.2, where the longer the duration, the higher the risk. If the driver has exhibited the same abnormal behavior more than 5 times in the past hour, it is multiplied by a frequency coefficient of 1.1, as the risk of repeated behaviors is cumulative. The duration coefficient and frequency coefficient are empirical and quantitative coefficients preset based on historical driving safety data, statistical analysis of the correlation between abnormal behavior and accidents, and the risk characteristics of driving scenarios. At the same time, the probability value is calibrated by comparing the matching degree between the model output probability and the actual labeled samples. The calibration formula is probability = original probability × accuracy. For example, the original fatigue driving probability is 0.9, which is adjusted to 0.738 after accuracy calibration to avoid risk assessment bias caused by model misjudgment.
[0029] The calculated values are weighted based on the modified probability of abnormal behavior and the risk coefficients of the associated environmental factors. The resulting values are then normalized, where [0, 0.4) represents low risk; [0.4, 0.8) represents medium risk; and [0.8, 1.0] represents high risk.
[0030] S105: Based on risk level, fused feature vectors, and the temporal context of multi-source data, abnormal behavior is analyzed to obtain the causes of abnormal behavior.
[0031] In this embodiment, based on the time when the abnormal behavior occurs, a multi-source data time sequence segment of the first preset time window before transmission is extracted, and the feature dimensions strongly related to the abnormal behavior in the feature vector are synchronously associated and fused, such as the sudden increase in heart rate and vehicle speed fluctuation before the abnormality occurs, to form a time sequence context including the vehicle operation sequence aligned with timestamps, the driver's state change curve, and the evolution trajectory of environmental parameters.
[0032] Secondly, a causal association mining mechanism is established. Specifically, the temporal context data is input into the improved causal discovery algorithm to calculate the conditional independence between abnormal behavior and each feature variable, and generate a weighted causal network graph, such as the causal path of strong light irradiation → line deviation → lane crossing. The weight represents the influence intensity, and the association strength threshold is set according to the risk level. For example, the threshold for high-risk behavior is set to 0.7, and the threshold for medium and low risk is set to 0.5, and the core association features are screened out.
[0033] The selected core correlation features are then matched against a pre-defined trigger database. Features with a match rate of ≥80% are directly identified as the corresponding triggers. For features with low match rates, the data is input into a BERT-based text inference model. This model, based on the semantic description of structured data, outputs the trigger type for the abnormal behavior, such as sudden braking caused by a pedestrian crossing suddenly. Cross-validation using a historical case database ensures the consistency of the trigger explanation. The pre-defined trigger database includes driver factors (fatigue or distraction), vehicle factors (braking delay), and environmental factors (sudden obstacles), etc.
[0034] S106: Based on the risk level and the trigger, match the prompt strategy from the preset intervention strategy library.
[0035] In this embodiment, firstly, a multi-level intervention strategy library is constructed based on the risk intervention characteristics of historical data. For example, a two-dimensional strategy matrix is established according to the risk level and the type of trigger. Then, prompting strategies are matched according to the intervention strategy library. For example, high risk corresponds to emergency intervention, medium risk corresponds to enhanced prompts, and low risk corresponds to gentle reminders. At the same time, targeted words are matched for each trigger. For example, fatigue trigger corresponds to suggesting to stop and rest, and strong light trigger corresponds to the automatic adjustment of the rearview mirror brightness.
[0036] When multiple intervention strategies are matched at the same time, they are sorted according to risk level > urgency of the trigger > historical effectiveness. Specifically, high-risk strategies are executed before low-risk strategies, and collision warning trigger strategies are executed before comfort reminders. At the same time, the driver's historical feedback data is called to adjust the prompt strategy, and the final prompt strategy is output.
[0037] S107: Generate corresponding prompt information based on the prompt strategy.
[0038] In this embodiment, the core parameters of the prompting strategy are extracted and personalized requirements are determined. For example, key elements are decomposed from the matched prompting strategy, including prompting modality (visual / auditory); information intensity (volume level, light brightness); and core content direction. Simultaneously, the driver's temporary personalized driving profile is retrieved to obtain their preference for prompting methods, commonly used language types, and past responses to different prompts, thus determining the basic framework of the prompts. Next, multimodal information is integrated to generate initial prompting content. Specifically, for visual prompts, differentiated interface elements are designed according to risk levels. For example, a red dynamically flashing icon is used for high-risk situations, a yellow static icon for medium-risk situations, and blue text for low-risk situations. For auditory prompts, a dedicated voice script is matched based on the trigger. For example, if strong light causes visual deviation, the voice content is: "Strong light ahead is affecting visibility; please turn on anti-glare mode, stay in your lane," and the speech rate and volume are adjusted according to preference, forming a multimodal collaborative initial prompt message.
[0039] Finally, the details of the prompts are optimized based on the scenario. Specifically, dynamic data of the current vehicle environment is obtained, and the initial prompts are adjusted according to the scenario. For example, when in a tunnel, the brightness of the visual prompts is reduced to avoid glare, while the volume of the auditory prompts is increased. When driving at high speed, the text prompts are simplified and short voice sentences are used to avoid distraction. In the end, prompts that conform to the real-time scenario, adapt to the driver's habits, and are highly matched with the prompt strategy are generated.
[0040] As can be seen from the above, this application comprehensively collects multi-source data from the vehicle, driver, and environment, and forms a complete fusion feature vector based on feature extraction and fusion technology, avoiding misjudgments caused by a single data dimension. By using a hybrid analysis model to accurately output abnormal behavior and probability, and then combining it with environmental data to conduct risk assessment, the risk level determination results are both consistent with the characteristics of the behavior itself and adapted to actual driving scenarios, ensuring the scientific nature of risk classification. Based on risk level, fusion feature vector, and time-series context, the application analyzes the causes of abnormalities from multiple dimensions, and matches prompt strategies through a preset intervention strategy library, then generates corresponding prompt information, realizing a closed-loop management of the entire process from data collection, anomaly identification, risk assessment, cause analysis to intervention prompts. The entire process takes into account the accuracy of anomaly identification, the rationality of risk assessment, and the targeting of intervention prompts. This application improves the accuracy of judging abnormal behavior during driving while reducing safety hazards.
[0041] In one embodiment of this application, feature extraction and fusion of multi-source data are performed to obtain a fused feature vector, including: The time-domain and frequency-domain features of multi-source data are extracted to obtain multi-source feature vectors; Based on a preset weight mapping table, weights are matched for multi-source feature vectors to obtain the basic weights of the multi-source feature vectors. The basic weights are adjusted based on the driving scenario to obtain the adjusted weights; The multi-source feature vectors are weighted and fused based on the adjusted weights to obtain the fused feature vector.
[0042] In this embodiment, firstly, the temporal and frequency domain features of the multi-source data are extracted. For example, for vehicle data, time-domain statistics such as mean and variance of vehicle speed, rate of change of acceleration, and time interval distribution of brake pedal operation are calculated. Frequency domain features such as the main frequency component of engine speed and the spectral energy distribution of gear shifting are extracted using Fourier transform. Simultaneously, for driver data, temporal features such as the temporal distribution of heart rate fluctuation amplitude and gaze deviation duration are extracted, along with frequency domain features of facial micro-expressions, such as the spectral analysis of blink frequency. Subsequently, temporal features such as the temporal trend of road curvature and the time-period distribution of traffic flow are calculated from the road environment data. Frequency domain energy features of rainfall intensity are extracted using wavelet transform, forming a multi-dimensional multi-source feature vector.
[0043] Secondly, a weight mapping table is constructed to obtain the basic weights for multi-source feature vector fusion. Specifically, a three-dimensional weight mapping table is constructed based on the importance of data parameters, the stability of feature vectors, and the contribution of historical data. For example, the basic weights of features directly related to driving safety, such as vehicle speed and brake pedal travel, are set to 0.8-0.9, while the weights of secondary parameters, such as lighting status features, are set to 0.2-0.3. The mapping table is updated periodically through a feature importance evaluation algorithm to match the corresponding basic weight value for each feature in the multi-source feature vector.
[0044] This embodiment adjusts the weights based on driving scenarios. Specifically, it uses the scenario type output by the scenario classification model to retrieve a preset scenario-weight adjustment coefficient matrix. For example, in a highway cruising scenario, the adjustment coefficient for lane line recognition features in the environment is increased; in an urban congestion scenario, the adjustment coefficient for the driver's following distance attention feature is increased. The adjusted weights are then multiplied by the corresponding scenario's adjustment coefficient to obtain the weights adapted to the current driving scenario. Next, the adjusted weights are weighted and calculated with the features of each dimension in the multi-source feature vector, fusing the feature vector. Principal component analysis is then used to reduce the dimensionality of the weighted high-dimensional features, retaining more than 95% of the feature information, ultimately forming a fused feature vector with unified dimensions and scenario adaptation.
[0045] As can be seen from the above, this embodiment extracts the time-domain and frequency-domain features of multi-source data, adjusts the weights according to preset weights and driving scenarios, and then fuses them with the adjusted vector of multi-source data. This preserves the multi-dimensional features of the data while avoiding feature homogenization through scenario-based weight adjustment, making the fused feature vector more consistent with actual driving scenarios and improving the accuracy of the input data for subsequent hybrid analysis models.
[0046] In one embodiment of this application, adjusting the base weights based on the driving scenario to obtain adjusted weights includes: By inputting environmental data and vehicle data into the scene classification model, the driving scene type can be obtained. Among them, the driving scenario types include at least: highway cruising, urban congestion, following another vehicle on a ring road, passing through an intersection, and parking. Based on the driving scenario and the preset scenario-weight adjustment coefficient table, match the adjustment coefficients of each feature in the corresponding multi-source data; Adjusted weights are obtained by adjusting the basic weights based on the adjustment coefficient.
[0047] In this embodiment, a Transformer-based deep learning model is used as the core of scene classification. The road type, traffic flow, and speed limit information in the environmental data are timestamped with the real-time vehicle speed, steering angle, driving trajectory and other features in the vehicle data. After extracting the spatiotemporal correlation features through a sliding window, the driving scene type is input. The driving scene type can be trained to output scene types such as highway cruising, urban congestion, ring road following, intersection passing, and parking conditions. The scene with the highest probability value is taken as the current driving scene type.
[0048] Secondly, feature adjustment rules are formulated for different types of driving scenarios. For example, in the high-speed cruise scenario, the adjustment coefficients for features such as lane line integrity and relative speed of surrounding vehicles in the environmental data are set to 1.2-1.5, and the adjustment coefficients for vehicle speed stability and steering smoothness in the vehicle data are set to 1.1. In the urban congestion scenario, the adjustment weights of the adjustment coefficients for the vehicle distance change rate in the environmental data and the adjustment coefficients for the following reaction time in the driver data are increased, forming a structured scenario-weight adjustment coefficient table, and the full adjustment coefficients for the corresponding scenario are automatically matched according to the classification results.
[0049] This embodiment adjusts and updates the weight of each feature in the multi-source feature vector by calculating the adjustment weight as the base weight multiplied by the scene adjustment coefficient. For example, the base weight of a certain vehicle acceleration feature is 0.6, and the adjustment coefficient in the loop following scenario is 1.2. Then its adjustment weight is 0.6 × 1.2 = 0.72, thus forming an adjustment weight set that is highly matched with the current scenario.
[0050] As can be seen from the above, this embodiment inputs environmental data and vehicle data into the scenario classification model to determine the scenario type. Then, it matches the corresponding adjustment coefficients to adjust the weights for each scenario type, making the fusion of multi-source feature vectors more suitable for specific driving scenarios. This solves the problem of the difference in feature importance under different scenarios, while improving the adaptability and accuracy of abnormal behavior analysis in diverse scenarios.
[0051] In one embodiment of this application, a method for analyzing abnormal driving behavior data of a driver further includes: An LSTM-based anomaly detection model is used to identify emergency events in multi-source data. When an emergency is identified, temporary adjustment instructions are generated based on the severity of the emergency. Based on temporary adjustment instructions, an attention mechanism is used to calculate the correlation between each data feature and the emergency event, and a correlation adjustment coefficient is generated. The correlation moderating coefficient and the moderating coefficient are combined to obtain the comprehensive moderating coefficient; The basic weights are adjusted based on a comprehensive adjustment coefficient to obtain the final weights applicable to emergency events.
[0052] In this embodiment, emergency event identification is first achieved using an LSTM-based anomaly event detection model. Specifically, multi-source data is segmented into time-series segments according to a preset sliding window and input into a pre-trained LSTM-based anomaly event detection model. The model identifies emergency events and outputs results, including forward collision risk and vehicle malfunction, along with a severity score corresponding to the emergency event. The severity score ranges from 0 to 10, with higher scores indicating a more dangerous emergency event, and 10 being the highest risk. At this point, the initial identification and classification of the emergency event is completed. This anomaly event detection model learns the characteristics of historical emergency events to capture abnormal fluctuations in the data, such as sudden drops in vehicle speed, sharp steering wheel turns, and millimeter-wave radar detecting nearby obstacles. Furthermore, it generates temporary adjustment commands based on the severity of the emergency event.
[0053] Specifically, a severity-instruction mapping rule is preset. For example, if the severity score is 8-10, indicating an impending collision, a high-priority temporary adjustment instruction is generated, requiring a focus on increasing the weight of environmental perception and vehicle operation-related features; if the severity score is 5-7, indicating sudden deceleration while following a vehicle at close range, a medium-priority temporary adjustment instruction is generated, focusing on optimizing the weight of features related to distance monitoring and driver reaction; if the severity score is 3-4, indicating minor road obstacles, a low-priority temporary adjustment instruction is generated, requiring fine-tuning the weight of road condition features in the environmental data to ensure the accuracy of matching the instruction with the urgency of the event. Subsequently, guided by the core focus dimensions in the temporary adjustment instructions, including high-priority instruction focus, obstacle distance, and braking response, an attention weight calculation network is constructed. This network performs correlation analysis between multi-source data features and emergency event features. For example, in a forward collision risk event, the real-time obstacle distance feature of the environmental data has a high correlation with the event, and the correlation adjustment coefficient is calculated to be 0.9 through the attention mechanism. The correlation coefficient of the "brake pedal response speed" feature of the vehicle data is 0.8, while the coefficient of non-core features such as "light status" is only 0.2, forming a set of correlation adjustment coefficients for all feature dimensions.
[0054] Finally, a weighted fusion formula is adopted: Comprehensive adjustment coefficient = Original scenario adjustment coefficient × 0.6 + Correlation adjustment coefficient × 0.4. The weights can be adjusted according to the severity of the event. The weight of the correlation coefficient is also increased in high-severity events. For example, in the high-speed cruise scenario, the original adjustment coefficient for the relative speed of surrounding vehicles is 1.3, while the correlation adjustment coefficient in the emergency event is 0.9. Substituting these values into the formula, the comprehensive adjustment coefficient is calculated as 1.3 × 0.6 + 0.9 × 0.4 = 1.08, achieving a scientific fusion of the two coefficients and taking into account both scenario adaptability and the specificity of the emergency event.
[0055] In this embodiment, the basic weights of the multi-source feature vectors are retrieved and calculated using the formula "adjustment weight = basic weight × comprehensive adjustment coefficient". For example, the basic weight of the "real-time distance to obstacles" feature is 0.7, and the comprehensive adjustment coefficient is 1.08, so its final weight is 0.7 × 1.08 = 0.756. At the same time, the final weight set is normalized to ensure that the sum of all feature weights is 1, forming weights that are adapted to the current emergency event and used for subsequent feature fusion.
[0056] As can be seen from the above, this embodiment identifies and generates temporary adjustment instructions through the LSTM model when an emergency occurs, and adjusts the weights according to the attention mechanism. This ensures that feature vectors highly correlated with the emergency are given priority in emergency situations, avoiding interference from conventional weights in emergency analysis and improving the timeliness and accuracy of abnormal behavior identification in emergency scenarios.
[0057] In one embodiment of this application, after inputting the fused feature vector into the hybrid analysis model and outputting the driver's abnormal behavior and probability, the method further includes: Create temporary personalized driving profiles for drivers, which are generated based on anonymized driving data; The inherent driving habits in the temporary personalized driving profile are assessed for normality, and normal driving habits are selected. The similarity between abnormal behavior and normal driving habits is calculated to obtain a behavior similarity score; Personalized correction coefficients are generated based on behavioral similarity scores using a logistic regression model. The probability output by the hybrid analysis model is corrected based on the personalized correction coefficient.
[0058] In this embodiment, the anonymized driving data collected from the driver over a continuous period is the driver's driving data with identity information removed. Driving behavior is extracted using a temporal clustering algorithm. Driving behavior includes regular speed ranges, such as a preference for 50-60 km / h on urban roads, brake pedal force distribution, such as the threshold of force for habitually light braking, and the matching relationship between steering angle and curve curvature. Based on driving time preferences, such as the frequency of travel during morning rush hour and the characteristics of frequently traveled routes, such as a preference for main roads or auxiliary roads, a temporary personalized driving profile is formed. This personalized temporary profile data is updated once every period.
[0059] Secondly, the normality of inherent driving habits is assessed. Specifically, driving habits in personalized temporary files are compared with traffic regulations and the baseline for safe driving in a group. Z-score standardization is used to calculate the deviation. Habits with a deviation ≤1.2 and that have not triggered safety warnings are considered normal driving habits. Habits with a deviation greater than 1.2 or associated with abnormal events are marked as unobservable, such as continuous lane changes without using turn signals. Normal driving habits are then selected as the comparison benchmark. Subsequently, the similarity between abnormal behavior and normal habits is calculated: from three dimensions—operational characteristics, environmental adaptability, and physiological response—a dynamic time warping algorithm is used to calculate the temporal similarity between the current abnormal behavior and normal driving habits. For example, the waveform of "sudden braking behavior" is compared with the "smooth deceleration mode" in normal habits, outputting a similarity score of 0-1, where 1 is a perfect match and 0 is a complete deviation. Operational characteristics include braking timing and steering amplitude; environmental adaptability includes speed adjustment strategies under different road conditions and road type adaptation behaviors; physiological responses include heart rate variability and head posture under different driving states.
[0060] Finally, the similarity score, the category of the abnormal behavior (e.g., operational error or abnormal state), and the driver's historical correction records are used as input features and substituted into a pre-trained logistic regression model. This model has been validated with data from over 5000 drivers and has an accuracy rate of 89%. The model outputs a personalized correction coefficient of 0.5-1.5. Specifically, when the similarity score is ≥0.8, it indicates that although the current behavior is initially judged as abnormal, it is actually a normal habit of the driver, such as an experienced driver's habit of quickly changing lanes. In this case, the coefficient is ≤0.7, which can reduce the initial probability of abnormality. For example, if the probability of this behavior being judged as abnormal is 80%, it can be reduced to 56% after this correction to avoid misjudgment of abnormal behavior. When the score is ≤0.3, the behavior deviates significantly from the driver's normal habits, such as a sudden emergency braking during normal driving. The coefficient is ≥1.2, which increases the initial probability of abnormality. For example, if the probability of this behavior being judged as abnormal is 60%, it can be increased to 72% after this correction.
[0061] As can be seen from the above, this embodiment constructs a temporary personalized driving profile and judges abnormal behavior based on the driver's inherent habits. By taking the differences in individual driver habits into account in the data analysis, it avoids misjudging personalized normal behavior as abnormal behavior. Furthermore, it generates a personalized correction coefficient through logistic regression, and uses this personalized correction coefficient to correct the probability of abnormal behavior output by the hybrid analysis model, thereby improving the personalization and accuracy of abnormal behavior judgment and reducing interference caused by misjudgments.
[0062] In one embodiment of this application, a personalized correction coefficient is generated based on a behavioral similarity score using a logistic regression model, including: Extract the frequency, intensity, and context of inherent behaviors from temporary personalized driving profiles during historical driving to construct a historical behavior baseline; Calculate the multidimensional similarity between the current abnormal behavior and the historical behavior baseline, including temporal similarity, intensity similarity, and scene similarity; A weighted sum of multidimensional similarity is used to calculate the behavioral consistency index; The behavioral consistency index is input into a pre-trained logistic regression model to generate personalized correction coefficients.
[0063] In this embodiment, based on the core dimensions of inherent behaviors in the temporary personalized driving profile, the core grayscale includes frequency, intensity, and contextual scenario dimensions. The frequency dimension counts the average daily occurrence and peak-hour percentage of inherent behaviors over the past month; inherent behaviors include overtaking on highways and low-speed following on urban roads. The intensity dimension records the extreme values and average intensity of parameters when inherent behaviors are executed; extreme values include the maximum acceleration during overtaking and the minimum safe following distance. The contextual scenario dimension annotates the environmental characteristics and vehicle status when inherent behaviors occur, and stores these in a time-series database according to inherent behavior type, forming a historical behavior baseline with three-dimensional features including frequency, intensity, and scenario. Each baseline dimension includes the data collection time range and sample size annotation. Environmental characteristics include weather, road type, and traffic flow; vehicle status includes speed range and vehicle gear.
[0064] Secondly, the multidimensional similarity between the current abnormal behavior and the historical behavior baseline is calculated. The multidimensional similarity includes: time similarity, intensity similarity and scene similarity.
[0065] Specifically, temporal similarity is calculated using a dynamic time warping algorithm, comparing the occurrence time distribution of the current abnormal behavior with that of similar behaviors in the historical baseline. For example, if both occur during the morning rush hour, the temporal similarity between the current abnormal behavior and the historical behavior baseline is high, resulting in a temporal similarity score of 0-1. Intensity similarity is calculated using Euclidean distance to determine the parameters of the current abnormal behavior. For instance, the deviation rate between the deceleration during sudden braking and the average intensity of the historical baseline is calculated. If the deviation rate is ≤10%, the intensity similarity is above 0.9; if the deviation rate is >30%, the similarity drops below 0.3. Scene similarity is calculated using a cosine similarity algorithm, matching the current scene feature vector with the historical baseline scene vector. When the overlap of scene elements is ≥80%, the scene similarity is above 0.8; when the overlap of scene elements is <50%, the scene similarity is <0.5, ultimately forming independent similarity scores across three dimensions. For example, in a driver's temporary personalized driving profile, the historical baseline intensity of sudden braking behavior in a rainy day + urban congestion + 20-30 km / h scenario is 4.2 m / s², which is the average deceleration of 15 sudden braking events in this scenario over the past 30 days. However, if the current driver experiences sudden braking in the same scenario, the measured deceleration is 4.5 m / s². The deviation rate is calculated as |4.5 - 4.2| / 4.2 × 100% ≈ 7.1%. Since the deviation rate is ≤10%, the intensity similarity score is 0. The similarity score is above 0.93. In terms of scene dimension, the current scene feature vector completely overlaps with the historical baseline scene vector, and the scene element overlap is between 80% and 100%, with a scene similarity score of 0.85. In terms of time dimension, if the sudden braking occurs at 8:00 AM during the morning rush hour, it matches the time distribution of 70% of sudden braking in this scene during the morning rush hour in the historical baseline, with a time similarity score of 0.82. Finally, three independent similarity scores are formed: time 0.82, intensity 0.93, and scene 0.85.
[0066] Finally, the consistency index is calculated using the formula: Behavioral Consistency Index = Time Similarity × Time Weight + Intensity Similarity × Intensity Weight + Scene Similarity × Scene Weight. For example, if a certain overtaking behavior has a time similarity of 0.8 and a weight of 0.3, an intensity similarity of 0.9 and a weight of 0.4, and a scene similarity of 0.7 and a weight of 0.3, then the consistency index is 0.8 × 0.3 + 0.9 × 0.4 + 0.7 × 0.3 = 0.81. The index range is controlled between 0 and 1, and the closer it is to 1, the more consistent the current behavior is with historical habits. Subsequently, the behavioral consistency index, driver behavior stability labels in the temporary personalized profile, and historical anomaly judgment correction records are used as joint input features and substituted into the pre-trained logistic regression model. This model learns the mapping relationship between the consistency index and correction coefficient of a large number of drivers and outputs a personalized correction coefficient of 0.5-1.5. For example, when the consistency index is ≥0.8, that is, when the current abnormal behavior is highly consistent with historical habits, a personalized correction coefficient of 0.5-0.7 is output to reduce the probability of false positives. When the consistency index is ≤0.3, that is, when the behavior seriously deviates from historical habits, a personalized correction coefficient of 1.3-1.5 is output to strengthen anomaly identification and ensure the reliability of the correction results.
[0067] As can be seen from the above, this embodiment extracts the frequency and intensity of inherent behaviors during historical driving, constructs a formal behavioral baseline based on this, calculates multidimensional similarity, and generates correction coefficients. By measuring behavioral consistency from multiple dimensions such as time, intensity, and scenario, the personalized correction coefficients better match the driver's actual driving habits, further improving the accuracy of correcting the probability of abnormal behavior.
[0068] In one embodiment of this application, a method for analyzing abnormal driving behavior data of a driver further includes: Calculate the completeness and statistical significance of the historical behavior baseline in the current triggering scenario; The completeness index and the statistical significance index are compared with preset standard thresholds, which include the completeness index threshold and the statistical significance index threshold. If both indicators are greater than or equal to the standard threshold, the confidence level of the personalized correction coefficient is deemed sufficient, and the personalized correction coefficient is used to correct the probability output by the mixed analysis model. If any indicator is less than the standard threshold, the confidence level of the personalized correction coefficient is deemed insufficient. The absolute differences between the completeness index and the statistical significance index and the standard threshold are then calculated, and the attenuation factor is determined. The personalized correction coefficient is adjusted based on the attenuation factor.
[0069] In this embodiment, the completeness index and statistical significance index of the historical behavior baseline under the current triggering scenario are calculated. Specifically, the completeness index is calculated by dividing the amount of valid historical data under the current scenario by the minimum amount of data required for the scenario. For example, in the current triggering scenario of rainy day + urban congestion, the historical baseline contains 12 valid emergency braking data points for this scenario, while the minimum data point required for the scenario is 10, resulting in a completeness index of 1.2. The statistical significance index is calculated using a t-test. Specifically, driving behavior parameters under the current scenario, such as emergency braking deceleration, are extracted from the historical baseline as a sample dataset. Assuming that the overall data follows a normal distribution, a null hypothesis is established that the sample mean is not significantly different from the theoretical mean. The t-statistic is calculated, and the corresponding p-value is obtained from the t-distribution table or statistical software. The smaller the p-value, the stronger the evidence to reject the null hypothesis, i.e., the better the centralization and representativeness of the sample data. If the p-value ≤ 0.05, it represents a typical significance level, and the statistical significance index is ≥ 0.8; if the p-value > 0.1, the statistical significance index is < 0.6, forming two quantitative evaluation indicators.
[0070] In this embodiment, standard thresholds are set and indicators are compared: the preset completeness index threshold is based on the amount of historical data that meets the minimum requirements for scenario analysis, and the statistical significance index threshold is based on the statistical significance of the historical data distribution; then, the calculated completeness index and statistical significance index are compared with the corresponding standard thresholds one by one to determine whether the two indicators meet the preset standards.
[0071] For example, if the completeness index is greater than or equal to the completeness index threshold, and the statistical significance index is greater than or equal to the statistical significance index threshold, it means that the historical baseline data is sufficient to support personalized correction, and the confidence level of the personalized correction coefficient is determined to be sufficient. The corresponding personalized correction coefficient is then directly substituted into the hybrid analysis model to correct the initial anomaly probability. If any index fails to reach the threshold, such as the completeness index being less than the completeness index threshold or the statistical significance index being less than the statistical significance index threshold, the confidence level of the corresponding personalized correction coefficient is determined to be insufficient, and the personalized correction coefficient adjustment process is initiated.
[0072] Specifically, for personalized correction coefficients with insufficient confidence, the differences between the two indicators and their corresponding standard thresholds are calculated separately. For example, the absolute difference between the completeness index of 0.8 and the completeness index threshold of 1.0 is 0.2, and the absolute difference between the statistical significance index of 0.6 and the statistical significance index threshold of 0.7 is 0.1. Secondly, the attenuation factor formula is used to calculate the attenuation factor: attenuation factor = 1 - (mean of the two absolute differences × 0.5). The personalized correction coefficient is then adjusted using the attenuation factor, i.e., the adjusted correction coefficient is the product of the personalized correction coefficient and the attenuation factor.
[0073] As can be seen from the above, this embodiment assesses the completeness and significance of historical data to determine and adjust the confidence level of the correction coefficient. This avoids the correction coefficient from becoming ineffective due to insufficient or insignificant historical data. By adjusting the personalized correction coefficient through a decay factor, the reliability of abnormal behavior probability correction is improved, ensuring that the personalized correction mechanism can function effectively under different data conditions.
[0074] In one embodiment of this application, a temporary personalized driving profile is constructed for the driver, including: Video data from driver data, and preprocessing the video data; Extracting driver behavior features from video data using computer vision models; Establish a temporal correlation between driver behavior characteristics and vehicle operation data and environmental data to form a sequence of driving habit characteristics; Pattern recognition and classification are performed on the driving habit feature sequence to obtain classification results, which include inherent driving habits that belong to abnormal behavior and inherent driving habits that do not belong to abnormal behavior. A temporary personalized driving profile is generated based on inherent driving habits that do not constitute abnormal behavior.
[0075] In this embodiment, target video data is obtained by preprocessing the driver's video data. The target video frames are then input into a pre-trained computer vision model, which can identify and extract core behavioral features. These core behavioral features include facial features, limb features, and visual features. Facial features include blinking frequency, pupil diameter changes, the degree of upturn / downturn of the corners of the mouth, and head turning angle. Limb features include the position of the hands on the steering wheel, the amplitude of arm swing, and the degree of forward / backward lean. Visual features include the area of focus of the gaze and the percentage of time the gaze deviates from the road conditions, forming a structured set of driver behavioral features.
[0076] Secondly, based on timestamps, the extracted driver behavior features are aligned with synchronously collected vehicle operation data (such as vehicle speed, steering angle, and brake pedal travel) and environmental data (such as road type, traffic flow, and weather conditions) to construct a three-dimensional temporal data chain of behavior features, vehicle operation, and environmental conditions. For example, the behavior feature of "head turning 30° + gaze deviating from road conditions for 2 seconds" is bound to vehicle and environmental data of the same period at a speed of 60 km / h + urban road + moderate traffic flow, and then linked in chronological order to form a continuous sequence of driving habit features.
[0077] Finally, an LSTM-based temporal pattern recognition model is employed. Inputting a sequence of driving habit features, the model learns the changing patterns and associations of the feature sequences, and, combined with pre-defined abnormal behavior judgment rules, classifies the driving habits. The classification results clearly categorize them into two types: one is inherent driving habits that constitute abnormal behavior; the other is inherent driving habits that do not constitute abnormal behavior.
[0078] This embodiment will filter out inherent driving habits that do not belong to abnormal behavior and integrate their core features. The core features of inherent driving habits include regular behavior patterns: average blinking frequency and common steering wheel grip posture during daily driving; scenario-adaptive behaviors: the proportion of eye focus time in high-speed scenarios and following operation habits in congested scenarios; operation-related features: the matching relationship between facial tension and braking force when braking. Based on the structure of behavior type - feature parameters - occurrence frequency - scenario adaptability, a temporary personalized driving profile including timestamps and data source annotations will be generated.
[0079] As can be seen from the above, this embodiment extracts driver behavior features through video data preprocessing and computer vision models, and constructs a time-series correlated driving habit feature sequence by combining vehicle operation data and environmental data. This not only achieves multi-dimensional and dynamic capture of the driver's driving habits, but also clearly distinguishes between abnormal behavior and non-abnormal inherent driving habits through identification and classification. This ensures that the temporary personalized driving profile only includes compliant and stable individual driving features, avoiding interference from abnormal behavior. At the same time, this temporary personalized profile is generated based on time-series correlated data and can represent the driver's operating patterns in different scenarios, providing reliable data support for subsequent personalized correction coefficient calculation and abnormal behavior judgment. This improves the pertinence and accuracy of driving behavior analysis and reduces the misjudgment problem caused by a one-size-fits-all approach.
[0080] In one embodiment of this application, abnormal behavior is analyzed based on risk level, fused feature vector, and temporal context of multi-source data to obtain the causes of the abnormal behavior, including: Extract the temporal context data of multi-source data within the first preset time window before the occurrence of abnormal behavior; The temporal context data is input into the causal discovery algorithm to obtain a multivariate causal graph; Multivariate causal graphs connect different data feature nodes through directed edges and assign a quantified causal relationship strength value to each directed edge; A first threshold is set based on the risk level; From the multivariate causal graph, the causal features corresponding to the directed edges that are connected to the abnormal behavior and whose causal relationship strength is greater than or equal to the first threshold are selected as candidate causal features. Match candidate trigger features with a pre-defined trigger knowledge base; If a match is found, the matched trigger type will be used as the trigger for the abnormal behavior. If a match fails, the corresponding candidate cause is input into a preset cause inference model for inference to obtain the cause type of abnormal behavior.
[0081] In this embodiment, firstly, the temporal context data of multi-source data prior to the occurrence of abnormal behavior is extracted. Specifically, taking the moment when the abnormal behavior is triggered as the endpoint, multi-source data within a first preset time window is extracted, including time-series data such as vehicle speed change curves, driver heart rate fluctuation data, obstacle trajectories in the environment, and lane line changes during this period, forming complete temporal context data of multi-source data to provide data support for causal analysis.
[0082] Secondly, a multivariate causal graph is generated using a causal discovery algorithm. For example, multi-source temporal context data is input into a causal Bayesian network algorithm. This algorithm analyzes the conditional independence between features of the multi-source temporal context data and identifies causal relationships between variables. For instance, nodes represent different data features, and directed edges represent the causal direction between variables. For example, "strong light → line of sight deviation" indicates that strong light caused line of sight deviation. Each directed edge is assigned a causal strength value of 0-1 through probability calculation, and finally, a structured multivariate causal graph is output. A higher causal strength value indicates a stronger causal relationship; nodes include strong light, line of sight deviation, and sudden drop in vehicle speed, etc.
[0083] In this embodiment, a first threshold is set according to the risk level of abnormal behavior. For example, high-risk behavior is set to 0.7, medium-risk behavior is set to 0.6, and low-risk behavior is set to 0.5. The core cause features are specifically selected from the multivariate causal graph. Only the features corresponding to the directed edges that are directly connected to the abnormal behavior node and whose causal relationship strength is greater than or equal to the first threshold are retained. For example, when the abnormal behavior is lane crossing, features such as line deviation with a causal strength of 0.75 and lane line blur with a causal strength of 0.68 are selected to form candidate cause features.
[0084] Secondly, the candidate causal features are precisely matched with a preset causal knowledge base. If a candidate feature completely matches an entry in the knowledge base, the entry is directly determined to be the causal factor for the abnormal behavior. For example, after a driver exhibits the abnormal behavior of crossing the lane line, the extracted core candidate feature is that an oncoming vehicle turns on its high beams, i.e., strong light → driver's gaze deviates from the road for 2 seconds → failure to correct direction in time, resulting in crossing the lane line. This candidate feature completely matches the entry in the knowledge base in terms of causal logic and key elements, so this candidate feature is directly determined to be the cause of the abnormal behavior, i.e., strong light causing gaze deviation. If no corresponding entry is found in the knowledge base, the candidate feature is input into a BERT-based causal reasoning model. This causal reasoning model outputs the type of causal factor that leads to the abnormal behavior based on historical cases and feature semantic analysis. The causal knowledge base includes three major categories: driver factors, vehicle factors, and environmental factors, such as fatigue driving, braking delay, and sudden obstacles.
[0085] As can be seen from the above, this embodiment extracts temporal context data within a first preset time window before the occurrence of abnormal behavior and generates a multivariate causal graph with quantified causal strength based on the causal discovery algorithm. This captures the dynamic causal relationship between abnormal behavior and multi-source data features, avoiding the one-sidedness of isolated analysis of single data. By screening highly correlated candidate trigger features through risk level adaptation to the first threshold, the core influencing factors can be focused on, improving the efficiency of trigger identification. Furthermore, the dual mechanism of matching with the preset trigger knowledge base and supplementing reasoning with the reasoning model ensures both the rapid determination of known triggers and the identification of triggers in unknown or complex scenarios, improving the utilization effect of trigger data and the accuracy and comprehensiveness of trigger identification.
[0086] In one embodiment of this application, if a match fails, the candidate trigger is input into a preset trigger inference model for inference to obtain the trigger type of the abnormal behavior, including: By inputting the risk level of candidate triggers, fused feature vectors, and multi-source data temporal context into the trigger inference model, the trigger type of abnormal behavior can be obtained. The type of trigger output by the trigger reasoning model is used as the trigger for abnormal behavior.
[0087] In this embodiment, the risk levels corresponding to candidate triggers are extracted, including high-risk, medium-risk, and low-risk levels, and quantified into numerical labels for risk levels 3, 2, and 1, where risk level 3 corresponds to high-risk. Simultaneously, the fusion feature vector of the first preset time window before the abnormal behavior occurs and the multi-source data temporal context are associated, and the three types of input data are standardized to form a model input dataset with a unified format. The fusion feature vector of the first preset time window before the abnormal behavior occurs includes weighted fusion features of vehicle, driver, and environment data, and the multi-source data temporal context consists of time-stamp-aligned sequence data such as vehicle speed changes, heart rate fluctuations, and environmental parameter evolution.
[0088] Secondly, the integrated risk level values, standardized fusion feature vectors, and temporal context data are jointly input into a pre-trained cause inference model. This model learns the correlation between risk level and cause, the key influence dimensions in the fusion features, and the causal logic of the temporal context. After analyzing the input data, it obtains the analysis results and outputs them as cause types. For example, high risk level + fusion feature: sudden decrease in obstacle distance + temporal context: slippery road in rainy weather → output cause type: sudden obstacle + slippery road surface.
[0089] Finally, the correlation between the cause type output by the cause inference model and the abnormal behavior type is verified. For example, the abnormal behavior of lane departure is matched with the cause type of line of sight departure. After confirming that there is no logical conflict between the cause type, the cause type is officially determined as the final cause of the current abnormal behavior and stored synchronously in the abnormal behavior analysis log. The corresponding risk level, fused feature vector and time series context data are associated to provide complete data support for subsequent intervention strategy matching and model optimization.
[0090] As can be seen from the above, this embodiment inputs the risk level of candidate triggers, fused feature vectors, and multi-source data temporal context into the trigger inference model. This incorporates the urgency weights corresponding to the risk level, and integrates comprehensive information from multi-dimensional features and causal correlation data at the temporal level. This allows the trigger inference model to reason based on more comprehensive and three-dimensional input information, avoiding misjudgments of triggers caused by a single data dimension. At the same time, directly using the trigger type output by the trigger inference model as the final trigger improves the efficiency of trigger identification. Furthermore, the richness and relevance of the input data enhance the accuracy of the inference results, ensuring that the triggers identified by the model not only fit the background of the abnormal behavior but also accurately point to the core cause. This provides reliable support for the formulation of subsequent targeted intervention strategies, further improving the scientific nature and effectiveness of driving safety management.
[0091] In one embodiment of this application, a corresponding prompt message is generated based on a prompting strategy, including: Determine the number of abnormal behaviors within the second preset time window; If the number of abnormalities is single, the level of the alert message is determined based on the risk level of the abnormal behavior; the content of the alert message is determined based on the cause of the abnormal behavior. Based on the temporary personalized driving profile, obtain the driver's preference history for prompting methods; Based on the level, content, and preference history of the prompts, personalized prompts that integrate visual and auditory modalities are generated. Personalized prompts are output through the vehicle's human-machine interface.
[0092] In this embodiment, the number of abnormal behaviors within a second preset time window is determined. For example, the second preset time window is set to 10 seconds. The second preset time window can be adjusted according to the driving scenario. For example, on highways, the second preset time window can be extended to 15 seconds. The number of abnormal behaviors identified within the second preset time window is statistically analyzed. Specifically, if only one independent abnormal behavior is detected, such as a single sudden braking or short-term lane departure, it is determined as a single abnormality, and this single abnormality directly enters the personalized prompt generation process. For example, if lane departure and sudden braking occur consecutively within the second preset time window of 10 seconds, a high-priority joint prompt mechanism is triggered, and the prompt information level is determined according to the risk level of the abnormal behavior. Specifically, a high-risk level corresponds to a strong warning prompt, a medium-risk level corresponds to a regular reminder prompt, and a low-risk level corresponds to a mild prompt. The corresponding prompt content is matched according to the identified cause. For example, when the cause is strong light causing visual deviation, the content focuses on visual correction and operation guidance; when the cause is fatigue driving, the content focuses on rest suggestions and safety tips, so that the prompt content specifically addresses the core problem.
[0093] For example, if a driver exhibits lane departure abnormality once within 15 seconds during a highway cruise, the risk level is determined to be medium risk. The cause is analyzed as "strong light causing visual deviation." Based on the risk level, the warning message level is determined to be a routine reminder. Based on the cause, the warning content focuses on visual correction and operation guidance, with the core statement being that strong light ahead is affecting visibility, and it is recommended to turn on the anti-glare mode and promptly straighten the steering wheel.
[0094] Personalized preference data is retrieved to adapt to driver habits: The driver's historical preference for prompting methods is extracted from temporary personalized driving profiles, including visual, auditory, and modal combination preferences. This is also correlated with past response data to different levels of prompts, such as the average response time to strong warning prompts, to form personalized adaptation parameters. For example, retrieving a driver's temporary personalized driving profile reveals a preference for Mandarin voice prompts combined with a static yellow icon on the dashboard. The resulting personalized prompts include a static yellow vision correction icon on the dashboard for visual prompts, and an auditory prompt indicating that strong glare ahead is affecting visibility and suggesting activating anti-glare mode and promptly straightening the steering wheel. These prompts are simultaneously output through the car audio system and dashboard, aligning with both the risk level and the trigger, and adapting to the driver's habits.
[0095] This embodiment constructs a multimodal prompting scheme based on the prompt level, core content, and preference parameters. Specifically, the strong warning level uses a red dynamic flashing icon + high-volume voice broadcast + eye-catching text on the central control screen; the regular reminder level uses a yellow static icon + medium-volume voice + concise text; and the mild prompt level only retains blue text + low-volume voice. For example, for drivers with high risk + strong light irradiation + preference for Mandarin, the generated prompt information is: Visual: The red vision correction icon on the dashboard is flashing; Auditory: Strong light ahead is affecting visibility, please turn on anti-glare mode and stay in your lane, ensuring efficient information delivery without distracting attention.
[0096] Secondly, the optimal interaction interface is selected based on the vehicle's hardware configuration. Visual prompts are prioritized through the instrument panel and head-up display, which are easily visible to the driver, while auditory prompts are delivered through the in-vehicle audio system. The use of the central control screen, which requires the driver to look down, as the primary prompt carrier is avoided. During the output, the current driving scenario is detected simultaneously, such as whether the driver is on a call. If the driver is on a call, the voice prompt volume is automatically increased to ensure that the driver can receive it clearly. The prompt output time and the driver's response behavior are recorded and fed back to the temporary personalized driving profile for subsequent preference optimization.
[0097] As can be seen from the above, this embodiment first filters individual abnormal behaviors through a second preset time window, avoiding chaotic prompts when multiple abnormalities overlap; when multiple abnormal behaviors occur simultaneously, the prompt level is matched based on the risk level, and the prompt intensity can be adjusted according to the degree of danger of the abnormal behavior, neither ignoring high-risk hazards nor causing excessive interference to driving due to low-risk prompts; the prompt content is customized according to the cause of the abnormality, so that the information directly addresses the core of the problem; at the same time, the prompt method preference is adapted according to the temporary personalized driving profile to meet the driver's personalized needs for visual and auditory modalities, and improve the driver's acceptance of abnormal behavior prompts; finally, through multimodal fusion and output through the vehicle human-machine interaction interface, the efficiency of information transmission is ensured, and it is in line with the attention distribution characteristics in the driving scenario, effectively balancing safety warnings and driving experience, and improving the response rate and intervention effect of prompts.
[0098] In one embodiment of this application, a method for analyzing abnormal driving behavior data of a driver further includes: If there are multiple abnormal behaviors, then a correlation analysis is performed on the multiple abnormal behaviors to detect whether there are any conflicts in their prompting strategies. When there is no conflict, the prompting strategy corresponding to the abnormal behavior is executed; When conflicts exist, the conflicting alert strategies are resolved based on preset priority rules. The priority rules are as follows: alert strategies triggered by high-risk abnormal behaviors take precedence over alert strategies triggered by low-risk abnormal behaviors. For abnormal behaviors with the same risk level, their priority is ranked according to the urgency of the trigger, with the following ranking rules: collision avoidance alerts are greater than normal driving alerts, which are greater than comfort alerts.
[0099] In this embodiment, correlation analysis and conflict detection are performed on multiple abnormal behaviors within the second preset time window. Specifically, multiple abnormal behaviors detected within the second preset time window are statistically analyzed. For example, if lane departure, close following, and excessive speed occur simultaneously, the causal or concurrent relationships between the abnormal behaviors are first analyzed using a behavior correlation algorithm. For example, does excessive speed lead to close following? Secondly, the prompting strategies corresponding to each abnormality are compared. Specifically, the prompting modality is checked, such as whether all require high-volume voice, the timing of output, and whether there is mutual exclusion or interference in the operation guidance, to determine whether the prompting strategies conflict. Among these, the timing of output includes whether pop-up windows are required simultaneously; whether there is mutual exclusion or interference in the operation guidance includes whether there are opposite instructions for acceleration and deceleration.
[0100] Secondly, the prompting strategy is determined based on the prompts in conflict-free scenarios. For example, if multiple detected abnormal behaviors do not interfere with each other in terms of modality, timing, and content (e.g., lane departure corresponds to a yellow icon and a gentle voice, and lights not being turned on corresponds to blue text and a low-volume voice), then all prompting strategies corresponding to the abnormalities are executed according to the principle of simultaneous concurrent execution without interference. Visual prompts are displayed in zones, for example, lane-related prompts are displayed on the left side of the instrument panel, and light-related prompts are displayed on the right side. Auditory prompts are broadcast sequentially according to time, with a broadcast interval of 0.5 seconds to avoid overlapping prompts, so that the driver can clearly receive multiple prompt messages.
[0101] In this embodiment, when conflicting prompting strategies exist—for example, if two high-risk anomalies both require high-volume voice broadcasts, or if operational instructions contradict each other—a priority determination mechanism is activated. Specifically, the prompting strategies for high-risk anomalies are prioritized over those for medium- and low-risk anomalies, and the prompts corresponding to high-risk anomalies are executed first. If multiple anomalies have the same risk level, such as both being medium-risk close following and lane departure, the prompting strategies are prioritized based on the urgency of the abnormal behavior triggers. Specifically, collision avoidance triggers have the highest priority, followed by proper driving triggers, and comfort reminder triggers have the lowest priority. Only the prompting strategy with the highest priority is executed, or conflict-free prompts are combined. For example, the same voice broadcast can integrate instructions to maintain a safe following distance and return to the lane to resolve the problem of multiple conflicting prompting strategies.
[0102] As can be seen from the above, this embodiment first detects conflicting prompt strategies through correlation analysis for multiple abnormal behavior scenarios, avoiding problems such as modal overlap and contradictory operation guidance caused by simultaneous output of multiple prompts, thus ensuring the orderly transmission of prompt information. When there is no conflict, each prompt strategy is executed synchronously, which can comprehensively cover multiple abnormal risk points and not miss any key safety prompts. When there is a conflict, the conflict is resolved based on the priority rules of risk level priority and cause urgency, so that prompts for high-risk abnormalities are delivered first. At the same time, abnormalities of the same risk level are sorted according to collision avoidance > normal driving > comfort reminders, focusing on the most urgent and core safety needs, avoiding non-urgent prompts from distracting the driver's attention, balancing the comprehensiveness and relevance of prompts, maximizing the effectiveness of prompt information in complex driving scenarios, providing scientific guidance for drivers to quickly respond to risks and avoid dangers, and further strengthening driving safety.
[0103] Corresponding to the driver's abnormal driving behavior data analysis method in the above embodiment, Figure 2 This is a structural block diagram of a driver's abnormal driving behavior data analysis system provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2The driver's abnormal driving behavior data analysis system 20 includes: a multi-source data acquisition module 21, a feature extraction and fusion module 22, an abnormal behavior identification module 23, a risk level assessment module 24, an abnormal cause analysis module 25, an intervention strategy matching module 26, and a prompt information generation module 27.
[0104] in, The multi-source data acquisition module 21 is used to acquire multi-source data during the driver's driving process, including vehicle data, driver data and environmental data. The feature extraction and fusion module 22 is used to extract and fuse features from multi-source data to obtain a fused feature vector. The abnormal behavior recognition module 23 is used to input the fused feature vector into the hybrid analysis model and output the abnormal behavior and probability of the driver. The risk level assessment module 24 is used to assess the risk of abnormal behavior based on the probability of abnormal behavior and environmental data, and obtain the risk level. The anomaly cause analysis module 25 is used to analyze abnormal behavior based on risk level, fused feature vector and time context of multi-source data to obtain the causes of abnormal behavior; The intervention strategy matching module 26 is used to match prompt strategies from a preset intervention strategy library based on risk level and trigger. The prompt message generation module 27 is used to generate corresponding prompt messages based on the prompt strategy.
[0105] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the multi-source data acquisition module 21, feature extraction and fusion module 22, abnormal behavior identification module 23, risk level assessment module 24, abnormal cause analysis module 25, intervention strategy matching module 26, and prompt information generation module 27 are shown.
[0106] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0107] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0108] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.
[0109] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in any embodiment of the driver's abnormal driving behavior data analysis method provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.
[0110] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0111] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0112] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0113] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0114] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0115] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0116] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0117] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for analyzing abnormal driving behavior data of drivers, characterized in that, include: Acquire multi-source data during the driver's driving process, wherein the multi-source data includes vehicle data, driver data, and environmental data; Feature extraction and fusion are performed on the multi-source data to obtain a fused feature vector; The fused feature vector is input into the hybrid analysis model, which outputs the driver's abnormal behavior and probability. Based on the probability of the abnormal behavior and the environmental data, a risk assessment is performed on the abnormal behavior to obtain a risk level; Based on the risk level, the fused feature vector, and the temporal context of the multi-source data, the abnormal behavior is analyzed to obtain the cause of the abnormal behavior; Based on the risk level and the trigger, a prompting strategy is matched from a preset intervention strategy library; Based on the aforementioned prompting strategy, corresponding prompting information is generated.
2. The method for analyzing abnormal driving behavior data of drivers according to claim 1, characterized in that, The step of extracting and fusing features from the multi-source data to obtain a fused feature vector includes: The time-domain and frequency-domain features of the multi-source data are extracted to obtain a multi-source feature vector; Based on a preset weight mapping table, weights are matched to the multi-source feature vectors to obtain the basic weights of the multi-source feature vectors. The basic weights are adjusted based on the driving scenario to obtain the adjusted weights; The multi-source feature vectors are weighted and fused based on the adjusted weights to obtain the fused feature vector.
3. The method for analyzing abnormal driving behavior data of drivers according to claim 2, characterized in that, The adjustment of the basic weights based on the driving scenario to obtain the adjusted weights includes: The environmental data and the vehicle data are input into a scene classification model to obtain the driving scene type; The driving scenario types include at least: highway cruising, urban congestion, following another vehicle on a ring road, passing through an intersection, and parking. Based on the driving scenario and the preset scenario-weight adjustment coefficient table, match the adjustment coefficients of each feature in the multi-source data. The basic weights are adjusted based on the adjustment coefficients to obtain the adjusted weights.
4. The method for analyzing abnormal driving behavior data of drivers according to claim 3, characterized in that, Also includes: The multi-source data is used to identify emergency events using an LSTM-based anomaly detection model. When an emergency is detected, a temporary adjustment instruction is generated based on the severity of the emergency; Based on the aforementioned temporary adjustment instructions, an attention mechanism is used to calculate the correlation between each data feature and the emergency event, and a correlation adjustment coefficient is generated. The correlation adjustment coefficient is fused with the adjustment coefficient to obtain the comprehensive adjustment coefficient; The basic weights are adjusted based on the comprehensive adjustment coefficient to obtain the final weights applicable to emergency events.
5. The method for analyzing abnormal driving behavior data of drivers according to claim 1, characterized in that, After inputting the fused feature vector into the hybrid analysis model and outputting the driver's abnormal behavior and probability, the method further includes: A temporary personalized driving profile is created for the driver, which is generated based on anonymized driving data; The inherent driving habits in the temporary personalized driving profile are assessed for normality, and normal driving habits are selected. The similarity between the abnormal behavior and the normal driving habits is calculated to obtain a behavior similarity score; Based on the behavioral similarity score, a personalized correction coefficient is generated using a logistic regression model; The probability output by the hybrid analysis model is corrected based on the personalized correction coefficient.
6. The method for analyzing abnormal driving behavior data of drivers according to claim 5, characterized in that, The process of generating personalized correction coefficients based on the behavioral similarity score using a logistic regression model includes: Extract the frequency, intensity, and context of inherent behaviors from the temporary personalized driving profile during historical driving to construct a historical behavior baseline; Calculate the multidimensional similarity between the current abnormal behavior and the historical behavior baseline, including temporal similarity, intensity similarity, and scene similarity; The behavioral consistency index is calculated based on the weighted sum of the multidimensional similarities. The behavioral consistency index is input into a pre-trained logistic regression model to generate the personalized correction coefficient.
7. The method for analyzing abnormal driving behavior data of drivers according to claim 6, characterized in that, Also includes: Calculate the completeness and statistical significance of the historical behavior baseline in the current triggering scenario; The completeness index and statistical significance index are compared with preset standard thresholds, which include: a completeness index threshold and a statistical significance index threshold. If both indicators are greater than or equal to the standard threshold, the confidence level of the personalized correction coefficient is determined to be sufficient, and the personalized correction coefficient is used to correct the probability output by the hybrid analysis model. If any indicator is less than the standard threshold, the confidence level of the personalized correction coefficient is deemed insufficient. The absolute differences between the completeness index and the statistical significance index and the standard threshold are then calculated, and the attenuation factor is determined. The personalized correction coefficient is adjusted based on the attenuation factor.
8. The method for analyzing abnormal driving behavior data of drivers according to claim 5, characterized in that, The process of creating a temporary personalized driving profile for the driver includes: The video data obtained from the driver's data is preprocessed. Driver behavior features are extracted from the video data based on a computer vision model; Establish a temporal correlation between the driver's behavioral characteristics and vehicle operation data and environmental data to form a driving habit characteristic sequence; The driving habit feature sequence is subjected to pattern recognition and classification to obtain classification results, which include inherent driving habits that belong to abnormal behavior and inherent driving habits that do not belong to abnormal behavior. The temporary personalized driving profile is generated based on the inherent driving habits that are not considered abnormal behavior.
9. The method for analyzing abnormal driving behavior data of drivers according to claim 1, characterized in that, The abnormal behavior is analyzed based on the risk level, the fused feature vector, and the temporal context of the multi-source data to obtain the causes of the abnormal behavior, including: Extract the temporal context data of multi-source data within a first preset time window before the occurrence of the abnormal behavior; The time-series context data is input into the causal discovery algorithm to obtain a multivariate causal graph; The multivariate causal graph connects different data feature nodes through directed edges, and assigns a quantified causal relationship strength value to each directed edge; Based on the aforementioned risk level, a first threshold is set; From the multivariate causal graph, the causal features corresponding to the directed edges that are connected to the abnormal behavior and whose causal relationship strength is greater than or equal to the first threshold are selected as candidate causal features. The candidate trigger features are matched with a preset trigger knowledge base; If a match is successful, the matched trigger type will be used as the trigger for the abnormal behavior. If a match fails, the corresponding candidate cause is input into a preset cause inference model for inference to obtain the cause type of the abnormal behavior.
10. A driver's abnormal driving behavior data analysis system, characterized in that, include: The multi-source data acquisition module is used to acquire multi-source data during the driver's driving process, wherein the multi-source data includes vehicle data, driver data and environmental data; The feature extraction and fusion module is used to extract and fuse features from the multi-source data to obtain a fused feature vector. An abnormal behavior identification module is used to input the fused feature vector into a hybrid analysis model and output the abnormal behavior and probability of the driver. The risk level assessment module is used to assess the risk of the abnormal behavior based on the probability of the abnormal behavior and the environmental data, and obtain the risk level. An anomaly cause analysis module is used to analyze the abnormal behavior based on the risk level, the fused feature vector, and the temporal context of the multi-source data to obtain the cause of the abnormal behavior. An intervention strategy matching module is used to match prompt strategies from a preset intervention strategy library based on the risk level and the trigger. The prompt message generation module is used to generate corresponding prompt messages based on the prompt strategy.