Driving risk dynamic prediction method based on multi-modal data, edge computing device and medium
Patent Information
- Application Number
- CN202511544717.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-10-27
AI Technical Summary
然而,这种方式主要聚焦于道路、其他车辆等外部因素带来的风险,忽略了驾驶员自身的情况对驾驶安全的影响,因此,在驾驶风险识别的准确性上存在不足
根据本公开的基于多模态数据的驾驶风险动态预测方法,通过获取驾驶员的多模态生理数据、车辆状态数据和行驶环境数据;对多模态生理数据、车辆状态数据和行驶环境数据进行特征提取并进行特征融合,得到融合特征向量;通过应激识别模型,基于融合特征向量,计算表征驾驶员应激状态的综合应激指数;通过车速预测模型,根据车辆的历史车速和融合特征向量,预测车辆在未来目标时间的车速;根据车辆在未来目标时间的车速与车辆的当前车速的差值、综合应激指数和行驶环境数据,确定驾驶风险。本公开实施例通过获取驾驶过程中多个不同维度的信息,并对这些信息进行特征提取和融合,使得多个不同维度的信息能够通过数据特征进行表达,应激识别模型基于融合特征向量识别出的综合应激指数,能够反映驾驶员在当前驾驶环境下的应激状态,车速预测模型可以根据历史车速和融合特征向量预测车辆在未来目标时间的车速,结合车速差值、综合应激指数和行驶环境数据来确定驾驶风险,既考虑了驾驶员自身状态,又兼顾了车辆与环境因素,使驾驶风险评估更加准确。
Smart Images

Figure CN121682065B_ABST
Abstract
Description
Technical Field
[0001] This disclosure pertains to the field of human factors intelligence, and particularly relates to a method for dynamic prediction of driving risks based on multimodal data, an edge computing device, and a medium. Background Technology
[0002] With the continuous increase in car ownership, road traffic safety has become an increasingly important concern. Various risks can arise at any time while a vehicle is in motion, threatening not only the lives of the driver and passengers but also causing harm to other pedestrians or vehicles on the road. In related technologies, driving risks are primarily assessed through the external environment. For example, cameras and radar installed on vehicles are used to capture road conditions in real time and monitor the operation of surrounding vehicles. This external information is then analyzed to determine the potential risks of collisions, rear-end collisions, etc. However, this approach mainly focuses on risks posed by external factors such as roads and other vehicles, neglecting the impact of the driver's own condition on driving safety. Therefore, it is insufficient in terms of the accuracy of driving risk identification. Summary of the Invention
[0003] One of the technical problems this disclosure aims to solve is that the accuracy of judging driving risks solely based on the external environment is insufficient.
[0004] To address the aforementioned technical problems, this disclosure provides a method for dynamic prediction of driving risks based on multimodal data, an edge computing device, and a medium to improve the accuracy of driving risk assessment.
[0005] Firstly, this disclosure provides a method for dynamic prediction of driving risk based on multimodal data, including: acquiring multimodal physiological data of the driver, vehicle state data, and driving environment data; extracting features from the multimodal physiological data, vehicle state data, and driving environment data and fusing them to obtain a fused feature vector; calculating a comprehensive stress index characterizing the driver's stress state based on the fused feature vector using a stress recognition model; predicting the vehicle's speed at a future target time using a vehicle speed prediction model based on the vehicle's historical speed and the fused feature vector; and determining the driving risk based on the difference between the vehicle's speed at the future target time and the vehicle's current speed, the comprehensive stress index, and the driving environment data.
[0006] In some embodiments, feature extraction and feature fusion are performed on multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector, including: extracting features from multimodal physiological data, vehicle state data, and driving environment data to obtain multimodal physiological feature vectors, vehicle state feature vectors, and driving environment feature vectors; concatenating the vehicle state feature vector and the multimodal physiological feature vector to obtain a concatenated feature vector; and using the driving environment feature vector as a query and the concatenated feature vector as a key-value pair for cross-attention calculation to obtain the fused feature vector.
[0007] According to some embodiments of this disclosure, the vehicle speed prediction model is trained as follows: A sample dataset is acquired; the sample dataset includes multimodal physiological data of the driver under different environments, vehicle state data, driving environment data, and historical vehicle speed; the sample dataset is labeled, and the labels are used to annotate the driver's comprehensive stress index and historical vehicle speed; features are extracted and fused from the multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector; the fused feature vector and historical vehicle speed are input into the vehicle speed prediction model, and the vehicle speed prediction model is trained using the following loss function: L_total = α×MSE(V_pred, V_true) + β×MSE(SSI_pred, SSI_true) Where L_total represents the loss function, MSE() represents the mean squared error, V_pred represents the predicted vehicle speed output by the vehicle speed prediction model, V_true represents the historical vehicle speed of the sample corresponding to the label of the sample dataset, α represents the first weight, SSI_pred represents the predicted comprehensive stress index output by the stress recognition model, SSI_true represents the true comprehensive stress index corresponding to the label of the sample dataset, and β represents the second weight.
[0008] According to some embodiments of this disclosure, the method further includes: extracting sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius from sample driving environment data; calculating an environmental risk indicator factor based on the sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius, wherein the environmental risk indicator factor characterizes the severity of environmental mutation; adjusting a second weight based on the environmental risk indicator factor; wherein the larger the environmental risk indicator factor, the larger the second weight.
[0009] According to some embodiments of this disclosure, driving risk is determined based on the difference between the vehicle's speed at a future target time and the vehicle's current speed, a comprehensive stress index, and driving environment data, including: Extract illumination abrupt change coefficient, blind spot risk value, and road curvature radius from driving environment data; based on the scene type corresponding to the driving environment data, weight and merge the extracted illumination abrupt change coefficient, blind spot risk value, and road curvature radius to obtain a scene risk coefficient; calculate the absolute value of the difference between the vehicle's speed at a future target time and the vehicle's current speed; calculate the driving risk value based on the absolute value of the difference, the comprehensive stress index, and the scene risk coefficient; determine the driving risk based on the driving risk value.
[0010] According to some embodiments of this disclosure, a driving risk value is calculated based on the absolute value of the difference, a comprehensive stress index, and a scenario risk coefficient, including: According to the formula: R = w1×ΔV_pred_norm + w2×SSI + w3×SRC ΔV_pred_norm = ΔV_pred / V_max Calculate driving risk value; Where R represents the driving risk value, ΔV_pred represents the absolute value of the difference, V_max represents the maximum permissible vehicle speed, SSI represents the comprehensive stress index, SRC represents the scenario risk coefficient, w1, w2 and w3 represent dynamic weights, and ΔV_pred_norm represents the deviation ratio.
[0011] According to some embodiments of this disclosure, the method further includes: matching the weight adjustment rules corresponding to the scene type; the weight adjustment rules include triggering conditions, and when the driving environment data meets the triggering conditions, adjusting and normalizing the basic weights corresponding to each dynamic weight by the weight adjustment amount corresponding to each dynamic weight to obtain the adjusted dynamic weights; and adjusting the dynamic weights according to the weight adjustment rules.
[0012] According to some embodiments of this disclosure, the method further includes: calculating the driver's reaction time based on a comprehensive stress index; calculating the critical safe speed in the current scenario based on the reaction time and driving environment data; correcting the critical safe speed based on the critical safe speed and the current risk level of the driving risk to obtain a corrected safe speed corresponding to the current risk level; and / or, when the current risk level is the target risk level, decelerating the vehicle to the corrected safe speed corresponding to the current risk level.
[0013] Secondly, this disclosure provides a dynamic prediction device for driving risk based on multimodal data, comprising: an acquisition module for acquiring multimodal physiological data of the driver, vehicle state data, and driving environment data; a fusion module for extracting features from the multimodal physiological data, vehicle state data, and driving environment data and fusing the features to obtain a fused feature vector; a calculation module for calculating a comprehensive stress index characterizing the driver's stress state based on the fused feature vector using a stress recognition model; a prediction module for predicting the vehicle's speed at a future target time using a vehicle speed prediction model based on the vehicle's historical speed and the fused feature vector; and a determination module for determining the driving risk based on the difference between the vehicle's speed at the future target time and the vehicle's current speed, the comprehensive stress index, and the driving environment data.
[0014] Thirdly, this disclosure provides an edge computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the driving risk dynamic prediction method based on multimodal data as described in the first aspect above.
[0015] Fourthly, this disclosure provides a computer-readable storage medium on which a program or instructions are stored, which, when executed by a processor, implements the driving risk dynamic prediction method based on multimodal data as described in the first aspect above.
[0016] Fifthly, this disclosure provides a chip including a processor and a communication interface, the communication interface and the processor being coupled together, the processor being used to run programs or instructions to implement the driving risk dynamic prediction method based on multimodal data as described in the first aspect above.
[0017] In a sixth aspect, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the driving risk dynamic prediction method based on multimodal data as described in the first aspect above.
[0018] The above-described one or more technical solutions in the embodiments of this disclosure have at least the following technical effects: According to the dynamic prediction method for driving risk based on multimodal data disclosed herein, the following steps are taken: First, multimodal physiological data of the driver, vehicle state data, and driving environment data are acquired. Then, features are extracted and fused from the multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector. Next, a stress recognition model is used to calculate a comprehensive stress index representing the driver's stress state based on the fused feature vector. Finally, a speed prediction model is used to predict the vehicle's speed at a future target time based on the vehicle's historical speed and the fused feature vector. Finally, driving risk is determined based on the difference between the vehicle's speed at the future target time and its current speed, the comprehensive stress index, and the driving environment data. This disclosure embodiment acquires information from multiple dimensions during the driving process, extracts and fuses features from this information, enabling the information from multiple dimensions to be expressed through data features. The stress recognition model, based on the comprehensive stress index identified by the fused feature vector, can reflect the driver's stress state in the current driving environment. The vehicle speed prediction model can predict the vehicle speed at a future target time based on historical vehicle speed and the fused feature vector. Combining the vehicle speed difference, the comprehensive stress index, and driving environment data, it determines driving risk, taking into account both the driver's own state and vehicle and environmental factors, making driving risk assessment more accurate.
[0019] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating the dynamic prediction method for driving risks based on multimodal data provided in this embodiment of the disclosure; Figure 2 This is a schematic diagram of the structure of the driving risk dynamic prediction device based on multimodal data provided in this embodiment of the disclosure; Figure 3 This is a schematic diagram of the structure of the edge computing device provided in the embodiments of this disclosure. Detailed Implementation
[0022] The technical solutions of the embodiments of this disclosure will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure are within the scope of protection of this disclosure.
[0023] The terms "first," "second," etc., used in this disclosure and in the claims are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this disclosure can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0024] It should be noted that, in the optional embodiments of this disclosure, the personnel information and other related data involved require the permission or consent of the personnel involved when the embodiments of this disclosure are applied to specific products or technologies. Furthermore, the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this disclosure involve personnel-related data, it must be obtained with the authorization and consent of the personnel, the authorization and consent of the relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments must also be implemented with the authorization and consent of the subject.
[0025] In related technologies, driving risks are primarily assessed through the external environment. For example, cameras and radar installed on vehicles are used to capture road conditions in real time and monitor the operation of surrounding vehicles. This external information is then analyzed to determine the potential risks of collisions, rear-end collisions, etc. However, this approach mainly focuses on risks posed by external factors such as roads and other vehicles, neglecting the driver's own subjective judgment of driving safety. Therefore, it is insufficient in the accuracy of driving risk identification, especially when external sensors (such as cameras and radar) malfunction.
[0026] To address at least one of the aforementioned technical problems, this disclosure provides a method for dynamic prediction of driving risks based on multimodal data, an edge computing device, and a medium. The following, in conjunction with the accompanying drawings, provides a detailed description of the method for dynamic prediction of driving risks based on multimodal data, the edge computing device, and the medium provided by this disclosure through specific embodiments and application scenarios.
[0027] The driving risk dynamic prediction method based on multimodal data can be applied to in-vehicle terminals, specifically executed by hardware or software within the in-vehicle terminal. This disclosure does not limit the implementation method of the terminal.
[0028] Specifically, during simulated driving training, the execution entity of the aforementioned dynamic prediction method for driving risks based on multimodal data can also be an edge computing device or a functional module or entity within an edge computing device capable of implementing the dynamic prediction method for driving risks based on multimodal data. (This disclosure embodiment) There are no restrictions on the implementation methods for edge devices. Certainly .
[0029] Taking an edge computing device as the execution subject as an example, the method for dynamic prediction of driving risks based on multimodal data provided in this disclosure embodiment will be described.
[0030] like Figure 1 As shown, the dynamic prediction method for driving risk based on multimodal data includes steps 110, 120, 130, 140 and 150.
[0031] Step 110: Obtain the driver's multimodal physiological data, vehicle status data, and driving environment data.
[0032] In this embodiment of the disclosure, multimodal physiological data refers to physiological data collected from different physiological dimensions. Physiological data includes various measurable data signals in the human body, including but not limited to electrocardiogram (ECG) signals, skin temperature (SKT) signals, photoplethysmogram (PPG) signals, electrodermal activity (EDA) signals, heart rate (HR) signals, electromyogram (EMG) signals, electroencephalogram (EEG) signals, peripheral capillary oxygen saturation (SPO2) signals, eye movement signals, etc.
[0033] Vehicle status data refers to various parameters used to characterize the vehicle's operating status during driving, including but not limited to the vehicle's current speed, acceleration, brake pedal travel, accelerator pedal opening, steering wheel angle, and lateral offset. Driving environment data refers to external environmental information related to vehicle operation. It clearly presents the external road scene and environmental conditions in which the vehicle is located. Driving environment data can include road type, road curvature radius, road speed limit information, obstacle distribution, weather conditions, and light intensity.
[0034] The aforementioned multimodal physiological data can be collected using wearable devices. For example, a smart bracelet with heart rate and blood oxygen detection functions can be worn on the driver's wrist to collect heart rate and blood oxygen saturation data; an EEG cap can be worn on the driver's head to obtain EEG signals; and a skin conductance sensor can be attached to the skin surface of the driver's hand to collect skin conductance response data. These wearable devices can have wireless transmission capabilities, allowing the collected physiological data to be transmitted to a data processing terminal in real time. Vehicle status data can be collected through the vehicle's own sensors and onboard systems. For example, the vehicle's speed sensor can collect the current vehicle speed; the acceleration sensor is used to obtain the vehicle's acceleration data, and so on. Driving environment data can be collected through onboard sensing devices. For example, vehicle-mounted cameras can capture images of road scenes, and image recognition technology can be used to extract data such as road type, road curvature radius, road speed limit signs, and pedestrian distribution; onboard radar can detect driving status data such as the position, speed, and distance of surrounding vehicles.
[0035] Step 120: Extract features from multimodal physiological data, vehicle status data, and driving environment data, and fuse the features to obtain a fused feature vector.
[0036] Specifically, before feature extraction, the raw data can be preprocessed to reduce noise interference and unify the data temporal sequence. For EEG and eye-tracking signals in multimodal physiological data, filtering and noise reduction can be performed. For example, wavelet denoising can be used for EEG signals to reduce power line interference and electromyographic noise while retaining effective EEG components; electrocardiogram, electrodermal activity, and other physiological signals can be filtered using low-pass filtering to eliminate high-frequency noise. For driving environment data, since it comes from different acquisition devices, spatiotemporal alignment can be performed to map data collected by different devices at the same time point to the same spatial coordinate system, improving the temporal and spatial consistency of the data.
[0037] It should be noted that different methods can be used for feature extraction of different types of data. For example, for EEG signals, sensitive frequency band features related to stress states and event-related potential components can be extracted. Sensitive frequency band features include alpha wave inhibition rate (the ratio of alpha wave power under stress to alpha wave power under rest) and the theta / beta wave power ratio. P300 amplitude features in event-related potentials can also be extracted; changes in P300 amplitude features deliberately characterize the driver's allocation of attentional resources.
[0038] In the process of extracting features from eye-tracking signals, the driver's pupil diameter can be recorded first, and the coefficient of variation of the pupil diameter, i.e., the ratio of the standard deviation of the pupil diameter to the mean, can be calculated. This coefficient of variation reflects the fluctuation of pupil size and is related to the driver's stress level. Alternatively, the fixation point dispersion, i.e., the standard deviation of the fixation point coordinates, can be calculated based on the fixation point coordinates obtained from eye tracking. This is used to measure the stability of the driver's gaze. Furthermore, the coefficient of variation of the pupil diameter and the fixation point dispersion can be fused to generate the Visual Load Index (VLI). The VLI can reflect the difficulty of visual information processing and the level of psychological stress of the driver in the current driving scenario.
[0039] For electrocardiogram (ECG) signals, the LF / HF (Low Frequency / High Frequency) ratio, which is the ratio of the power of the low-frequency component to the power of the high-frequency component, can be extracted from heart rate variability (HRV); for electrodermal signals, the rise slope characteristics can be analyzed.
[0040] Feature extraction from driving environment data primarily targets key environmental factors affecting driving safety. These can include the illumination abrupt change coefficient, which is the change in light intensity per unit time (in lux / s). An excessively large illumination abrupt change coefficient (such as when entering or exiting a tunnel) can affect the driver's visibility. The road curvature radius (in meters) relates to the difficulty and safety of vehicle turns. Blind spot risk values, calculated based on radar detection range, range from 0 to 1; higher values indicate a higher risk of undetected areas in the surrounding area, potentially concealing collision hazards.
[0041] Feature extraction of vehicle status data can include extracting current vehicle speed, acceleration, frequency of brake and accelerator pedal operation, steering angle, etc.
[0042] Because different types of features have significantly different numerical ranges, direct fusion may lead to an imbalance in feature weights. Therefore, a standardization method can be used to process the features, converting each feature value into a standard value with a mean of 0 and a standard deviation of 1, so that all types of features are on the same order of magnitude. Then, the standardized EEG features, eye-tracking features, ECG / skin conductance features, scene features, and vehicle state features can be horizontally concatenated in a preset order to form a fused feature vector, Context, containing multi-dimensional information. Alternatively, a weighted summation method can be used, assigning a weight to each feature based on the degree of influence of each data point on driving risk, and then performing a weighted summation to obtain the fused feature vector, Context.
[0043] Step 130 uses a stress identification model to calculate a comprehensive stress index that characterizes the driver's stress state based on a fused feature vector.
[0044] Specifically, the Comprehensive Stress Index (SSI) is a scalar quantity used to quantify a driver's stress state, with a value ranging from [0,1]. The closer the value is to 0, the more calm and relaxed the driver is, indicating a low level of stress; the closer the value is to 1, the more stressed the driver is, indicating a strong stress response.
[0045] It should be noted that the Comprehensive Stress Index (SSI) can be obtained through a stress recognition model. This model can pre-learn the mapping relationship between the fused feature vector (Context) and the driver's actual stress state, thereby enabling the prediction of stress levels.
[0046] The stress recognition model described above can be implemented using a fully connected neural network layer or a lightweight regressor. Taking a fully connected neural network as an example, the calculation process of the Comprehensive Stress Index (SSI) can be expressed as follows: SSI = Sigmoid (FC_SSI (Context)) Where “Context” represents the input fusion feature vector, “FC_SSI” represents the computation of the fully connected layer, and the Sigmoid function is used to compress the output result to the range [0, 1].
[0047] The stress recognition model described above can be trained using supervised learning, with the label (SSI_true) representing the overall stress level serving as the supervisory signal. The label can be obtained as follows: Under specific environmental abrupt changes (such as sudden obstacles or abrupt changes in road conditions), multimodal physiological signals of the driver are collected, including pupillary change rate, specific EEG power change rate, skin conductance slope, and heart rate variability indicators. Then, during the collection of the driver's multimodal physiological signals, expert assessments of the driver's stress state, the driver's self-reported stress level, and behavioral responses observed by experts (such as sudden braking and steering wheel operation amplitude) are combined to comprehensively analyze these multimodal physiological signal features. Through standardization and experimental calibration, the analysis results are converted into normalized values within the range [0, 1], thus obtaining the label SSI_true.
[0048] During the training phase of the stress recognition model, the mean squared error (MSE) can be used as the loss function. This involves calculating the mean squared difference between the predicted comprehensive stress index (SSI_pred) and the true label (SSI_true), and continuously adjusting the model parameters through backpropagation until the loss function converges. In practical applications, the trained stress recognition model can receive the fused feature vector Context as input and output the comprehensive stress index SSI.
[0049] Step 140: Using the vehicle speed prediction model, predict the vehicle speed at a future target time based on the vehicle's historical speed and fused feature vector.
[0050] The vehicle speed prediction model can combine the vehicle's historical speed, such as the historical speed from the current moment to the past 5 seconds or the past 10 seconds, and fuse the feature vector Context to predict the vehicle's speed at a future target time, such as predicting the speed in the next 2 seconds or the next 3 seconds.
[0051] The aforementioned vehicle speed prediction model can be based on a neural network architecture. The input layer of the vehicle speed prediction model includes the vehicle's historical speed sequence, such as the speed data within 10 seconds prior to the current time point t; it can also include a fused feature vector Context.
[0052] In the hidden layers of the vehicle speed prediction model, the main task is to extract and process features from the historical vehicle speed sequence and the fused feature vector Context input from the input layer, and output the vehicle speed at the future target time in the output layer. For example, the fused feature vector Context can be input into a refinement network, such as a small fully connected layer FC_GateControl. After processing by this network, a gate control vector GateControlVector is output. The dimension of the gate control vector GateControlVector can be 16 or 32. The gate control vector GateControlVector encodes the core physiological load type and intensity information caused by the current environmental changes to the driver. For example, it can reflect key information such as "stress intensity caused by strong light in the tunnel" and "intensity of spatial pressure caused by sharp bends".
[0053] Next, feature fusion is performed, concatenating the extracted gate control vector (GateControlVector) with the vehicle speed data from each time step in the historical vehicle speed sequence. For example, for the vehicle speed data v_t at time step t, the concatenated extended input vector is [v_t, GateControlVector]; for the vehicle speed data v_t-1 at time step t-1, the concatenated vector is [v_t-1, GateControlVector]. In this way, the vehicle speed data at each time step incorporates key information about the current scene and driver state, enabling the vehicle speed prediction model to combine real-time dynamic conditions when analyzing historical vehicle speeds.
[0054] The core of the hidden layer is a bidirectional LSTM (Long Short-Term Memory) layer, which processes the expanded input sequence after feature fusion to learn the temporal patterns of vehicle speed changes. The bidirectional LSTM layer, through its internal forget gate, input gate, and output gate mechanisms, can focus on both recent details of vehicle speed changes and capture long-term speed trends. Furthermore, since the expanded input sequence contains a gate control vector, it can serve as contextual information to guide the LSTM's learning process: when the gate control vector indicates a high level of environmental risk (e.g., high driver stress), the LSTM reduces its reliance on historical vehicle speed patterns unrelated to the current risk through the forget gate, focusing more on features reflecting real-time stress states. This makes the vehicle speed prediction model more realistic in complex scenarios.
[0055] After dynamic processing by the bidirectional LSTM layer, the output layer of the vehicle speed prediction model generates the vehicle speed prediction value at the future target time.
[0056] To improve the accuracy of vehicle speed prediction models, a large amount of historical driving data can be used for training during the model's training phase. Training data can include historical vehicle speed sequences under different scenarios, corresponding fused feature vectors, and actual future vehicle speeds as labels. By continuously adjusting the parameters of the vehicle speed prediction model, the error (e.g., mean squared error) between the predicted and actual vehicle speeds is minimized.
[0057] Step 150: Determine driving risk based on the difference between the vehicle's speed at the future target time and the vehicle's current speed, the Comprehensive Stress Index (SSI), and driving environment data.
[0058] Specifically, the difference between the predicted vehicle speed at a future target time and the current vehicle speed reflects the short-term trend of vehicle speed changes. For example, if the predicted speed will decrease or increase significantly in the next 5 seconds, it indicates that the driver is about to brake suddenly or there is a risk of speeding. The comprehensive stress index reflects the driver's current stress level. Driving environment data can include information such as road curvature, light abrupt change coefficient, and blind spot risk value, reflecting the safety level of the external environment in which the vehicle is located.
[0059] The driving risk value R can be calculated using a weighted summation method. Weights can be assigned to the three types of input items mentioned above. Different input items have different degrees of influence on driving risk; for example, corresponding weights can be set according to the characteristics of actual driving scenarios. For instance, in scenarios with significant road curvature, the weight of driving environment data is relatively high; while when the driver's stress index is high, the weight of the comprehensive stress index will increase. The weight settings can be pre-calibrated using experimental data or dynamically adjusted according to different driving scenarios; this disclosure does not limit this approach.
[0060] Different risk levels, such as low risk, medium risk, and high risk, can be set based on the magnitude of the driving risk value R. For example, when the risk value is in the range of [0, 0.3], it can be judged as low risk; when the risk value is in the range of [0.3, 0.7], it can be judged as medium risk; and when the risk value exceeds 0.7, it can be judged as high risk.
[0061] According to the dynamic prediction method for driving risk based on multimodal data disclosed herein, multimodal physiological data of the driver, vehicle state data, and driving environment data are acquired; features are extracted and fused from the multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector; a stress recognition model is used to calculate a comprehensive stress index representing the driver's stress state based on the fused feature vector; a vehicle speed prediction model is used to predict the vehicle's speed at a future target time based on the vehicle's historical speed and the fused feature vector; and driving risk is determined based on the difference between the vehicle's speed at the future target time and the vehicle's current speed, the comprehensive stress index, and the driving environment data. This disclosure embodiment acquires information from multiple dimensions during the driving process, extracts and fuses features from this information, enabling the information from multiple dimensions to be expressed through data features. The stress recognition model, based on the comprehensive stress index identified by the fused feature vector, can reflect the driver's stress state in the current driving environment. The vehicle speed prediction model can predict the vehicle speed at a future target time based on historical vehicle speed and the fused feature vector. Combining the vehicle speed difference, the comprehensive stress index, and driving environment data, it determines driving risk, taking into account both the driver's own state and vehicle and environmental factors, making driving risk assessment more accurate.
[0062] In some embodiments of this disclosure, traditional feature fusion methods typically simply concatenate these data, ignoring the interrelationships and importance between different data types. For example, a driver's physiological state may be closely related to the current driving environment, while the vehicle's driving state is influenced by both driver operation and environmental conditions. Therefore, feature extraction and feature fusion are performed on multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector Context, including: Feature extraction was performed on multimodal physiological data, vehicle state data, and driving environment data to obtain multimodal physiological feature vector PhysioEncoded, vehicle state feature vector VehGroup, and driving environment feature vector EnvGroup. The vehicle state feature vector VehGroup and the multimodal physiological feature vector PhysioEncoded are concatenated to obtain the concatenated feature vector PhysioVehConcat; The driving environment feature vector EnvGroup is used as the query, and the concatenated feature vector PhysioVehConcat is used as the key-value pair for cross-attention calculation to obtain the fused feature vector Context.
[0063] In this embodiment, different types of features can be encoded and mapped to a higher-dimensional, more expressive feature space to obtain feature vectors. For example, the extracted features can be divided into three groups: Physiological features (PhysioGroup): including physiological indicators reflecting driver cognitive load, visual stress, and stress arousal, such as visual load index, θ / β wave power ratio, P300 amplitude change, GSR rise slope, and LF / HF ratio; Vehicle features (VehGroup): including vehicle dynamics parameters such as real-time vehicle speed, acceleration, steering wheel angle, and lateral offset; and Environmental features (EnvGroup): including environmental parameters such as illumination abrupt change coefficient, reciprocal of road curvature radius (1 / R), blind spot risk value, and current scene type encoding (e.g., tunnel exit = 1, underground parking garage exit = 2, sharp bend = 3).
[0064] Each set of features can be input into an independent feature encoding sub-network, such as a small multilayer perceptron (MLP) or a one-dimensional convolutional neural network (1D-CNN). These sub-networks are responsible for mapping the original features to a higher-dimensional, more expressive feature space. For example: multimodal physiological feature vector PhysioEncoded = MLP_Physio(PhysioGroup); vehicle state feature vector VehEncoded = MLP_Veh(VehGroup); driving environment feature vector EnvEncoded = MLP_Env(EnvGroup).
[0065] In this embodiment, the encoded vehicle state feature vector VehEncoded and the multimodal physiological feature vector PhysioEncoded can be concatenated to form a concatenated feature vector PhysioVehConcat.
[0066] The feature fusion process can utilize a cross-attention mechanism. The driving environment feature vector `EnvEncoded` is used as the query, and the concatenated feature vector `PhysioVehConcat` is used as the key-value pair, forming a common source for both the key and value vectors. Cross-attention calculation is then performed to obtain the fused feature vector `Context`. Specifically, the query vector `Query = EnvEncoded` can be used to query the key vector `Key = PhysioVehConcat`. The similarity between each element in `Query` and `Key` is calculated, and the attention weights are obtained by normalizing using the Softmax function. This weight distribution represents which physiological and vehicle state features (represented by `Key`) are most critical for risk assessment under the current specific environmental risk conditions (defined by `Query`). The formula is as follows: Attention Weights = Softmax( (Query×Key^T) / sqrt(d_k) ) Where d_k is the dimension of the Key vector, used for scaling, and sqrt() represents the square root.
[0067] The calculated attention weights are used to sum the value vector Value = PhysioVehConcat to generate a dynamically weighted fused feature vector Context, i.e., fused feature vector Context = Attention Weights × Value.
[0068] The following example illustrates the technical effectiveness of generating the fused feature vector Context using the above method. Taking a tunnel exit scenario as an example, the illumination abrupt change coefficient is very high in the driving environment features. Special attention needs to be paid to the following in the spliced feature vector: Eye movement signals: the recovery rate after a sharp contraction of the pupil diameter, and whether the fixation point dispersion suddenly increases; Electroencephalogram (EEG) signals: whether the alpha wave inhibition deepens, and whether the P300 amplitude is abnormal due to unexpected strong light stimulation; Electrodermal activity (EGA) signals: whether the GSR rise slope is steep; Vehicle status: whether the current vehicle speed is too high, and whether there is a tendency for sudden deceleration.
[0069] The attention mechanism plays a crucial role: Under the strong query signal of EnvEncoded (high light intensity abrupt change), the cross-attention mechanism assigns higher weights to the physiological characteristics closely related to "visual glare stress" (pupil changes, alpha inhibition, P300 abnormality, GSR spike) and the corresponding vehicle responses (deceleration behavior). The output dynamic feature vector strongly represents the driver's visual stress and fright state due to the glare abrupt change (reflected through high-weighted physiological signals) and the vehicle's initial response (reflected through high-weighted vehicle signals), and is deeply coupled with the driving environment characteristics of the high light intensity abrupt change. This allows the fused feature vector to more accurately predict the risky behaviors that the driver may exhibit under glare, such as control errors due to temporary visual blindness.
[0070] In this embodiment, feature vectors are obtained by extracting features from multimodal physiological data, vehicle state data, and driving environment data respectively. By concatenating the vehicle state feature vector and the driving environment feature vector, the vehicle state and the driver's physiological state can be initially integrated. This allows for the full extraction of the interrelationships and influences between the driving environment, vehicle state, and driver's physiological state by using the driving environment feature vector as a query and the concatenated feature vector as a key-value pair for cross-attention calculation. This strengthens the expression of the interrelationships between the data of each dimension corresponding to the fused feature vector, and further improves the accuracy of subsequent driving risk assessment.
[0071] In some embodiments of this disclosure, if the vehicle speed prediction model is trained solely for the vehicle speed prediction task during training, the correlation between driver stress and vehicle speed changes may be overlooked, leading to insufficient accuracy in vehicle speed prediction under complex driving scenarios. Considering that changes in driver stress directly affect the driver's control of vehicle speed, and that drastic changes in vehicle speed may conversely exacerbate the driver's stress response, the vehicle speed prediction model is trained in the following manner: Obtain the sample dataset; the sample dataset includes multimodal physiological data of drivers under different environments, vehicle status data, driving environment data, and historical vehicle speed; the sample dataset has corresponding labels, which are used to label the driver's comprehensive stress index and historical vehicle speed. Feature extraction and feature fusion are performed on multimodal physiological data, vehicle status data, and driving environment data of the samples to obtain a sample fusion feature vector; The sample fusion feature vector and the historical vehicle speed of the samples are input into the vehicle speed prediction model, and the following loss function is used to train the vehicle speed prediction model: L_total = α×MSE(V_pred, V_true) + β×MSE(SSI_pred, SSI_true) Where L_total represents the loss function, MSE() represents the mean squared error, V_pred represents the vehicle speed prediction value output by the vehicle speed prediction model, V_true represents the historical vehicle speed of the sample corresponding to the label of the sample dataset, α represents the first weight, SSI_pred represents the comprehensive stress index prediction value output by the stress recognition model, SSI_true represents the true value of the comprehensive stress index corresponding to the label of the sample dataset, and β represents the second weight. The vehicle speed prediction model includes a feature extraction layer, a feature fusion layer, and a bidirectional LSTM layer. The feature extraction layer processes the sample fusion feature vector through the extraction network and outputs a gate control vector (GateControlVector). The feature fusion layer concatenates the gate control vector (GateControlVector) with the historical vehicle speed of the sample to form an expanded input vector. The bidirectional LSTM layer outputs the predicted vehicle speed value based on the expanded input vector.
[0072] In this embodiment, the sample dataset may include various types of data from drivers in different environments to improve the comprehensiveness and generalization ability of the vehicle speed prediction model training. The sample dataset includes sample multimodal physiological data (such as EEG signals, eye movement features, ECG and skin conductance responses in different scenarios), sample vehicle state data (such as vehicle speed, acceleration, steering wheel angle in different road conditions), sample driving environment data (such as scenario information such as different lighting conditions, road curvature, blind spot risk, etc.), and sample historical vehicle speeds (i.e., the speed sequence of the vehicle over a period of time).
[0073] The sample dataset can also be labeled, with the labels used to annotate the driver's true comprehensive stress index (SSI_true) and the vehicle's true speed at a future target time, i.e., the sample's historical vehicle speed (V_true) at a certain time in the sample dataset. The true comprehensive stress index can be determined by combining expert evaluation and standardization of the sample's multimodal physiological data with driver self-reports and behavioral observations.
[0074] In this embodiment, the stress recognition model and the vehicle speed prediction model can be combined to achieve collaborative training through joint optimization.
[0075] The first term α×MSE(V_pred, V_true) in the formula is used to measure the error of vehicle speed prediction, and the second term β×MSE(SSI_pred, SSI_true) is used to measure the error of comprehensive stress index prediction.
[0076] During training, features are extracted and fused from the sample dataset to obtain a sample fused feature vector. This vector is then used as input for training to obtain the outputs of the vehicle speed prediction model and the stress recognition model. The total loss value is then calculated using the aforementioned loss function. Finally, based on the total loss value, the parameters of the stress recognition model and the vehicle speed prediction model are adjusted using a backpropagation algorithm to gradually reduce the total loss value.
[0077] The values of α and β can be set according to actual training needs. For example, in scenarios where the accuracy of vehicle speed prediction is of greater concern, the value of α can be appropriately increased; in scenarios where the impact of driver stress on vehicle speed needs to be considered, the value of β can be increased.
[0078] In this embodiment, the sample dataset includes various sample data of drivers in different environments and their corresponding labels, which enables the model to better learn the relationship between various factors and vehicle speed and comprehensive stress index. The loss function used for training considers the mean square error between the predicted vehicle speed value and the historical vehicle speed corresponding to the label, as well as the mean square error between the predicted comprehensive stress index value and the actual comprehensive stress index value corresponding to the label. The influence of the two is balanced by the first weight and the second weight, so that the vehicle speed prediction model can not only focus on the prediction accuracy of the vehicle speed itself during the learning process, but also consider that the driver's stress state will affect the driving operation and thus be related to the vehicle speed, thereby improving the prediction accuracy of the vehicle speed prediction model.
[0079] Specifically, the vehicle speed prediction model can be based on a neural network architecture. The input layer of the vehicle speed prediction model includes the vehicle's historical speed sequence, such as the speed data within 10 seconds prior to the current time point t; it can also include sample fusion feature vectors.
[0080] The hidden layers of the vehicle speed prediction model include a feature extraction layer, a feature fusion layer, and a bidirectional LSTM layer. The feature extraction layer is used to extract and process features from the fused feature vector Context input from the input layer. For example, the feature extraction layer may include a small fully connected layer FC_GateControl, which outputs a gate control vector GateControlVector. The dimension of GateControlVector can be 16 or 32. GateControlVector encodes the type and intensity of the core physiological load caused to the driver by the current environmental changes. For example, it can reflect key information such as "stress intensity caused by strong tunnel light" and "intensity of spatial pressure caused by sharp curves."
[0081] Next, feature fusion is performed through a feature fusion layer, concatenating the extracted gate control vector (GateControlVector) with the vehicle speed data from each time step in the historical vehicle speed sequence. For example, for the vehicle speed data v_t at time step t, the concatenated extended input vector is [v_t, GateControlVector]; for the vehicle speed data v_t-1 at time step t-1, the concatenated vector is [v_t-1, GateControlVector]. In this way, the vehicle speed data at each time step incorporates key information about the current scene and driver state, enabling the vehicle speed prediction model to combine real-time dynamic conditions when analyzing historical vehicle speeds.
[0082] The core of the hidden layer is a bidirectional LSTM (Long Short-Term Memory) layer, which processes the expanded input sequence after feature fusion to learn the temporal patterns of vehicle speed changes. The bidirectional LSTM layer, through its internal forget gate, input gate, and output gate mechanisms, can focus on both recent details of vehicle speed changes and capture long-term speed trends. Furthermore, since the expanded input sequence contains a gate control vector, it can serve as contextual information to guide the LSTM's learning process: when the gate control vector indicates a high level of environmental risk (e.g., high driver stress), the LSTM reduces its reliance on historical vehicle speed patterns unrelated to the current risk through the forget gate, focusing more on features reflecting real-time stress states. This makes the vehicle speed prediction model more realistic in complex scenarios.
[0083] After dynamic processing by the bidirectional LSTM layer, the output layer of the vehicle speed prediction model generates the vehicle speed prediction value at the future target time.
[0084] In some embodiments of this disclosure, the importance of driver stress state to vehicle speed prediction varies under different driving environments. The more drastic the environmental changes, the more significant the impact of driver stress response on vehicle speed control. If a fixed value is used for β, the accuracy of the vehicle speed prediction model may be insufficient in some scenarios. Therefore, the method further includes: Extract the sample illumination mutation coefficient, sample blind spot risk value, and sample road curvature radius from the sample driving environment data; Based on the sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius, the environmental risk indicator factor EnvRiskIndicator is calculated, where the environmental risk indicator factor EnvRiskIndicator characterizes the severity of environmental mutation. The second weight is adjusted based on the environmental risk indicator factor EnvRiskIndicator; the larger the environmental risk indicator factor EnvRiskIndicator, the larger the second weight.
[0085] In this embodiment, the severity of environmental changes can be determined by a combination of factors such as changes in lighting, road conditions, obstacles, and blind spots. The more severe the environmental changes, the more significant the impact of driver stress on vehicle speed control may be. Therefore, a higher second weight can be used to make the vehicle speed prediction model pay more attention to the relationship between the driver's stress state and the environment during the training process.
[0086] In this embodiment, considering that the more drastic the environmental changes, the more likely the driver's stress state will fluctuate significantly, which will also affect the driver's control of vehicle speed, by increasing the second weight when the environmental changes are more drastic, the vehicle speed prediction model can focus more on learning the matching degree between the comprehensive stress index and the actual environment during training, thereby improving the training quality and prediction accuracy of the vehicle speed prediction model.
[0087] To quantify the severity of environmental abrupt changes, sample illumination abrupt change coefficient, sample blind spot risk value, and sample road curvature radius can be extracted from sample driving environment data. The illumination abrupt change coefficient reflects the magnitude of change in light intensity per unit time, such as sudden changes in strong light at tunnel entrances and exits, or changes in headlight illumination when meeting oncoming traffic at night. Changes in illumination can affect the driver's vision and easily trigger a stress response. The blind spot risk value can be calculated based on the radar detection range, with values between [0,1]. A larger value indicates a larger undetected area around the vehicle, posing a potential collision risk. The road curvature radius reflects the degree of road curvature; a smaller radius indicates a sharper curve, requiring the driver to frequently adjust steering, increasing the difficulty of handling and the level of stress.
[0088] The environmental risk indicator (EnvRiskIndicator) can be calculated based on the extracted sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius. After normalization, the environmental risk indicator takes a value between [0,1], with a larger value indicating a higher degree of environmental mutation. The calculation process involves first standardizing the sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius, and then obtaining the environmental risk indicator (EnvRiskIndicator) through a fusion method.
[0089] Specifically, the sample illumination change coefficient can be normalized by combining it with a preset maximum possible value, MaxLightChange, such as 10000 lux / s. The formula is: Illumination indicator factor = Sample illumination change coefficient / MaxLightChange. For example, when the sample illumination change coefficient is 2500 lux / s, the illumination indicator factor is 2500 / 10000 = 0.25, which characterizes the relative intensity of the illumination change.
[0090] For the radius of curvature of the sample road, its reciprocal can be calculated, i.e., 1 / radius of curvature of the sample road, and then combined with the preset maximum curvature reciprocal value MaxCurvature, such as 1 / 20 m. -1 After normalization, the curvature indicator factor is calculated as (1 / road curvature radius) / MaxCurvature. Taking a sample road with a curvature radius of 40m as an example, the reciprocal of the curvature is 0.025m. -1 If MaxCurvature is 0.05 m -1 The curvature indicator factor is 0.025 / 0.05=0.5, which characterizes the relative steepness of road curves.
[0091] The sample blind zone risk value is already in the [0,1] range, so no additional normalization is needed and it can be directly used as the corresponding indicator factor.
[0092] After the parameters are standardized, they can be merged by taking the maximum value or by weighted averaging to obtain the environmental risk indicator factor EnvRiskIndicator.
[0093] In some embodiments of this disclosure, the formula can be used: β = β_base × (1 + γ ×EnvRiskIndicator) Adjust the second weight; Wherein, β represents the second weight, β_base represents the base value of the second weight, EnvRiskIndicator represents the environmental risk indicator factor, and γ represents the amplification factor. For example, the value of γ can be 1.0 or 1.5, which is used to control the sensitivity of β to environmental risk.
[0094] The following examples illustrate the effect of the second weight adjustment.
[0095] Example 1: Tunnel exit, dominated by sudden environmental changes: Environmental characteristics: The sample light intensity abrupt change coefficient = 2500 lux / s. Assuming MaxLightChange = 10000 lux / s, the light intensity indicator factor = 2500 / 10000 = 0.25. Assuming other environmental risks are low, EnvRiskIndicator ≈ 0.25.
[0096] Set β_base = 0.3 and γ = 1.0.
[0097] β = 0.3 × (1 + 1.0 × 0.25) = 0.3 × 1.25 = 0.375 Results: β increased from the baseline value of 0.3 to 0.375. The accuracy requirements for the predicted Comprehensive Stress Index (SSI_pred) in the loss function were significantly increased. This forced the vehicle speed prediction model to more accurately capture the physiological stress responses (such as pupil constriction and alpha wave inhibition) caused by strong light stimulation in scenarios of sudden changes in lighting, and to reflect these responses in vehicle speed prediction.
[0098] Example 2: Sharp bends in mountainous areas (operationally limited scenario): Environmental characteristics: The radius of curvature of the sample road is R = 40m, then the reciprocal of the curvature = 1 / 40 = 0.025 m -1 Assume MaxCurvature = 1 / 20 = 0.05 m -1 Therefore, the curvature indicator factor = 0.025 / 0.05 = 0.5. Assume the blind zone risk value = 0.2.
[0099] Set EnvRiskIndicator = max(0.5, 0.2) = 0.5 (or a weighted average). Set β_base = 0.3, γ = 1.0.
[0100] β = 0.3× (1 + 1.0 ×0.5) = 0.3 × 1.5 = 0.45 Effect: β increased to 0.45. In sharp curve scenarios, the vehicle speed prediction model must pay close attention to the driver's stress state caused by increased spatial pressure and control tension (which may manifest as increased heart rate variability LF / HF, increased steering wheel grip force related electromyographic signals not listed here but the principle is the same), to ensure that the predicted vehicle speed can reflect the driver's actual control ability and potential risks in the curve.
[0101] Example 3: Ordinary urban roads (low environmental risk): Environmental characteristics: stable lighting, straight roads, and small blind spots. EnvRiskIndicator ≈ 0.
[0102] β = 0.3 × (1 + 1.0 × 0) = 0.3 Results: β remains at its baseline value of 0.3. The loss function primarily focuses on the accuracy of vehicle speed prediction, with relatively lower requirements for the accuracy of physiological stress prediction (though these still exist), consistent with normal driving conditions.
[0103] By dynamically amplifying β through EnvRiskIndicator, the loss function significantly strengthens the requirement for the accuracy of the vehicle speed prediction model in predicting the driver's physiological stress state when environmental risks increase. This forces the vehicle speed prediction model to closely integrate and accurately reflect the driver's real-time physiological response caused by sudden environmental changes when predicting future vehicle speeds, thereby improving the reliability and safety of predictions in critical scenarios.
[0104] In this embodiment, the sample illumination mutation coefficient, sample blind spot risk value, and sample road curvature radius are extracted from the sample driving environment data. The illumination mutation coefficient reflects the degree of light change during driving, which affects the driver's vision; the blind spot risk value reflects the collision risk during driving; and the road curvature radius reflects the degree of road shape change, which affects the vehicle's speed and stability. The environmental risk indicator factor calculated by combining these three parameters can accurately quantify the degree of environmental mutation.
[0105] In some embodiments of this disclosure, when determining driving risk, the risk weights of environmental factors such as lighting, blind spots, and road conditions differ under different driving scenarios. If these factors are combined with fixed weights, the risk assessment may not match the actual scenario. Therefore, driving risk is determined based on the difference between the vehicle's speed at a future target time and the vehicle's current speed, the Comprehensive Stress Index (SSI), and driving environment data, including: Extract illumination abrupt change coefficient, blind spot risk value, and road curvature radius from driving environment data; Based on the scene type corresponding to the driving environment data, the extracted illumination abruptness coefficient, blind spot risk value and road curvature radius are weighted and merged to obtain the scene risk coefficient SRC. Calculate the absolute value ΔV_pred of the difference between the vehicle's speed at a future target time and the vehicle's current speed; The driving risk value R is calculated based on the absolute value of the difference ΔV_pred, the comprehensive stress index, and the scenario risk coefficient. The driving risk is determined based on the driving risk value R.
[0106] In this embodiment, the risk impact of illumination abrupt change coefficient, blind spot risk value, and road curvature radius varies under different scenarios. By matching weights to scenario types, risk quantification can be made more closely aligned with actual scenarios. Scenario types can be automatically identified based on driving environment data. For example, a sudden increase in illumination abrupt change coefficient can be identified as a "tunnel entrance / exit scenario"; a very small road curvature radius can be identified as a "sharp bend scenario"; and a persistently high blind spot risk value located on an urban road can be identified as a "complex intersection scenario," etc.
[0107] Different weighting rules can be set for different scenario types. For example, in the "tunnel entrance / exit scenario", the impact of sudden changes in illumination on driving safety is the most prominent, so the weight of the sudden change in illumination coefficient can be set to the highest, such as 0.6, the weight of blind spot risk value is the second highest, such as 0.3, and the weight of road curvature radius is the lowest, such as 0.1. In the "sharp bend scenario", the weight of road curvature radius can be increased, such as 0.7, and the weights of sudden change in illumination coefficient and blind spot risk value can be set to 0.1 and 0.2 respectively. In the "complex intersection scenario", the blind spot risk value has the greatest impact, so its weight can be set to 0.6, and the weights of sudden change in illumination coefficient and road curvature radius can be set to 0.2.
[0108] Before weighted merging, the illumination mutation coefficient and road curvature radius can be normalized, and then the weighted sum can be calculated according to the weight corresponding to the scene type to obtain the scene risk coefficient SRC. The larger the value, the higher the environmental risk of the current scene.
[0109] After calculating the scenario risk coefficient, the driving risk value R can be calculated based on the absolute value ΔV_pred of the difference between the vehicle's speed at the future target time and the vehicle's current speed, the comprehensive stress index SSI, and the scenario risk coefficient SRC.
[0110] In some embodiments, the driving risk value R can be calculated by averaging, taking the maximum value, weighted summation, etc.
[0111] Taking weighted summation as an example, we can use the formula: R = w1×ΔV_pred_norm + w2×SSI + w3×SRC ΔV_pred_norm = ΔV_pred / V_max Calculate the driving risk value R; Where R represents the driving risk value, ΔV_pred represents the absolute value of the difference, V_max represents the maximum permissible vehicle speed, SSI represents the comprehensive stress index, SRC represents the scenario risk coefficient, and w1, w2 and w3 represent dynamic weights.
[0112] In this embodiment, the driving risk value is obtained by weighting and summing the absolute value of the speed difference, the comprehensive stress index, and the scenario risk coefficient according to dynamic weights. This fully considers the contribution of different factors to driving risk and can adjust the importance of each factor according to different situations. This allows the assessment of driving risk to be adapted to the needs of different driving scenarios, and the resulting driving risk value can be more in line with the actual safety situation.
[0113] In some embodiments, risk level classification thresholds can be preset. For example, when the driving risk value is in the range of [0, 0.4], it is determined to be "low risk", indicating that the current driving state is stable; when it is in the range of [0.4, 0.7], it is determined to be "medium risk", prompting the driver to be more vigilant; when it exceeds 0.7, it is determined to be "high risk", indicating that there is an immediate risk of collision or operational error.
[0114] In this embodiment, the illumination mutation coefficient, blind spot risk value and road curvature radius are extracted from driving environment data. Then, the scene risk coefficient is obtained by weighted merging based on the scene type corresponding to the driving environment data. The differences in the importance of each risk factor under different scenarios are taken into account, so that the scene risk coefficient can more accurately represent the overall risk level of the current driving scenario and further improve the accuracy of driving risk assessment.
[0115] In some embodiments of this disclosure, the sources of risk may differ in different specific scenarios. If a fixed weight is used to calculate the risk value, the risk assessment will not match the actual scenario. For example, for driver visual stress at tunnel exits and environmental hazards in underground parking garages, the weights of different calculation indicators need to be adjusted accordingly. Therefore, the method further includes: Match the weight adjustment rules corresponding to the scenario type; the weight adjustment rules include the trigger conditions. When the driving environment data meets the trigger conditions, the basic weights corresponding to each dynamic weight are adjusted and normalized by the weight adjustment amount corresponding to each dynamic weight to obtain the adjusted dynamic weights. The dynamic weights are adjusted according to the weight adjustment rules.
[0116] In this embodiment, multiple sets of weight adjustment rules can be pre-set, each set corresponding to a type of scenario, and these rules are stored together. After collecting driving environment data, the weight adjustment rules corresponding to the scenario type can be matched.
[0117] Taking the "environmental change-dominated scenario" as an example, the trigger condition is that the light change coefficient is greater than the preset threshold C1, such as 1000 lux / s. If the light change coefficient in the current driving environment reaches 2500 lux / s, the trigger condition is met, and weight adjustment rule 1 is matched. For the "space-constrained scenario", the trigger condition is that the blind spot risk value is greater than the threshold C2, such as 0.7, and the effective road width is less than the threshold C3, such as 3.5m. If the blind spot risk value is 0.8 and the effective road width is 3m in the underground parking garage scenario, weight adjustment rule 2 is matched. The trigger condition for the "operation-constrained scenario" is that the road curvature radius is less than the threshold C4, such as 50m, and the curvature radius at a sharp bend in a mountainous area is 40m, and weight adjustment rule 3 is matched.
[0118] The weight adjustment rules can be roughly divided into two steps: dynamically calculating the intermediate weight values w1_interim, w2_interim, and w3_interim, and then obtaining the final weights w1, w2, and w3 applied to the risk value calculation through normalization, so that w1 + w2 + w3 = 1.
[0119] In one example, the weight adjustment rules are as follows: Rule 1: Triggering condition: When the light intensity abrupt change coefficient is greater than the preset threshold C1, for example, 1000 lux / s.
[0120] Intermediate weight calculation: w2_interim = w2_base + Δw2_rule1 (e.g., base weight w2_base = 0.4, adjustment Δw2_rule1 = 0.2, then w2_interim = 0.6); w3_interim = w3_base - Δw3_rule1 (e.g., base weight w3_base = 0.3, adjustment Δw3_rule1 = 0.05, then w3_interim = 0.25); w1_interim = w1_base (e.g., base weight w1_base = 0.3, remains unchanged).
[0121] Normalization calculation of final weights: Calculate the intermediate weights and sum S = w1_interim + w2_interim + w3_interim (e.g., 0.3 + 0.6 + 0.25 = 1.15); w1 = w1_interim / S (e.g., 0.3 / 1.15 ≈ 0.261); w2 = w2_interim / S (e.g., 0.6 / 1.15 ≈ 0.522); w3 = w3_interim / S (e.g., 0.25 / 1.15 ≈ 0.217).
[0122] Effect: Significantly increases the weight w2 of the driver's comprehensive stress index and slightly reduces the weight w3 of the scenario risk coefficient SRC, because the driver's visual stress contributes dramatically to the risk at this time.
[0123] Rule 2: Triggering conditions: When the blind spot risk value is greater than the preset threshold C2, for example, 0.7, and the effective road width is less than the preset threshold C3, for example, 3.5m.
[0124] Intermediate weight calculation: w3_interim = w3_base + Δw3_rule2 (e.g., w3_base = 0.3, Δw3_rule2 = 0.2, then w3_interim = 0.5); w1_interim = w1_base (e.g., 0.3); w2_interim = w2_base (e.g., 0.4).
[0125] The final weights are calculated using normalization: S = w1_interim + w2_interim + w3_interim (e.g., 0.3 + 0.4 + 0.5 = 1.2); w1 = w1_interim / S (0.3 / 1.2 = 0.25); w2 = w2_interim / S (0.4 / 1.2 ≈ 0.333); w3 = w3_interim / S (0.5 / 1.2 ≈ 0.417).
[0126] Effect: Significantly increases the weight w3 of the scenario risk coefficient SRC, because the danger of the environment itself is the main source of risk.
[0127] Rule 3: Triggering condition: When the radius of curvature of the road is less than the preset threshold C4, for example, 50m.
[0128] Intermediate weight calculation: w1_interim = w1_base + Δw1_rule3 (e.g., if w1_base = 0.3, Δw1_rule3 = 0.1, then w1_interim = 0.4); w3_interim = w3_base + Δw3_rule3 (e.g., if w3_base = 0.3, Δw3_rule3 = 0.1, then w3_interim = 0.4); w2_interim = w2_base (e.g., 0.4).
[0129] The final weights are calculated using normalization: S = w1_interim + w2_interim + w3_interim (e.g., 0.4 + 0.4 + 0.4 = 1.2); w1 = w1_interim / S (0.4 / 1.2 ≈ 0.333); w2 = w2_interim / S (0.4 / 1.2 ≈ 0.333); w3 = w3_interim / S (0.4 / 1.2 ≈ 0.333).
[0130] Effect: It simultaneously increases the weights of vehicle speed prediction deviation ΔV_pred and scenario risk coefficient SRC, because vehicle control stability and road geometry together constitute risk.
[0131] Where w1_base, w2_base, and w3_base are the basic weights determined through experiments or simulations, satisfying w1_base + w2_base + w3_base = 1.
[0132] Δw1_rule1, Δw2_rule1, Δw3_rule1, Δw3_rule2, Δw1_rule3, and Δw3_rule3 represent the preset weight adjustment amounts corresponding to each rule. Positive values indicate an increase in weight, while negative values indicate a decrease in weight. These can be determined through experiments or simulations. Thresholds C1, C2, C3, and C4 can also be determined through experiments or simulations.
[0133] In this embodiment, by matching the weight adjustment rules corresponding to the scenario type, it is possible to automatically identify and apply appropriate weight configurations based on the characteristics of different driving scenarios, so that the calculation of driving risk value can more accurately reflect the actual risk in the current scenario.
[0134] In some embodiments of this disclosure, providing warnings solely based on driving risk levels may not offer drivers specific safe operating guidelines. Therefore, the method further includes: The driver's reaction time is calculated based on the Comprehensive Stress Index (SSI). Calculate the critical safe speed Vsafe in the current scenario based on reaction time and driving environment data; Driving intervention is performed based on the critical safe speed Vsafe and the current risk level of driving risk.
[0135] In this embodiment, the driver's reaction time can be calculated based on the Comprehensive Stress Index (SSI). Specifically, considering the correlation between the driver's stress state and reaction time, driver fatigue, lack of concentration, excessive tension, and panic can all lead to prolonged reaction time.
[0136] In this embodiment, a model relating the Comprehensive Stress Index (SSI) to reaction time is pre-established through experiments and analysis. For example, when the SSI is in the moderate stress range of 0.2-0.4, the driver reacts most quickly, with a reaction time of 0.8-1.0 seconds. When the SSI is below 0.2, it indicates that the driver may be fatigued, and the reaction time will be extended to more than 1.5 seconds. When the SSI is above 0.6, the driver may experience operational delays due to excessive tension, and the reaction time will also increase to 1.3-1.6 seconds.
[0137] In some embodiments, the formula can be used: t = t_min + (t_max - t_min)×Sigmoid(k × (SSI - SSI_mid)) Calculate the driver's reaction time; Where t represents reaction time, t_min represents minimum reaction time, t_max represents maximum reaction time, SSI represents the Comprehensive Stress Index, SSI_mid represents the median value of the Comprehensive Stress Index, the reaction time corresponding to SSI_mid is approximately (t_min + t_max) / 2, and k represents the parameter controlling the steepness of the curve. k is a positive number, and the larger the value, the steeper the curve, and the faster the reaction time changes around SSI_mid.
[0138] In this embodiment, the Sigmoid function can simulate the nonlinear relationship between driver reaction time and the comprehensive stress index. By setting the minimum and maximum reaction times, as well as the parameters for controlling the steepness of the curve, it is possible to flexibly adapt to the reaction characteristics of different drivers and adjust the calculation accuracy of reaction time according to actual needs, thereby enabling a more accurate assessment of driving risks.
[0139] The critical safe speed, Vsafe, refers to the maximum safe speed at which a driver can avoid risks in a timely manner by operating the vehicle with a calculated reaction time in the current environment. It can be calculated by combining reaction time and driving environment data.
[0140] In some embodiments, the formula can be used: Vsafe = min( sqrt(2×a_max×S_available) ) S_available = k_scene×W×(1 - L_factor) Calculate the critical safe speed Vsafe in the current scenario; Where Vsafe represents the critical safe speed, sqrt() represents the square root, a_max represents the maximum safe deceleration, S_available represents the scene-dependent available safe distance, and S_available is greater than the reaction distance S_reaction = v_current × t and the braking distance S_braking = v_current. 2 The sum of (2×a_max) / (2×a_max), where v_current represents the current vehicle speed, t represents the reaction time, W represents the effective width of the road, k_scene represents the scene-based risk factor, and different scene types correspond to different k_scenes. For example, k_scene is 0.8 for tunnel exits, 0.6 for underground parking garages, and 0.7 for mountain curves. L_factor represents the scene coefficient.
[0141] In this embodiment, the available safe distance is calculated using the effective road width, scenario-based risk factors, and scenario coefficients. The effective road width directly reflects the driving space conditions, while the scenario risk factors and scenario coefficients adapt to the different safety requirements of different scenarios, making the calculation of the available safe distance more in line with the characteristics of the current environment, thereby improving the accuracy of the calculation of the critical safe speed in the current scenario.
[0142] After calculating the critical safe speed, driving intervention can be implemented based on the critical safe speed and the current risk level of driving risk. The driving intervention adopts a layered and progressive intervention strategy, which can reduce the risk caused by insufficient intervention and prevent excessive intervention from affecting the driving experience.
[0143] For example, when the current risk level is "low risk" and the actual vehicle speed is below the critical safe speed, no active intervention is required. The in-vehicle display simply shows the current risk level, reaction time, and critical safe speed in real time. If the actual speed is slightly above the critical safe speed, a mild intervention can be triggered, with a voice prompt stating, "Current speed is slightly high; it is recommended to reduce to XX km / h." When the current risk level is "medium risk," if the actual vehicle speed is above the critical safe speed, in addition to the voice prompt, a moderate intervention is initiated. This includes a prominent deceleration reminder displayed on the in-vehicle screen and slight vibration of the seat to reinforce the warning. If the driver does not decelerate in time, the drive system's power output can be briefly reduced to assist in slowing down. When the current risk level is "high risk," if the actual vehicle speed is above the critical safe speed, an emergency voice alarm can be triggered immediately, and the brake assist system can be automatically activated to quickly reduce the vehicle speed below the critical safe speed.
[0144] In this embodiment, the driver's reaction time is calculated based on the comprehensive stress index, which fully considers the relationship between the driver's physiological state and actual operating ability. The critical safe speed is calculated by combining reaction time and driving environment data, taking into account the driver's operating limits and the current road, lighting and other environmental characteristics, so that the calculated critical speed is more in line with actual safety requirements. Corresponding intervention measures are taken according to the current risk level and the critical safe speed, which reduces the probability of danger and further improves the safety of the driving process.
[0145] In some embodiments of this disclosure, after determining the critical safe speed and the current risk level, directly using the critical safe speed as the intervention standard may not adequately meet the safety requirements under different risk levels. For example, in high-risk scenarios, maintaining the critical safe speed may still pose hidden dangers, while in low-risk scenarios, overly stringent standards may affect driving efficiency. Therefore, driving intervention based on the critical safe speed Vsafe and the current risk level of driving risk includes: The critical safe speed Vsafe is adjusted based on the current risk level of driving risk to obtain the adjusted safe speed corresponding to the current risk level. If the current risk level is the target risk level, reduce the vehicle to the corrected safe speed corresponding to the current risk level.
[0146] In this embodiment, the correspondence between risk levels and correction coefficients can be preset. The higher the risk level, the smaller the correction coefficient, and the lower the corrected safe speed, thus reserving sufficient safety redundancy. For example, driving risks can be divided into three levels: low risk, medium risk, and high risk, with corresponding correction coefficients set to 1.1, 1, and 0.8, respectively. It should be noted that the correction coefficients can be calibrated through extensive experiments and simulations to adapt to the safety requirements of different risk scenarios.
[0147] The corrected safe speed is calculated using the formula: "Corrected safe speed = Critical safe speed × Correction coefficient corresponding to risk level". For example, if the critical safe speed for the current scenario is 10 km / h, then if the risk level is low, the corrected safe speed = 10 × 1.1 = 11 km / h; if it's medium risk, the corrected safe speed = 10 × 1 = 10 km / h; and if it's high risk, the corrected safe speed = 10 × 0.8 = 8 km / h.
[0148] If the risk level is the target risk level, reduce the vehicle to the corrected safe speed corresponding to the current risk level. The target risk level can be high risk.
[0149] In this embodiment, by correcting the critical safe speed according to the current risk level of driving risk, the safe speed can be dynamically adjusted according to the level of risk. This reduces the problem of insufficient protection or excessive intervention when using a fixed critical safe speed in different risk scenarios. When the current risk level reaches the target risk level, the vehicle is decelerated to the corresponding corrected safe speed to reduce the speed in time and avoid danger.
[0150] The driving risk dynamic prediction method based on multimodal data provided in this disclosure can be executed by a driving risk dynamic prediction device based on multimodal data. This disclosure uses the driving risk dynamic prediction device based on multimodal data executing the driving risk dynamic prediction method based on multimodal data as an example to illustrate the driving risk dynamic prediction device based on multimodal data provided in this disclosure.
[0151] This disclosure also provides a driving risk dynamic prediction device based on multimodal data.
[0152] like Figure 2 As shown, the driving risk dynamic prediction device based on multimodal data includes: The acquisition module 210 is used to acquire the driver's multimodal physiological data, vehicle status data, and driving environment data; The fusion module 220 is used to extract features from multimodal physiological data, vehicle state data, and driving environment data and fuse them to obtain a fused feature vector. The calculation module 230 is used to calculate a comprehensive stress index that characterizes the driver's stress state based on the fused feature vector through a stress recognition model; The prediction module 240 is used to predict the vehicle speed at a future target time by using a vehicle speed prediction model based on the vehicle's historical speed and fused feature vector. The determination module 250 is used to determine driving risks based on the difference between the vehicle's speed at a future target time and the vehicle's current speed, the comprehensive stress index, and driving environment data.
[0153] According to the dynamic prediction device for driving risk based on multimodal data disclosed herein, the device acquires multimodal physiological data of the driver, vehicle state data, and driving environment data; extracts and fuses features from the multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector; calculates a comprehensive stress index characterizing the driver's stress state based on the fused feature vector using a stress recognition model; predicts the vehicle's speed at a future target time using a vehicle speed prediction model based on the vehicle's historical speed and the fused feature vector; and determines the driving risk based on the difference between the vehicle's speed at the future target time and the vehicle's current speed, the comprehensive stress index, and the driving environment data. This disclosure embodiment acquires information from multiple dimensions during the driving process, extracts and fuses features from this information, enabling the information from multiple dimensions to be expressed through data features. The stress recognition model, based on the comprehensive stress index identified by the fused feature vector, can reflect the driver's stress state in the current driving environment. The vehicle speed prediction model can predict the vehicle speed at a future target time based on historical vehicle speed and the fused feature vector. Combining the vehicle speed difference, the comprehensive stress index, and driving environment data, it determines driving risk, taking into account both the driver's own state and vehicle and environmental factors, making driving risk assessment more accurate.
[0154] In some embodiments, the fusion module 220 can also be used for: Feature extraction is performed on multimodal physiological data, vehicle state data, and driving environment data to obtain a multimodal physiological feature vector PhysioEncoded, a vehicle state feature vector VehGroup, and a driving environment feature vector EnvGroup. The vehicle state feature vector VehGroup and the multimodal physiological feature vector PhysioEncoded are concatenated to obtain a concatenated feature vector PhysioVehConcat. The driving environment feature vector EnvGroup is used as the query, and the concatenated feature vector PhysioVehConcat is used as the key-value pair for cross-attention calculation to obtain the fused feature vector Context.
[0155] In some embodiments, the driving risk dynamic prediction device based on multimodal data may further include: The training module is used to acquire the sample dataset. The sample dataset includes multimodal physiological data of drivers under different environments, vehicle state data, driving environment data, and historical vehicle speeds. Each sample dataset has corresponding labels used to annotate the driver's comprehensive stress index and historical vehicle speed. Feature extraction and feature fusion are performed on the multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector. The fused feature vector and historical vehicle speeds are then input into the vehicle speed prediction model, and the model is trained using the following loss function: L_total = α×MSE(V_pred, V_true) + β×MSE(SSI_pred, SSI_true) Where L_total represents the loss function, MSE() represents the mean squared error, V_pred represents the predicted vehicle speed output by the vehicle speed prediction model, V_true represents the historical vehicle speed corresponding to the label of the sample dataset, α represents the first weight, SSI_pred represents the predicted comprehensive stress index output by the stress recognition model, SSI_true represents the true comprehensive stress index corresponding to the label of the sample dataset, and β represents the second weight; the vehicle speed prediction model includes a feature extraction layer, a feature fusion layer, and a bidirectional LSTM layer; the feature extraction layer is used to process the sample fused feature vector through the extraction network and output a gate control vector GateControlVector; the feature fusion layer is used to concatenate the gate control vector GateControlVector with the historical vehicle speed of the sample to form an expanded input vector; the bidirectional LSTM layer is used to output the predicted vehicle speed based on the expanded input vector.
[0156] In some embodiments, the training module can also be used for: Extract the sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius from the sample driving environment data; calculate the environmental risk indicator EnvRiskIndicator based on the sample illumination mutation coefficient, sample blind zone risk value, and sample road curvature radius; the environmental risk indicator EnvRiskIndicator characterizes the severity of environmental mutations; The second weight is adjusted based on the environmental risk indicator factor EnvRiskIndicator; the larger the environmental risk indicator factor EnvRiskIndicator, the larger the second weight.
[0157] In some embodiments, the determining module 250 may also be used for: The illumination abrupt change coefficient, blind spot risk value, and road curvature radius are extracted from the driving environment data. Based on the scene type corresponding to the driving environment data, the extracted illumination abrupt change coefficient, blind spot risk value, and road curvature radius are weighted and merged to obtain the scene risk coefficient SRC. The absolute value ΔV_pred of the difference between the vehicle's speed at a future target time and the vehicle's current speed is calculated. The driving risk value R is calculated based on the absolute value ΔV_pred, the comprehensive stress index, and the scene risk coefficient. The driving risk is determined based on the driving risk value R.
[0158] In some embodiments, the determining module 250 may also be used for: According to the formula: R = w1×ΔV_pred_norm + w2×SSI + w3×SRC ΔV_pred_norm = ΔV_pred / V_max Calculate the driving risk value; where R represents the driving risk value, ΔV_pred represents the absolute value of the difference, V_max represents the maximum permissible vehicle speed, SSI represents the comprehensive stress index, SRC represents the scenario risk coefficient, w1, w2 and w3 represent dynamic weights, and ΔV_pred_norm represents the deviation ratio.
[0159] In some embodiments, the determining module 250 may also be used for: Match the weight adjustment rules corresponding to the scenario type; the weight adjustment rules include trigger conditions. When the driving environment data meets the trigger conditions, the basic weights corresponding to each dynamic weight are adjusted and normalized by the weight adjustment amount corresponding to each dynamic weight to obtain the adjusted dynamic weights; adjust the dynamic weights according to the weight adjustment rules.
[0160] In some embodiments, the driving risk dynamic prediction device based on multimodal data may further include: The intervention module is used to calculate the driver's reaction time based on the Comprehensive Stress Index (SSI); calculate the critical safe speed (Vsafe) for the current scenario based on the reaction time and driving environment data; correct the critical safe speed (Vsafe) based on the current risk level of the driving risk to obtain the corrected safe speed corresponding to the current risk level; and / or correct the critical safe speed (Vsafe) based on the current risk level of the driving risk to obtain the corrected safe speed corresponding to the current risk level; and decelerate the vehicle to the corrected safe speed corresponding to the current risk level when the risk level is the target risk level.
[0161] The driving risk dynamic prediction device based on multimodal data in this disclosure can be an electronic device. The electronic device can be an edge computing device, a non-edge computing device, or a component within either an edge computing device or a non-edge computing device, such as an integrated circuit or a chip. The electronic device can be a terminal or any other device besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, eye tracker, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This disclosure does not impose specific limitations.
[0162] The driving risk dynamic prediction device based on multimodal data in this embodiment can be a device with an operating system. This operating system can be Microsoft Windows, Android, iOS, or other possible operating systems; this embodiment does not specifically limit the specific operating system.
[0163] In some embodiments, such as Figure 3 As shown, this disclosure also provides an edge computing device 300, including a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the program is executed by the processor 301, it implements the various processes of the above-described embodiments of the dynamic prediction method for driving risks based on multimodal data and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0164] It should be noted that the electronic devices in this embodiment include the mobile edge computing devices and non-mobile edge computing devices described above.
[0165] In optional examples, edge deployment schemes that deploy edge computing devices or components on the edge can migrate at least some computing tasks to edge devices, leveraging the advantages of edge computing to improve data processing efficiency, reduce latency and network resource consumption and dependence, and also benefit data security and privacy protection.
[0166] This disclosure also provides a computer-readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the various processes of the above-described embodiments of the dynamic prediction method for driving risks based on multimodal data and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0167] The processor is the processor in the edge computing device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0168] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described dynamic prediction method for driving risks based on multimodal data.
[0169] The processor is the processor in the edge computing device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0170] This disclosure also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described embodiments of the dynamic prediction method for driving risks based on multimodal data, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0171] It should be understood that the chip mentioned in the embodiments of this disclosure may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0172] It should be noted that the scope of the methods and apparatus in this disclosure is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.
[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0174] The embodiments of this disclosure have been described above with reference to the accompanying drawings. However, this disclosure is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this disclosure without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this disclosure.
[0175] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least some embodiments or examples of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0176] Although embodiments of this disclosure have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. A method for dynamic prediction of driving risk based on multimodal data, characterized in that, include: Acquire multimodal physiological data of the driver, vehicle status data, and driving environment data; The process involves extracting and fusing features from the multimodal physiological data, vehicle state data, and driving environment data to obtain a fused feature vector. This includes: extracting features from the multimodal physiological data, vehicle state data, and driving environment data to obtain a multimodal physiological feature vector, a vehicle state feature vector, and a driving environment feature vector; concatenating the vehicle state feature vector and the multimodal physiological feature vector to obtain a concatenated feature vector; and using the driving environment feature vector as a query and the concatenated feature vector as a key-value pair to perform cross-attention calculation to obtain the fused feature vector. Using a stress recognition model, a comprehensive stress index characterizing the driver's stress state is calculated based on the fused feature vector. The vehicle speed prediction model predicts the vehicle speed at a future target time based on the vehicle's historical speed and the fused feature vector. Determining driving risk based on the difference between the vehicle's speed at a future target time and its current speed, the comprehensive stress index, and the driving environment data includes: extracting a sudden change in illumination, a blind spot risk value, and a road curvature radius from the driving environment data; weighting and combining the extracted sudden change in illumination, blind spot risk value, and road curvature radius based on the scene type corresponding to the driving environment data to obtain a scene risk coefficient; calculating the absolute value of the difference between the vehicle's speed at a future target time and its current speed; calculating a driving risk value based on the absolute value of the difference, the comprehensive stress index, and the scene risk coefficient; and determining the driving risk based on the driving risk value.
2. The method according to claim 1, characterized in that, The vehicle speed prediction model was trained in the following manner: Obtain a sample dataset; the sample dataset includes multimodal physiological data of drivers under different environments, vehicle status data, driving environment data, and historical vehicle speed; the sample dataset is labeled, and the labels are used to annotate the driver's comprehensive stress index and historical vehicle speed. Feature extraction and feature fusion are performed on the multimodal physiological data, vehicle state data, and driving environment data of the samples to obtain a sample fusion feature vector; The sample fusion feature vector and the sample historical vehicle speed are input into the vehicle speed prediction model, and the vehicle speed prediction model is trained using the following loss function: L_total = α×MSE(V_pred, V_true) + β×MSE(SSI_pred, SSI_true) Where L_total represents the loss function, MSE() represents the mean squared error, V_pred represents the predicted vehicle speed output by the vehicle speed prediction model, V_true represents the historical vehicle speed of the sample corresponding to the label of the sample dataset, α represents the first weight, SSI_pred represents the predicted comprehensive stress index output by the stress recognition model, SSI_true represents the true comprehensive stress index corresponding to the label of the sample dataset, and β represents the second weight.
3. The method according to claim 2, characterized in that, The method further includes: Extract the sample illumination mutation coefficient, sample blind spot risk value, and sample road curvature radius from the sample driving environment data; Based on the sample illumination mutation coefficient, the sample blind zone risk value, and the sample road curvature radius, an environmental risk indicator factor is calculated, wherein the environmental risk indicator factor characterizes the severity of environmental mutation. The second weight is adjusted according to the environmental risk indicator factor; wherein, the larger the environmental risk indicator factor, the larger the second weight.
4. The method according to claim 1, characterized in that, The calculation of the driving risk value based on the absolute value of the difference, the comprehensive stress index, and the scenario risk coefficient includes: According to the formula: R = w1×ΔV_pred_norm + w2×SSI + w3×SRC ΔV_pred_norm = ΔV_pred / V_max Calculate driving risk value; Where R represents the driving risk value, ΔV_pred represents the absolute value of the difference, V_max represents the maximum permissible vehicle speed, SSI represents the comprehensive stress index, SRC represents the scenario risk coefficient, w1, w2 and w3 represent dynamic weights, and ΔV_pred_norm represents the deviation ratio.
5. The method according to claim 4, characterized in that, The method further includes: Match the weight adjustment rule corresponding to the scenario type; the weight adjustment rule includes a trigger condition. When the driving environment data meets the trigger condition, the basic weights corresponding to each dynamic weight are adjusted and normalized by the weight adjustment amount corresponding to each dynamic weight to obtain the adjusted dynamic weights. The dynamic weights are adjusted according to the weight adjustment rules.
6. The method according to claim 1, characterized in that, The method further includes: The driver's reaction time is calculated based on the comprehensive stress index. Calculate the critical safe speed in the current scenario based on the reaction time and the driving environment data; Based on the critical safe speed and the current risk level of the driving risk, the critical safe speed is corrected to obtain the corrected safe speed corresponding to the current risk level; and / or, if the current risk level is the target risk level, the vehicle is decelerated to the corrected safe speed corresponding to the current risk level.
7. An edge computing device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Vehicle risk assessment method and device, vehicle and storage medium
CN120525327A
Driving route prediction apparatus for a vehicle, a driving route prediction method thereof, and a system including the same
US20240426615A1