Automobile driving safety early warning method and system based on driver emotion

By integrating attention mechanisms based on vehicle driving, facial, and voice features to identify driver emotions and generate reasonable early warning strategies, this technology addresses the problem of insufficient driver emotion monitoring in existing technologies, thereby improving the accuracy and safety of early warnings.

CN121849174AInactive Publication Date: 2026-04-14BEIJING JIUZHOU ANHUA INFORMATION SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing automotive safety warning technologies mostly focus on monitoring vehicle driving status or road environment, paying insufficient attention to the driver's emotional state. This results in low accuracy and poor robustness in emotion recognition, making it impossible to effectively warn of dangerous driving behaviors caused by negative emotions in drivers.

Method used

By extracting features from multi-source data, including vehicle driving features, facial features, and voice features, and fusing them using an attention mechanism, combined with an emotion recognition model, early warning information and risk values ​​are generated. This distinguishes between immediate dangerous emotions and persistent negative emotions, and matches appropriate early warning strategies accordingly.

Benefits of technology

It enables accurate identification of the type and intensity of drivers' emotions, avoids excessive or delayed warnings, improves the rationality and safety of warnings, and reduces the probability of traffic accidents caused by negative emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121849174A_ABST
    Figure CN121849174A_ABST
Patent Text Reader

Abstract

The invention provides an automobile driving safety early warning method and system based on driver emotion, and relates to the technical field of automobile safety driving, and the method comprises the steps: carrying out the feature extraction of multi-source data, and obtaining multi-modal feature data; fusing the multi-modal feature data based on an attention mechanism to obtain a target fusion feature vector; the target fusion feature vector is input into an emotion recognition model, an emotion recognition result is obtained, and the emotion recognition result comprises an emotion type and emotion intensity; in response to matching of the emotion recognition result and a first preset emotion state, generating early warning prompt information; when the emotion recognition result of the driver in the first time is matched with a second preset emotion state, calculating a driving safety risk value according to the emotion intensity and the confidence coefficient of the multi-modal feature data; and determining a corresponding early warning strategy based on the driving safety risk value. According to the invention, the comprehensiveness of driving safety protection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automotive safe driving technology, and in particular to a method and system for automotive driving safety warning based on driver emotions. Background Technology

[0002] Vehicle driving safety is a core issue in the transportation sector. Every year in my country, traffic accidents cause numerous injuries and property losses, with abnormal driver emotions being a key contributing factor. Research shows that negative emotions such as anger and anxiety shorten drivers' attention spans, delay their reaction times, and significantly increase the incidence of dangerous driving behaviors. Existing vehicle safety warning technologies primarily focus on monitoring vehicle driving status or road conditions, paying insufficient attention to the driver's emotional state. Some existing emotion-related warning solutions rely solely on a single data source for emotion recognition, resulting in low accuracy and poor robustness.

[0003] There is an urgent need for a safety warning method based on driver emotions, which can achieve accurate identification of emotions through multimodal fusion, thereby assessing risks and matching warning strategies to improve the comprehensiveness of driving safety protection. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a method and system for vehicle driving safety warning based on driver emotions.

[0005] A first aspect of this application provides a vehicle driving safety warning method based on driver emotions, including: Feature extraction is performed on multi-source data to obtain multimodal feature data; the multimodal feature data includes: vehicle driving features, facial features, and voice features; The multimodal feature data are fused based on an attention mechanism to obtain a target fused feature vector; The target fusion feature vector is input into the emotion recognition model to obtain the emotion recognition result, which includes emotion type and emotion intensity. In response to the emotion recognition result matching with a first preset emotion state, an early warning message is generated; In response to the driver's emotion recognition result matching the second preset emotion state in the first time, the driving safety risk value is calculated based on the emotion intensity and the confidence level of the multimodal feature data. The corresponding early warning strategy is determined based on the driving safety risk value.

[0006] A second aspect of this application provides a vehicle driving safety warning system based on driver emotions, comprising: The feature extraction module is used to extract features from multi-source data to obtain multimodal feature data; the multimodal feature data includes: vehicle driving features, facial features, and voice features; The feature fusion module is used to fuse the multimodal feature data based on an attention mechanism to obtain a target fused feature vector; An emotion recognition module is used to input the target fused feature vector into an emotion recognition model to obtain an emotion recognition result, which includes emotion type and emotion intensity. The first matching module is used to generate an early warning message in response to the matching of the emotion recognition result with the first preset emotion state; The second matching module is used to calculate the driving safety risk value based on the emotional recognition result of the driver in the first time and the second preset emotional state in response to the driver's emotional intensity and the confidence level of the multimodal feature data. The early warning strategy module is used to determine the corresponding early warning strategy based on the driving safety risk value.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described vehicle driving safety warning method based on driver emotions.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described vehicle driving safety warning method based on driver emotions.

[0009] The beneficial effects of the vehicle driving safety early warning method and system based on driver emotions provided in this application are as follows: Firstly, this application integrates multi-source features and strengthens feature weights based on an attention mechanism, solving the problem of single data being easily interfered with, making the identification of emotion type and intensity more accurate and providing a reliable basis for early warning. Secondly, it distinguishes between immediate dangerous emotions and persistent negative emotions, providing rapid early warning when immediately matching the first preset state, and calculating the risk value by combining emotion intensity and data confidence when continuously matching the second preset state, avoiding over-warning or delayed warnings, and improving the rationality of the early warning. The risk value calculation takes into account both the core impact of emotions and data reliability, making the early warning strategy more consistent with the actual risk level, and can intervene in advance in dangerous driving behaviors caused by high-risk emotions such as anger and fatigue, further reducing the probability of accidents. The multi-source data in this application is easily collected through in-vehicle devices, the attention mechanism and emotion model are adapted to in-vehicle computing scenarios, and the hierarchical strategy takes into account both safety and driving experience, facilitating practical application. Attached Figure Description

[0010] Figure 1 A flowchart illustrating a vehicle driving safety warning method based on driver emotions provided in an embodiment of this application; Figure 2 A structural block diagram of a vehicle driving safety warning system based on driver emotions provided in an embodiment of this application; Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0012] To make the purpose, technical solution, and advantages of this application clearer, the following will be described in conjunction with the appendix. Figure 1-3 The following is an explanation using specific examples.

[0013] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a vehicle driving safety warning method based on driver emotions provided in an embodiment of this application. The method includes: S101: Extract features from multi-source data to obtain multimodal feature data; the multimodal feature data includes: vehicle driving features, facial features and voice features.

[0014] In this embodiment, vehicle driving data is collected in real time via the vehicle's CAN bus, OBD interface, and onboard sensors. This vehicle driving data includes: driving speed, acceleration, brake pedal travel, throttle opening, steering wheel angle, gear information, current lane position, following distance, and vehicle light status. The driver's facial data is collected by an in-vehicle infrared camera mounted above the dashboard in front of the steering wheel. This driver facial data includes: facial image sequences, eye status, and facial muscle movements.

[0015] The driver's voice data is collected through an in-vehicle microphone array. The collection range includes the driver's normal speech and speech when emotionally agitated, and the decibel value of the in-vehicle ambient noise is recorded simultaneously during the collection.

[0016] This embodiment employs targeted preprocessing techniques to remove invalid data and standardize the data based on the noise characteristics of different data types. Specifically, Vehicle driving data preprocessing uses the 3σ criterion to remove outliers, linear interpolation to fill in short-term data gaps, and smooths continuous data such as acceleration and brake pedal travel. Finally, the data is standardized to the [0,1] interval.

[0017] Facial data preprocessing first uses a multi-task convolutional neural network to detect and align faces, cropping out standard facial regions. Then, contrast-limited adaptive histogram equalization is used to enhance facial feature details. Images with closed eyes or a side profile angle greater than 30° are marked as invalid data, and invalid data segments are filled by feature interpolation of adjacent valid frames.

[0018] The speech data preprocessing uses Wiener filtering to suppress environmental noise, removes silent segments by using a short-time energy threshold, and performs pre-emphasis, framing, and Hanning windowing on the effective speech segments.

[0019] This embodiment employs different technical features for feature extraction of different types of data. Specifically, vehicle driving feature extraction involves extracting 12-dimensional features from the preprocessed data, including statistical features, behavioral features, and state features. Statistical features include average speed over 5 seconds, speed standard deviation, maximum acceleration, and braking frequency. Behavioral features include accelerator pedal travel fluctuation rate, steering wheel angle change rate, following distance and speed matching degree, and lane departure frequency. State features include current gear, percentage of headlight usage time, and continuous driving duration.

[0020] Facial feature extraction: A pre-trained ResNet-50 model was used as the backbone network. The parameters of the first 10 layers were frozen and adjusted on the facial image dataset. The 256-dimensional features output by the fully connected layer were extracted. At the same time, 8-dimensional physiological features, such as blinking frequency, average eyelid opening and closing, corner of mouth angle, and frown coefficient, were manually extracted. The two were concatenated to obtain a 264-dimensional facial feature vector.

[0021] Speech feature extraction: Extracting acoustic and emotional features of speech, including: Mel frequency cepstral coefficients (MFCC, 13-dimensional, including first and second order differences, totaling 39 dimensions), fundamental frequency (3-dimensional), short-time energy (2-dimensional), and spectral entropy (1-dimensional), for a total of 45-dimensional speech feature vectors.

[0022] S102: Based on the attention mechanism, multimodal feature data are fused to obtain the target fused feature vector.

[0023] In this embodiment, a Temporal Independent Convolutional Network (TCN) is constructed to extract enhanced temporal features based on the temporal dynamic characteristics of each modality, thereby strengthening the capture of temporal dimension information of emotional changes. The network structure design includes: configuring independent TCN networks for vehicle driving features (12-dimensional), facial features (264-dimensional), and speech features (45-dimensional). The network structure is consistent but the parameters are trained independently. Specifically, it includes: an input layer (adapted to each modality feature dimension), two temporal convolutional layers (convolution kernel size 3, stride 1, padding='same', activation function is GELU), one pooling layer (max pooling, pooling kernel size 2), and one batch normalization layer.

[0024] The extracted modal features are organized into temporal feature matrices (vehicle: [50×12], face: [50×264], voice: [50×45]) according to a time step T=50 (number of samples within 5 seconds). These matrices are then input into the corresponding TCN networks to output the enhanced temporal feature matrices for each modality (with dimensions uniformly set to [25×128], where 25 is the time step after pooling and 128 is the feature dimension).

[0025] Using vehicle driving characteristics as the core query benchmark, key-value pairs are constructed based on facial and voice features, and dynamic masking is introduced to adapt to scenario requirements, thereby achieving accurate cross-modal weighted fusion.

[0026] S103: Input the target fusion feature vector into the emotion recognition model to obtain the emotion recognition result, which includes emotion type and emotion intensity; In this embodiment, the emotion recognition model adopts a dual-task model structure of feature adaptive adjustment and classification regression, as detailed below: The input layer receives the 256-dimensional target fusion feature vector from the output. The adaptive feature adjustment layer consists of one fully connected attention network (with 256 hidden layers, adjusting feature weights through an attention mechanism) and one batch normalization layer (to prevent overfitting).

[0027] The dual-task output layer includes an emotion classification branch and an emotion intensity regression branch. The emotion classification branch uses a 2-layer fully connected network with 128 hidden layers and the ReLU activation function; the output layer has 8 dimensions, corresponding to 8 emotion types: calm, joy, anger, anxiety, fear, sadness, surprise, and fatigue, and uses the softmax activation function to output the probability of each emotion type.

[0028] The emotion intensity regression branch uses a 1-layer fully connected network with a hidden layer dimension of 64 and an activation function of ReLU; the output layer dimension is 1 and uses the Sigmoid activation function to output the emotion intensity value (range [0,1], 0 is the weakest and 1 is the strongest).

[0029] The emotion recognition model in this embodiment employs the following training strategy: The dataset construction includes: collecting driver data from different age groups and driving scenarios, labeling 8 emotion types and emotion intensities, and constructing a training set with 100,000 samples, a validation set with 20,000 samples, and a test set with 30,000 samples.

[0030] Loss function design: A joint loss function is adopted, with the classification loss being the cross-entropy loss (weight 0.6) and the regression loss being the mean squared error loss (weight 0.4).

[0031] Training parameter settings: Adam optimizer is used, with an initial learning rate of 0.001, which decays to 0.8 every 5 epochs; batch size is 32, training epochs are 50, and an early stopping strategy is adopted (training stops if the validation set loss does not decrease for 3 consecutive epochs); overfitting is suppressed by Dropout (probability 0.2) and L2 regularization (coefficient 0.001).

[0032] In this embodiment, the model outputs complete recognition results and confidence information after inference. Once the emotion type is determined, the emotion type corresponding to the maximum probability output of the emotion classification branch is selected as the recognition result. If the maximum probability is less than 0.5 (low confidence), it is marked as an uncertain emotion.

[0033] Once the emotion intensity is determined, the output value of the emotion intensity regression branch is directly taken as the emotion intensity. If the emotion type is uncertain, the intensity value is set to 0.

[0034] S104: In response to the emotion recognition result matching the first preset emotion state, generate an early warning message.

[0035] In this embodiment, the first preset emotional state refers to an emotion that would immediately lead to dangerous driving behavior, including: anger (emotional intensity ≥ 0.7), fear (emotional intensity ≥ 0.6), and surprise (emotional intensity ≥ 0.8). An immediate warning is triggered when any of the emotional types and corresponding intensity conditions are met.

[0036] S105: In response to the driver's emotion recognition result matching the second preset emotion state in the first time, the driving safety risk value is calculated based on the emotion intensity and the confidence level of the multimodal feature data. In this embodiment, the term refers to emotions that continuously accumulate safety risks, including: anxiety (intensity ≥ 0.5), fatigue (intensity ≥ 0.6), and sadness (intensity ≥ 0.5). When the driver's emotion recognition result continuously matches this state within the first time frame (≥ 5 consecutive recognitions, with each recognition interval of 100ms), a risk value calculation is triggered. The timing of this first time frame can be adjusted according to different driving scenarios.

[0037] This embodiment calculates the driving safety risk value based on the intensity of emotion and the confidence level of multimodal feature data. Specifically, a baseline risk value is first calculated, which is the product of the intensity value corresponding to the second preset emotional state and the confidence level of the multimodal features. A correction coefficient is calculated based on the current vehicle driving data, and the baseline risk value is corrected according to the correction coefficient to obtain the final driving safety risk value.

[0038] S106: Determine the corresponding early warning strategy based on the driving safety risk value.

[0039] In this embodiment, four warning levels are defined based on the final obtained driving safety risk value, corresponding to different warning methods and intervention measures: Level 1 warning: The driving safety risk value is less than 0.3, which is considered low risk. The warning method is a text prompt on the vehicle's central control screen (green font, please keep a calm mood), without sound prompts to avoid interfering with driving; at the same time, the driver's emotional state is recorded and changes are continuously monitored.

[0040] Level 2 warning: The driving safety risk value is greater than or equal to 0.3 and less than 0.6, which is classified as low to medium risk. The warning is given by text and a gentle voice. The central control screen displays a green text prompt, and the car audio plays a gentle prompt (Your current mood may affect driving, so it is recommended to relax) at 70% of the current music volume.

[0041] A Level 3 warning indicates a driving safety risk value greater than or equal to 0.6 and less than 0.9, classifying it as medium to high risk. Warnings are issued via text, voice, vibration, and lights. The central control screen displays a flashing yellow text alert, and a clear voice prompt is given ("We have detected your continued anxiety. Please pay attention to driving safety and we recommend reducing your speed appropriately."). The instrument panel displays a flashing yellow warning light.

[0042] Level 4 warning indicates a driving safety risk value greater than or equal to 0.9, classified as high risk. It employs multi-dimensional strong warning + auxiliary intervention, with red text continuously flashing on the central control screen, high-frequency voice broadcast (high risk, it is recommended to immediately drive to a service area to rest), and the red warning light on the instrument panel remaining on, accompanied by slight seat vibration.

[0043] To avoid false and invalid warnings, this embodiment establishes a feedback adjustment mechanism. Specifically, driver feedback: a shortcut key on the steering wheel allows the driver to provide feedback on the warning results. If three consecutive false warnings are received, the recognition threshold for the current emotion type is lowered.

[0044] If the driver takes effective measures (such as slowing down and adjusting their posture) within 10 seconds after the warning and the emotional intensity decreases by ≥0.2, the warning level will be lowered; if the emotional intensity does not change or increases within 15 seconds after the warning, the warning level will be upgraded, up to a maximum of level four.

[0045] As can be seen from the above, this application integrates multi-source features and strengthens feature weights based on an attention mechanism, solving the problem of single data being easily interfered with, making the identification of emotion type and intensity more accurate and providing a reliable basis for early warning. Secondly, it distinguishes between immediate dangerous emotions and persistent negative emotions, providing rapid early warning when immediately matching the first preset state, and calculating the risk value by combining emotion intensity and data confidence when continuously matching the second preset state, avoiding over-warning or delayed warnings and improving the rationality of the warnings. The risk value calculation takes into account both the core impact of emotions and data reliability, making the early warning strategy more consistent with the actual risk level, and enabling early intervention in dangerous driving behaviors caused by high-risk emotions such as anger and fatigue, further reducing the probability of accidents. The multi-source data in this application is easily collected through in-vehicle devices, the attention mechanism and emotion model are adapted to in-vehicle computing scenarios, and the hierarchical strategy takes into account both safety and driving experience, facilitating practical application.

[0046] In one embodiment of this application, multimodal feature data is fused based on an attention mechanism to obtain a target fused feature vector, including: Vehicle driving features, facial features, and voice features are input into independent temporal convolutional networks to extract enhanced temporal features for each modality; The enhanced temporal features of each modality are encoded using a multi-head self-attention mechanism to obtain the context-aware feature vectors of each modality. The context-aware feature vectors of each modality are concatenated to obtain the feature vector; The feature vectors are passed through a fully connected layer and a softmax function to generate a cross-modal attention matrix; The feature vectors are weighted and fused based on the cross-modal attention matrix to obtain a preliminary fused vector; Based on the contextual information formed by vehicle driving status and driving environment information, the initial fusion vector is weighted through an attention gating mechanism to generate the target fusion feature vector.

[0047] In this embodiment, an intramodal attention module is constructed to address the differences in importance of different features within a single modality. An independent temporal convolutional network is constructed to extract enhanced temporal features based on the temporal dynamic characteristics of each modality feature, thereby strengthening the capture of temporal dimension information of emotional changes.

[0048] Specifically, the dimensions of vehicle driving features, facial features, and voice features are unified. The three are mapped to a 128-dimensional feature space through 1×1 convolution to obtain a modal feature matrix with a dimension of [T×128], where T is the time step and the number of samples within 5 seconds is taken, i.e., T=50.

[0049] The fully connected layer maps vehicle driving features, facial features, and voice features into query vectors and key vectors (both with dimensions [T×64]), and outputs intramodal enhanced features.

[0050] The temporal convolutional network structure is designed such that vehicle driving features, facial features, and speech features are configured with independent temporal convolutional networks. The structure of each network is consistent, but the parameters are trained independently. Specifically, it includes: an input layer (adapted to each modality feature dimension), two temporal convolutional layers (convolutional kernel size 3, stride 1, padding='same', activation function is GELU), one pooling layer (max pooling, pooling kernel size 2), and one batch normalization layer.

[0051] The extracted modal features are organized into temporal feature matrices with a time step of T=50, namely vehicle [50×12], face [50×264], and voice [50×45]. These matrices are then input into the corresponding temporal convolutional networks, and the output is the enhanced temporal feature matrix for each modality. The dimensions are uniformly set to [25×128], where 25 is the time step after pooling and 128 is the feature dimension.

[0052] Based on the complementarity between modalities, this embodiment constructs a single-modal feature enhanced by cross-modal attention module fusion: a multi-head self-attention mechanism is used to encode the temporal features enhanced by each modality, capturing the contextual association of features at different time steps within the modality.

[0053] Specifically, the enhanced features of the three modalities are concatenated into a global feature matrix, and global average pooling is used to obtain the global feature vectors of each modality, all with dimensions [1×128].

[0054] Using facial features as the core reference modality (the most direct reflection of driver emotions), the cosine similarity between the global feature vectors of each modality is calculated to obtain vehicle-face weights and voice-face weights; the final modality weights are obtained by normalization using the softmax function.

[0055] The enhanced features of the three modalities are multiplied by their corresponding modal weights, and then concatenated to obtain a fusion feature matrix. This matrix is ​​then reduced in dimensionality by a two-layer fully connected network (hidden layer dimension 256, activation function ReLU) to finally output a target fusion feature vector sequence of [T×128].

[0056] The number of attention heads is set to 8. The temporal feature matrix of each modality enhancement [25×128] is split into 8 sub-feature matrices of [25×16] according to the number of heads. The attention weights (the dimensions of query Q, key K, and value V are all [25×16]) are calculated for each sub-matrix and then weighted and summed to obtain the single-head encoded features.

[0057] The eight single-head encoded features are concatenated into a feature matrix of [25×128], and then integrated through a 1-layer fully connected network (output dimension 128, activation function GELU) to obtain the context-aware feature vector sequence of each modality (all dimensions are [25×128]). The global average pooling result of each sequence is taken as the context-aware feature vector of each modality (all dimensions are [1×128]).

[0058] This embodiment, based on the temporal continuity of driver emotions, supplements a temporal enhancement step. The target fusion feature vector sequence is input into a bidirectional LSTM network (128 hidden layer nodes, dropout probability 0.2) to extract the temporal features of emotion changes. The output of the last time step of the LSTM network is taken as the final target fusion feature vector (dimension [1×256]). Cross-modal feature fusion is achieved through concatenation and attention matrix weighting, highlighting complementary information between modalities.

[0059] Specifically, the context-aware feature vectors of vehicles, faces, and voices are concatenated sequentially to obtain a global feature vector of dimension [1×384]. This global feature vector is then input into a two-layer fully connected network (the first layer has a dimension of 256 and uses the GELU activation function; the second layer has a dimension of 384 and no activation function). The output is then normalized using the Softmax function to generate a cross-modal attention matrix of dimension [1×384]. Each element in the matrix corresponds to the weight coefficient of each dimension in the concatenated feature vector.

[0060] The global feature vector is multiplied element-wise with the cross-modal attention matrix to obtain a weighted feature vector of dimension [1×384]. Then, a dimensionality reduction is performed through a 1-layer fully connected network (output dimension 256, activation function GELU) to generate a preliminary fusion vector.

[0061] In this embodiment, a context information vector is constructed using vehicle driving status information and driving environment information to represent the characteristics of the current driving scenario.

[0062] The vehicle driving status information is obtained from the real-time data collected in step S101, including the current driving speed, longitudinal acceleration, braking frequency, steering wheel angle change rate, and following distance level, totaling 5 dimensions; the driving environment information is obtained through on-board sensors and the vehicle network, including ambient light intensity, road condition level, traffic flow level, weather type, and time period type, totaling 5 dimensions. The raw context information of 10 dimensions is input into a fully connected layer with batch normalization (input dimension 10, output dimension 256, activation function is LeakyReLU). A non-linear transformation maps the scene information to a feature space consistent with the temporal augmentation feature vector, resulting in a standardized context vector of dimension [1×256]. The context vector is then subjected to a Sigmoid activation function to constrain the output values ​​to the [0,1] interval, resulting in a gated weight vector of dimension [1×256]. Each element in the vector represents the weight coefficient of the corresponding feature dimension. The initial fusion vector is multiplied element-wise with the gating weight vector, and the feature components that match the current driving scenario are strengthened through the attention gating mechanism, finally generating a target fusion feature vector of [1×256].

[0063] As can be seen from the above, this embodiment extracts enhanced temporal features of each modality through independent temporal convolutional networks, effectively capturing the patterns of driver emotions changing over time and adapting to the continuous changes in emotions during driving. Secondly, the encoded context-aware feature vector can associate the preceding and following temporal information of each modality's data, avoiding misjudgments of emotions caused by isolated analysis of single-frame data. Adjusting the fusion weights based on vehicle driving status and driving environment information allows the target fused feature vector to adapt to different driving scenarios, improving the robustness of emotion recognition in complex scenarios.

[0064] In one embodiment of this application, a preliminary fused vector is obtained by weighted fusion of feature vectors based on a cross-modal attention matrix, including: Generate a query vector from vehicle driving features using linear transformation; generate key vectors and value vectors from facial features and voice features using linear transformation, respectively. A dynamic mask is constructed, which is used to adjust the weights of each component of facial and speech features in the key vector based on the vehicle's driving status and driving environment information. Calculate the attention score matrix based on the query vector, key vector, and the dimensions of the key vector; The attention score matrix is ​​corrected using a dynamic mask to obtain the corrected attention score matrix; The value vectors are weighted and fused based on the corrected attention score matrix to obtain weighted multimodal features; The weighted multimodal features and vehicle driving features are concatenated and dimensionality reduced to obtain a preliminary fusion vector.

[0065] In this embodiment, independent linear transformation matrices are designed for the enhanced features within the modality. Vehicle driving features are input into the linear transformation layer to generate a [T×64] query vector, used to focus on emotional needs associated with driving status. Facial features are transformed into facial key vectors through the linear transformation layer, and speech features are transformed into speech key vectors through the linear transformation layer. The facial key vector and the speech key vector are concatenated to obtain a [T×128] key vector. Similarly, facial features generate facial value vectors, and speech features generate speech value vectors. The facial value vector and the speech value vector are concatenated to obtain a [T×128] value vector. This embodiment collects real-time vehicle driving status information (speed V, longitudinal acceleration A, braking frequency F) and driving environment information (light intensity L, road condition level R, traffic flow D), which are then standardized to form a [T×6] scene feature matrix. The scene feature matrix is ​​input into a two-layer fully connected network (first layer with 128-dimensional ReLU activation, second layer with 128-dimensional Sigmoid activation) to generate a [T×128] dynamic mask, where mask element values ​​∈ [0,1]. High-value regions (>0.7) correspond to highly correlated features in the scene; for example, enhancing the weight of facial eyelid opening and closing features during high-speed driving. Low-value regions (<0.3) suppress noise features; for example, reducing the weight of non-emotional speech features in noisy environments. Based on the query vector, key vector, and key vector dimension, the original attention score matrix is ​​calculated to represent the association strength of features at different temporal positions. The dynamic mask is expanded by row copying to a mask matrix corresponding to the original attention score matrix, and then multiplied element-wise with the original score matrix to achieve correction, resulting in the corrected score matrix. The Softmax function is applied to the corrected score matrix to obtain a normalized attention weight matrix, with the sum of elements in each row being 1. Weighted multimodal features, which incorporate scene adaptation information from facial and speech data, are calculated using this normalized attention weight matrix. The weighted multimodal features are concatenated with vehicle driving features to form a fusion matrix, which is then input to a batch-normalized dimensionality reduction layer (1×1 convolutional kernel, 128 output channels, ReLU activation) to obtain a preliminary fusion vector.

[0066] As can be seen from the above, this embodiment uses vehicle driving features as the query vector and facial and voice features as the key and value vectors, respectively, to construct a clear modal interaction logic. Secondly, attention scores are corrected based on driving status and environmental information, strengthening effective features and weakening interfering features to address the problem of partial modal failure in different scenarios. Weighted multimodal features are fused with vehicle driving features and their dimensionality reduced, preserving key information while reducing data redundancy and improving the computational efficiency and recognition accuracy of subsequent emotion recognition models.

[0067] In one embodiment of this application, constructing a dynamic mask includes: In response to a vehicle speed greater than or equal to a first preset threshold, a first sub-mask is generated. The first sub-mask is used to assign a first enhancement coefficient to the key vector elements in the facial features corresponding to the eye region, and to assign a first weakening coefficient to all key vector elements corresponding to the speech features. In response to a steering wheel angle being greater than or equal to a second preset threshold, a second sub-mask is generated. The second sub-mask is used to assign a second enhancement coefficient to the key vector elements in the facial features corresponding to the mouth region. In response to the ambient light intensity being less than a third preset threshold, a third sub-mask is generated. The third sub-mask is used to assign a third enhancement coefficient to all key vector elements corresponding to the facial features. In response to an ambient noise intensity greater than a fourth preset threshold, a fourth sub-mask is generated. The fourth sub-mask is used to assign a fourth weakening coefficient to all key vector elements corresponding to the speech features. The first, second, third, and fourth sub-masks are merged to generate the final dynamic mask matrix, where the values ​​of each strengthening coefficient are greater than 1 and the values ​​of each weakening coefficient are between 0 and 1.

[0068] In this embodiment, based on key parameters of vehicle driving status and core indicators of driving environment, sub-masks are generated and fused according to different scenarios to achieve scenario-based adjustment of feature weights. The specific steps are as follows: The threshold values ​​for four preset scenarios were calibrated through real vehicle testing and statistical analysis. The first preset vehicle speed threshold is 100km / h (high-speed driving scenario), the second preset steering wheel angle threshold is 30° (sharp turning scenario), the third preset light intensity threshold is 30lux (low light scenario), and the fourth preset environmental noise intensity threshold is 60dB (high noise scenario). The coefficient range is defined as follows: the enhancement coefficient is 1.2-1.5, and the weakening coefficient is 0.3-0.7.

[0069] The method for generating the first sub-mask includes real-time acquisition of vehicle speed. If the vehicle speed is ≥100km / h, a first sub-mask of [T×128] is generated. Specifically, the feature dimension corresponding to the eye region in the facial key vector is assigned a first enhancement coefficient of 1.5, all 64 dimensions corresponding to the speech key vector are assigned a first weakening coefficient of 0.3, and the coefficients for the remaining dimensions are set to 1.0.

[0070] The method for generating the second sub-mask includes real-time acquisition of the steering wheel angle θ. If θ ≥ 30°, a second sub-mask of [T×128] is generated. The feature dimension corresponding to the mouth region in the facial key vector is assigned a second enhancement coefficient of 1.3, while the coefficients for the other dimensions are set to 1.0.

[0071] The method for generating the third submask includes real-time acquisition of ambient light intensity. If the ambient light intensity is <30 lux, a third submask of [T×128] is generated. All 64 dimensions corresponding to the facial key vector are assigned a third enhancement coefficient of 1.4, while the coefficients for the remaining dimensions are set to 1.0, to compensate for the loss of facial feature recognition under low light conditions.

[0072] The fourth sub-mask generation method includes real-time acquisition of ambient noise intensity. If the ambient noise intensity is >60dB, a fourth sub-mask of [T×128] is generated. In this sub-mask, all dimensions corresponding to the speech key vector are assigned a fourth weakening coefficient of 0.4, while the coefficients for the remaining dimensions are set to 1.0, to suppress the interference of high noise on speech features.

[0073] The sub-masks are fused by multiplying coefficients element by element. All elements of the sub-mask corresponding to the untriggered scene are set to 1.0, which does not affect the fusion result. The fused result is a final dynamic mask matrix of [T×128], with element values ​​ranging from [0.3, 1.5]. This embodiment achieves accurate enhancement of target features and effective suppression of noise features under different scenarios.

[0074] As can be seen from the above, this embodiment designs exclusive sub-masks for different driving states and environmental conditions. For example, it strengthens eye features when driving at high speeds and strengthens overall facial features in low-light environments, compensating for feature blurring caused by insufficient lighting and making feature weight adjustments more aligned with actual scenario requirements. Secondly, it weakens speech features when there is high environmental noise and when driving at high speeds, avoiding interference from unreliable modal data on the fusion results and improving the anti-interference capability of emotion recognition. Finally, based on multi-dimensional scene information such as vehicle speed, steering wheel angle, lighting, and noise, a dynamic mask matrix is ​​generated, enabling feature weights to adjust various variables in complex driving scenarios, further enhancing the versatility of the method.

[0075] In this embodiment, a vehicle driving safety warning method based on driver emotions further includes: The first image sharpness, first occlusion rate, and first illumination uniformity of the eye region in the facial features are calculated, and the eye confidence score is generated based on the weighted product of the first image sharpness, first occlusion rate, and first illumination uniformity. The first illumination uniformity is measured by calculating the variance of the pixel values ​​in the eye region. The smaller the variance, the more uniform the illumination. The second image sharpness, second occlusion rate, and second illumination uniformity of the mouth region in the facial features are calculated, and the mouth confidence score is generated based on the weighted product of the second image sharpness, second occlusion rate, and second illumination uniformity, wherein the second illumination uniformity is measured by calculating the variance of the pixel values ​​of the mouth region. The signal-to-noise ratio, audio amplitude stability, and background noise type of speech features are calculated. Speech confidence is generated based on the weighted product of the signal-to-noise ratio, audio amplitude stability, and background noise type. Audio amplitude stability is measured by calculating the coefficient of variation of the speech signal amplitude. Background noise type is identified by a pre-trained noise classification model and matched with a preset interference noise type. The higher the matching degree, the lower the confidence. The probability of emotion recognition based on eye features, mouth features, and voice features under similar driving conditions is extracted from historical data and ultimately verified as correct. This probability is used as the historical efficacy weight of each modality feature, and the historical efficacy weight is updated through an exponential decay model. The first enhancement coefficient is obtained by multiplying the base enhancement coefficient by the ocular confidence score and the ocular historical efficacy weights, and then non-linearly scaling it using the sigmoid function. The second enhancement coefficient is obtained by multiplying the base enhancement coefficient by the mouth confidence level and the mouth historical performance weights, and then nonlinearly scaling it using the same sigmoid function; The third enhancement coefficient is obtained by multiplying the base enhancement coefficient by the overall confidence of the facial features and the weight of the facial history efficacy. The overall confidence of the facial features is the geometric mean of the confidence of the eyes and the confidence of the mouth, and is non-linearly scaled by the sigmoid function. The first weakening coefficient is obtained by multiplying the base weakening coefficient by the inverse of the speech confidence and the inverse of the speech history effectiveness weight, and is non-linearly scaled by the sigmoid function to ensure that the coefficient value is between 0 and 1. The fourth weakening coefficient is obtained by multiplying the base weakening coefficient by the inverse of the speech confidence and the inverse of the speech history effectiveness weight, and then non-linearly scaling it using the same sigmoid function; Among them, the basic reinforcement coefficient and the basic weakening coefficient are dynamically initialized according to the preset driving safety level. The higher the driving safety level, the larger the basic reinforcement coefficient and the smaller the basic weakening coefficient.

[0076] In this embodiment, the first image sharpness is calculated by using the Laplacian operator to calculate the edge gradient variance of the eye region, normalized to [0,1], with higher values ​​indicating greater sharpness. The first occlusion rate is calculated by using a semantic segmentation model (U-Net) to identify occluders in the eye region, calculating the occlusion area ratio, and normalizing to [0,1], with higher values ​​indicating more severe occlusion; the effective coefficient is 1 - the first occlusion rate. The first illumination uniformity is calculated by using the variance of pixel grayscale values ​​in the eye region, normalized to [0,1], with smaller variance indicating greater uniformity. The confidence score of the eye region is obtained through weighted calculation. The weighting coefficients in the above weighting are based on historical experience, and the weighting coefficients emphasize sharpness and occlusion, as these two factors have a greater impact on the effectiveness of eye features.

[0077] The second image sharpness is calculated using the same method as for the eyes, employing the Laplacian operator to calculate the edge gradient variance of the mouth region, normalized to [0,1]. The second occlusion rate is calculated by identifying occluders in the mouth region using a semantic segmentation model (U-Net), calculating the occlusion area percentage, and normalizing it to [0,1]. The effective coefficient is then 1 - the second occlusion rate. The second illumination uniformity is calculated by determining the pixel grayscale variance of the mouth region. The mouth region is larger and less sensitive to illumination. A weighted average is used to obtain the confidence score for the mouth region. The weighting coefficients in the above weighting are based on historical experience and emphasize occlusion, as masks and other objects significantly affect mouth features.

[0078] The speech signal and background noise were separated by spectral analysis, and the signal-to-noise ratio (SNR) was calculated and normalized to [0,1]. The coefficient of variation (COP) of the short-time amplitude of the speech signal was calculated and normalized to [0,1], with higher values ​​indicating greater stability. A pre-trained noise classification model (CNN+LSTM) identified five types of noise: wind noise, engine noise, horn honking, passenger conversation, and no noise. The preset interference noise types were horn honking and wind noise; the higher the matching degree, the greater the confidence loss. The confidence of the speech features was calculated by weighting the SNR, COP, and noise type, with the weighting coefficients based on historical experience.

[0079] This embodiment extracts the accuracy of eye / mouth / voice features used individually for emotion recognition under similar driving conditions from historical data. Here, the historical performance weight of the eye is the probability of correct recognition of eye features; the historical performance weight of the mouth is the probability of correct recognition of mouth features; the overall historical performance weight of the face is the average of the historical performance weights of the eye and mouth; and the historical performance weight of the voice is the probability of correct recognition of voice features.

[0080] In this embodiment, the driving safety level is divided into 3 levels, which are determined based on a comprehensive assessment of vehicle speed, road conditions, and weather. Specifically, Level 1 (low risk) has a basic reinforcement coefficient of 1.2 and a basic weakening coefficient of 0.8; Level 2 (medium risk) has a basic reinforcement coefficient of 1.4 and a basic weakening coefficient of 0.6; and Level 3 (high risk) has a basic reinforcement coefficient of 1.6 and a basic weakening coefficient of 0.5.

[0081] The coefficients in this embodiment not only depend on the quality of real-time features, but also refer to the effectiveness of the features in historical data, such as whether eye features are correctly identified in similar scenarios, thus avoiding misjudgment based on a single dimension. Secondly, the sigmoid function is used to prevent the coefficients from becoming too heavy due to extreme values, thus maintaining the stability of the fusion. The base coefficients are more aggressive in high-risk scenarios to meet the priority requirements of safety warnings.

[0082] In one embodiment of this application, the value vectors are weighted and fused based on the corrected attention score matrix to obtain weighted multimodal features, including: Calculate the attention contribution rate for each channel in the corrected attention score matrix. The attention contribution rate is the ratio of the attention score of each channel to the total score. Channels with attention contribution rates less than or equal to a preset contribution rate threshold are identified as low contribution channels and removed to obtain the second attention matrix; Weighted multimodal features are calculated based on the second attention matrix.

[0083] In this embodiment, the corrected score matrix is ​​divided into T channels (corresponding to T temporal positions) by columns. The total score of each channel is calculated, and the attention contribution rate of each channel is the ratio of the attention score of each channel to the total score, with a value range of [0,1]. A preset contribution rate threshold Th=0.01 is set (verified by 100,000 sets of samples, removing channels with a proportion of <1% does not affect the recognition accuracy). All channels are traversed, and channels with an attention contribution rate less than or equal to the preset contribution rate threshold are identified as low contribution channels and removed. If the number of channels after removal is less than T / 2, only the top 20% of channels are removed in ascending order, finally obtaining the second attention matrix of [T×T'].

[0084] As can be seen from the above, this embodiment removes channels with extremely low attention contribution rates and eliminates invalid or interfering features, reducing the burden of redundant data on subsequent calculations and improving the model's computational efficiency. Secondly, by retaining high-contribution channels, the weighted multimodal features are more concentrated on core emotional information, further improving the accuracy of emotion recognition.

[0085] In one embodiment of this application, after generating a warning message in response to the emotion recognition result matching a first preset emotion state, the method further includes: Acquire at least one confirmation signal from the driver's head movements, voice input, or gestures; If the driver makes a confirmation signal that matches the preset confirmation action in the second time, the reliability of the confirmation is calculated based on the standardization of the confirmation signal, environmental interference factors and signal quality. If the confidence level is greater than or equal to the first confidence level threshold, then the warning level will be reduced. If the driver fails to provide a valid confirmation signal within the second time, the upgraded driving safety risk value is calculated based on the duration of the emotional intensity and the confidence level of the multimodal feature data, using a risk decay model that is adjusted based on the historical trend of emotional intensity and the environmental risk index. Based on the risk range of the upgraded driving safety risk value, the corresponding upgrade warning strategy is triggered.

[0086] In this embodiment, after the warning is triggered, multi-source confirmation signal acquisition is initiated simultaneously, supporting three types of signal input. These include head movements, such as those identified by an in-vehicle infrared camera, with a preset confirmation action of two consecutive nods (amplitude ≥30°, interval ≤1s). Voice input, such as that acquired through a microphone array, with preset confirmation commands like "received" or "understood," is also included. Gestures, identified by a steering wheel infrared sensor, are preset confirmation actions such as pressing a shortcut key twice or raising the thumb for 1 second, with a acquisition duration of a second time period (preset 5s). If a confirmation signal matching the preset action is acquired within 5s, the credibility is determined from three dimensions.

[0087] Specifically, the signal conformity score has a weight of 0.4; the evaluation is differentiated by signal type: head movements: calculate amplitude compliance rate (actual amplitude / standard amplitude), time interval deviation rate (|actual interval - standard interval| / standard interval), and movement fluency (standard deviation of angle change between adjacent frames); voice input: calculate MFCC feature matching degree, speech rate compliance rate (actual speech rate / standard speech rate), and the proportion of silence; gesture operation: calculate action accuracy, pressure compliance rate, and response delay (1 point for operation triggering to system recognition time ≤ 0.2s), and the scores of the three types of signals are normalized to [0,1].

[0088] The environmental adaptation coefficient, with a weight of 0.25, is dynamically adjusted based on real-time environmental parameters and corrected using a piecewise function: Head movements: The environmental adaptation coefficient is 1.0 when the light intensity is greater than or equal to 100 lux; when 30 lux ≤ light intensity < 100 lux, the environmental adaptation coefficient = 0.9 + 0.1 × (L-30) / 70; when the light intensity is less than 30 lux, the environmental adaptation coefficient is 0.6; Voice input: The environmental adaptation coefficient is 1.0 when the noise intensity is ≤ 40 dB, K_dist = 0.9 - 0.015 × (N-40) when 40 dB < N ≤ 60 dB, and K_dist = 0.6 when N > 60 dB (0 when disabled); Gesture operation: When the sensitivity of the capacitive sensor decreases in a low-temperature environment, the environmental adaptation coefficient is 0.9; when the steering wheel vibration frequency is > 5 Hz (bumpy road surface), the environmental adaptation coefficient is 0.85; in a normal environment, the environmental adaptation coefficient is 1.0.

[0089] The spatiotemporal consistency score has a weight of 0.2. This score assesses the spatiotemporal matching of the signal and the warning scenario: Temporal consistency: 1.0 point is awarded for a confirmation signal trigger time ≤ 1s and a warning notification time ≤ 3s, 0.8 points for 1s < Δt ≤ 3s, and 0.5 points for Δt > 3s; Spatial consistency: Head movements must ensure the face is centered within the camera's viewfinder, gestures must be within the sensor's effective recognition area, and voice must be located in the driver's seat area. The percentage of spatial compliance items is used to determine spatial consistency. The final spatiotemporal consistency score is determined based on both temporal and spatial consistency.

[0090] Historical matching weight, weight 0.15: Extract historical confirmation signal data of the driver in similar scenarios within the past 30 days, calculate the historical matching success rate (number of successful confirmations / total number of confirmations) and the average standard score. The formula is historical matching weight = 0.6 × historical success rate + 0.4 × (historical standard score / current standard score). If the historical data is less than 10, then take 0.8.

[0091] Overall credibility = (0.4 × signal standardization score + 0.25 × environment adaptation coefficient + 0.2 × spatiotemporal consistency score + 0.15 × historical matching weight) × signal validity coefficient, where single-source validity is 1.0, multi-source consistency is 1.1, and multi-source conflict is taken as the highest single-source score × 0.9, with a value range of [0, 1.1]. Values ​​greater than 1.0 are counted as 1.0.

[0092] If the driver's two consecutive confirmation actions are inconsistent, the credibility is reduced by 0.1; when the warning level is level four, the correction coefficient for high-risk scenarios is 1.05; when the equipment failure warning is issued, the credibility is multiplied by 0.9.

[0093] In this embodiment, the warning downgrade determination logic is as follows: A preset level-based credibility threshold is set, linked to the warning level: A Level 1 warning requires no confirmation; a Level 2 warning requires a confirmation credibility ≥ 0.6 to be downgraded; a Level 3 warning requires a confirmation credibility ≥ 0.7; and a Level 4 warning requires a confirmation credibility ≥ 0.8 (the threshold is increased by 0.05 for high-risk scenarios). If the confirmation credibility does not reach the corresponding threshold, it is determined as invalid confirmation, triggering a secondary prompt and extending the data collection time by 2 seconds.

[0094] The first confidence threshold is preset to 0.7. If the confidence level is ≥0.7, it is considered a valid confirmation, and the warning level is downgraded: the original level 3 warning is downgraded to level 2, and the original level 4 warning is downgraded to level 3. At the same time, the visual prompt is retained, and the voice prompt is turned off after 5 seconds. If the confidence level is <0.7, it is considered an invalid confirmation, the original warning level is maintained, and a prompt is made to confirm in a standardized manner.

[0095] Risk escalation calculation without confirmation: If no valid confirmation signal is collected within 5 seconds, the upgraded driving safety risk value is calculated using the risk attenuation model. Specifically, the input parameters include the duration of emotion intensity (the first preset emotion state matching duration, in seconds), multimodal feature confidence, emotion intensity change rate, and environmental risk index (1.2 for highway / congestion / nighttime scenarios, and 1.0 for others). The upgraded driving safety risk value was calculated using a risk attenuation model. ; in, The original risk value is 0.05×t, which is used to reflect the continuous accumulation of risk. max(r,0) is the maximum value between the rate of change of emotion intensity r and 0, considering only the risk contribution of the increase in intensity. This is the environmental risk index; if the confidence level of the multimodal features is less than 0.6, it is multiplied by a correction factor of 0.95 (to reduce the upgrade weight of low-confidence data). The upper limit is set to 1.5 (not exceeding the level 4 warning threshold); t is the duration of the emotion matching the second preset emotional state; r is the rate of change of emotion intensity.

[0096] As can be seen from the above, this embodiment allows drivers to provide feedback through head movements, voice, etc., avoiding a one-size-fits-all mandatory warning and improving driver acceptance and cooperation with the warning system. Secondly, based on signal standardization, environmental interference, and the reliability of signal quality calculations, it avoids incorrect warning adjustments caused by erroneous actions or environmental interference, balancing flexibility and accuracy. This embodiment also calculates the upgraded driving safety risk value using a risk attenuation model for situations where timely confirmation is not available, triggering corresponding strategies based on the risk range to ensure effective intervention of persistent feelings of danger and prevent risk accumulation.

[0097] In one embodiment of this application, if the driver issues a confirmation signal matching a preset confirmation action within a second time period, the confirmation confidence level is calculated based on the standardization of the confirmation signal, environmental interference factors, and signal quality. In response to a confirmation confidence level greater than or equal to a first confidence threshold, a warning level reduction process is executed, including: The pre-trained action recognition model is used to extract the feature vector of the driver's head movements, including nodding frequency, head shaking amplitude, head rotation angular velocity, and action duration. The similarity between the feature vector and the preset confirmation action template is calculated to obtain the action matching degree; Calculate the credibility of action confirmation based on action matching degree and environmental interference factors; If the confidence level of the action confirmation is greater than or equal to the first confidence threshold, the current warning level will be immediately reduced by one level. In response to an action confirmation confidence level falling between the first confidence level threshold and the second confidence level threshold, the warning interval is extended. The warning is cancelled when the credibility of the action confirmation is greater than or equal to the second credibility threshold for N consecutive monitoring periods. The first and second credibility thresholds are set individually based on the driver's historical behavior data.

[0098] In this embodiment, after the warning is triggered, the driver's head movement sequence is captured by an infrared camera (the second time is preset to 5 seconds), and input into a pre-trained 3D-CNN action recognition model (pre-trained on the Kinetics dataset, fine-tuned on the driver's head movement dataset, with a recognition accuracy of ≥95%). The model outputs a 4-dimensional core feature vector [f1, f2, f3, f4], where: f1 is the nodding frequency, in times / s, only counting effective nodding with an amplitude ≥30°; f2 is the head shaking amplitude, in degrees, taking the average of the maximum left and right head shaking angles; f3 is the head rotation angular velocity, in degrees / s, taking the peak value during the action; and f4 is the action duration, in seconds, the duration of continuous effective frames from the start to the end of the action.

[0099] The preset confirmation action template is called from the action template library. The matching degree between the feature vector and the template is calculated using cosine similarity. The value range is [0,1], and the closer it is to 1, the higher the matching degree.

[0100] Based on the driver's historical confirmation data over the past 3 months, thresholds are set individually using statistical methods: the first credibility threshold is 1.1 times the average credibility of historical valid confirmation actions, and not less than 0.7; the second credibility threshold is 0.8 times the first credibility threshold, and not less than 0.5; when new users have not accumulated historical data, the default thresholds are used, that is, the first credibility threshold is 0.7 and the second credibility threshold is 0.6.

[0101] If the credibility of the action confirmation is greater than or equal to the first credibility threshold, it is determined to be a high-credibility confirmation, and the current warning level is immediately reduced by one level (e.g., level three is downgraded to level two, level four is downgraded to level three).

[0102] If the credibility of the action confirmation is between the first credibility threshold and the second credibility threshold, it is determined to be a medium credibility confirmation, and the warning interval is extended from the original 5s to 10s (level 3 warning) or from 3s to 6s (level 4 warning).

[0103] The preset monitoring cycle is 2 seconds. If the action confirmation credibility is greater than or equal to the second credibility threshold for N=3 consecutive cycles (6 seconds in total), it is judged as a continuous credible confirmation, and the warning is cancelled. The central control screen displays that the warning has been cancelled. All warning signs are turned off within 10 seconds after the safe driving prompt. If the action confirmation credibility is less than the second credibility threshold in any cycle during this period, the count is reset, and the extended warning interval is maintained.

[0104] As can be seen from the above, this embodiment extracts multi-dimensional head movement features such as nodding frequency and head shaking amplitude, and performs similarity matching with preset templates to accurately determine the driver's confirmation intention, reducing false alarms. Secondly, it distinguishes between three scenarios—lowering the warning level, extending the warning interval, and canceling the warning—based on a credibility threshold, avoiding the limitations of a single processing method, neither ignoring potential risks nor excessively interfering with driving. Furthermore, it customizes the first and second credibility thresholds based on the driver's historical behavior data to adapt to different drivers' movement habits, further enhancing personalization and applicability.

[0105] In one embodiment of this application, based on the risk range of the upgraded driving safety risk value, a corresponding upgrade warning strategy is triggered, including: When the upgraded driving safety risk value is in the first-level risk range, the intensity of the warning prompt is enhanced, and the prompt volume and vibration frequency are increased based on the first step length. When the upgraded driving safety risk value is in the level 2 risk range, the vehicle intervention system is activated to limit the vehicle's maximum acceleration; When the upgraded driving safety risk value is in the level three risk range, the vehicle intervention system is activated to limit the vehicle's maximum speed and send a confirmation request to the driver.

[0106] In this embodiment, based on the risk value range corresponding to the original warning level, the upgraded risk range is set as follows: Level 1 risk range: the upgraded driving safety risk value is less than 1.1 times the original warning level risk limit; Level 2 risk range: the original limit of 1.1 times ≤ the upgraded driving safety risk value < the original limit of 1.3 times; Level 3 risk range: the upgraded driving safety risk value is greater than or equal to the original limit of 1.3 times, ensuring that the upgrade range matches the risk increase.

[0107] The upgraded driving safety risk value is in the first-level risk range strategy, which determines that the risk has increased slightly and executes the warning intensity enhancement operation: increase the voice warning volume based on the first step length (preset 5dB-85dB) and optimize the sound quality (improve mid-to-high frequency clarity); increase the vibration frequency based on the first step length (preset 0.5Hz-3Hz); add risk increase prompts to the voice content, and darken the text color of the prompt on the central control screen from the original level corresponding color.

[0108] The upgraded driving safety risk value is in the level 2 risk range, which is judged as a moderate increase in risk. The vehicle intervention system is activated to implement acceleration limitation, dynamically setting the maximum acceleration threshold based on the current vehicle speed. At the same time, the central control screen displays that acceleration has been limited and please pay attention to the emotional prompt; the voice prompt interval has been shortened from the original 5 seconds to 3 seconds.

[0109] The upgraded driving safety risk value is in the level three risk range, which is considered a severe increase in risk. A high-level intervention and secondary confirmation are initiated. The vehicle intervention system limits the maximum speed to 90% of the current speed and not lower than 80% of the current road speed limit. A confirmation request is sent with a highly recognizable prompt tone and voice broadcast. At the same time, the backlight of the steering wheel confirmation button flashes to indicate the prompt, and the collection time is extended to 4 seconds.

[0110] If the emotional intensity decreases by more than or equal to 0.2 within any monitoring period (2s) after the upgrade, or if the driver takes proactive measures such as slowing down (vehicle speed reduced by ≥10%) or turning on hazard lights, the intervention will be lifted.

[0111] As can be seen from the above, this embodiment only enhances the warning prompts for Level 1 risks, limits acceleration for Level 2 risks, and limits vehicle speed and requests confirmation for Level 3 risks, gradually increasing the intensity of intervention to ensure safety while preserving the driver's autonomy to the greatest extent. Differentiated strategies are developed for different risk ranges to avoid excessive intervention in low-risk scenarios and insufficient intervention in high-risk scenarios, thus improving the targeting and effectiveness of warnings. In high-risk situations, proactive intervention measures such as limiting acceleration and vehicle speed directly reduce the possibility of dangerous operations caused by the driver's emotional instability, further strengthening safety assurance.

[0112] In one embodiment of this application, a vehicle driving safety warning method based on driver emotions further includes: Acquire vehicle external environment data, including traffic density, weather conditions, and road type; Calculate the environmental risk index based on environmental data; Based on the environmental risk index, the classification threshold of the emotion recognition model is adjusted. The higher the environmental risk index, the higher the sensitivity of the emotion recognition model. This is achieved by lowering the confidence threshold for emotion type judgment.

[0113] In this embodiment, the environmental risk index is calculated using a weighted summation model based on preprocessed external environmental data. , ; Among them: Weight setting (calibrated based on real vehicle accident data): Traffic density weight =0.4, Weather Condition Weight =0.3, Road type weight =0.3; ②Indicator value: Traffic density after standardization (0 is the lowest, 1 is the highest). The standardized weather risk value (0 for sunny, 1 for heavy rain / snow). The standardized road risk value (0 for residential roads, 1 for highways); the environmental risk index ranges from [0,1], and is divided into low environmental risk (...). <0.3), medium environmental risk (0.3≤ <0.7), high environmental risk ( (≥0.7) Three levels.

[0114] The classification confidence threshold is adjusted based on the environmental risk index to achieve sensitivity adjustment for environmental adaptation. Specifically, the classification confidence threshold is increased by 0.05 for low environmental risk (reducing sensitivity and minimizing false alarms), the baseline threshold is used for medium environmental risk, and the classification confidence threshold is decreased by 0.05 for high environmental risk (increasing sensitivity and avoiding missed alarms).

[0115] As can be seen from the above, this embodiment calculates the environmental risk index based on traffic density, weather, and road type, enabling the early warning strategy to adapt to driving environments with different levels of danger. Secondly, adjusting the risk interval threshold according to the environmental risk index avoids the inapplicability of fixed thresholds in different environments, making risk assessment more closely aligned with real-world scenarios. In high-risk environments, the model's sensitivity is increased to identify dangerous emotions earlier; in low-risk environments, sensitivity is appropriately reduced to decrease false alarms, further achieving precise adaptation between strict early warning for high-risk environments and minimal interference for low-risk environments.

[0116] In one embodiment of this application, a vehicle driving safety warning method based on driver emotions further includes: The driver's first audio is segmented to obtain multiple first audio segments; features are extracted from the multiple first audio segments to obtain multiple first feature segments; dimensionality reduction is performed on the multiple feature segments to obtain multiple target feature segments; the multiple target feature segments are concatenated to obtain the first audio feature. The identity information of the first user is determined based on the matching degree between the first user's first audio feature and the preset audio feature; Based on the driver's identity information, load the corresponding personalized emotion baseline and warning parameters; When identity verification fails, default warning parameters are used, and anonymous driving data is recorded.

[0117] In this embodiment, the first audio triggered after the driver gets into the vehicle is segmented, and a sliding window method (window length 0.5s, step size 0.25s) is used to obtain 6 first audio segments. For each audio segment, 39-dimensional Mel-frequency cepstral coefficients (MFCC, including first and second order differences) and 13-dimensional linear predictive cepstral coefficients (LPCC) are extracted to form a 52-dimensional first feature segment. Principal component analysis (PCA) is used for dimensionality reduction, retaining 95% of the feature information, resulting in a 20-dimensional target feature segment. Audio feature templates for preset users are retrieved from the vehicle identity feature library (each user stores 3-5 baseline features from different scenarios, optimized using Euclidean distance clustering). The cosine similarity (value [0,1]) between the first audio features and each template is calculated. The preset matching threshold is 0.85. If there is a template with a cosine similarity greater than or equal to the preset matching threshold, the user corresponding to the template with the highest matching degree is the first user, and the identity information is output. If all cosine similarities are less than the preset matching threshold, the identity recognition is deemed to have failed.

[0118] After successful identity verification, the corresponding parameters are loaded from the user's personalized database; Personalized emotional baseline: The baseline intensity range and characteristic distribution of eight types of emotions, such as calmness and pleasure, for this user in different scenarios (highway / city / traffic jam); Warning parameters include preferred warning methods (voice timbre / visual location), emotion threshold offset, and warning message style.

[0119] When identity verification fails, load the default warning parameters and automatically record anonymous driving data.

[0120] A first preset emotional state triggers an immediate alert. If identity verification is successful, the driver's personalized emotional intensity threshold is loaded; if identity verification fails, the default threshold is used. An immediate alert is triggered when any emotional type and corresponding intensity condition are met.

[0121] The second preset emotional state applies a personalized threshold to the driver if identity recognition is successful; otherwise, a default threshold is used. If the driver's emotional recognition result consistently matches this state within the first period, a risk value calculation is triggered.

[0122] As can be seen from the above, this embodiment achieves driver identification through audio segmentation, feature extraction, and matching, requiring no additional hardware, thus reducing deployment costs, and the operation is non-intrusive. Secondly, by loading a unique emotion baseline and warning parameters based on the driver's identity, the problem of insufficient adaptability of general parameters is avoided. When identification fails, default parameters are used and anonymous data is recorded, which neither affects the normal operation of the system nor compromises driver privacy, further enhancing user trust.

[0123] In one embodiment of this application, in response to the driver's emotion recognition result matching a second preset emotion state within a first time period, a driving safety risk value is calculated based on the emotion intensity and the confidence level of multimodal feature data, including: Monitor the duration of emotion recognition results within the first time frame and calculate the rate of change in emotion intensity; The weight of emotion intensity in risk calculation is adjusted based on duration and rate of change, with a larger weight adjustment coefficient for longer duration or higher rate of change. Based on the adjusted emotion weights and the confidence scores of multimodal feature data, a nonlinear regression model is used to calculate the driving safety risk value, where the confidence scores are determined based on the attention scores of each modality feature during the fusion process.

[0124] As can be seen from the above, this embodiment, based on the duration and rate of change of emotions, aligns with the cumulative and sudden characteristics of emotions, making the risk value calculation more closely reflect the actual level of danger. Secondly, the longer the duration and the higher the rate of change, the higher the weight is given to the emotion intensity, highlighting the core risk variable and improving the rationality and relevance of the risk value. Based on the adjusted emotion weights and the confidence level of multimodal features, driving safety risks are quantified through a nonlinear model, providing reliable data support for the formulation of early warning strategies.

[0125] In one embodiment of this application, with the driver's authorization, the driver's historical emotion recognition results and driving behavior data are collected to construct a personalized emotion baseline and driving habit model, wherein the data collection is anonymized. During the real-time warning process, the deviation between the current emotion recognition result and the personalized emotion baseline is calculated. The deviation is calculated using statistical bias based on the frequency and intensity of the emotion type. Based on the deviation and driving habit model, the driving safety risk value is corrected, and for drivers who habitually drive aggressively, the correction coefficient of the risk value is increased; Based on the corrected risk value, a matching warning method is selected from the warning strategy library, prioritizing the triggering of graded voice prompts and visual warnings, and triggering active vehicle control only when the risk value exceeds the high-risk threshold.

[0126] In this embodiment, the driver is prompted to authorize personalized data collection, and is clearly informed of the data's purpose (only for building a personal model), storage method (local encrypted storage + cloud backup encryption), and retention period (default 1 year, adjustable). After the driver authorizes, their historical emotion recognition results (emotion type, intensity, recognition time, scene tags) and driving behavior data (vehicle speed, acceleration, brake / accelerator operation, steering habits, following distance, etc.) are collected, with the collection frequency consistent with real-time data. Data anonymization is performed using a "de-identification + hash mapping" method: direct identifiers such as name and ID number are removed, and the driver ID is hashed using SHA-256 to ensure that the data cannot be reverse-linked to personal identity. For unauthorized drivers, only anonymous driving data is collected, without associating it with personal identifiers.

[0127] This embodiment is based on 30 consecutive days of authorized driving data, categorized and statistically analyzed according to scenario dimensions (highway, urban main road, congested, rural road). Specifically, the frequency distribution of emotion types is calculated by determining the percentage of occurrence of eight emotion types in each scenario; the emotion intensity benchmark is calculated by determining the mean and standard deviation of the corresponding emotion intensity in each scenario; the duration distribution of emotions is statistically analyzed using a sliding window method (10-minute window) to determine the normal duration threshold for each emotion; ultimately, a scenario-based personalized emotion baseline is formed.

[0128] This embodiment uses historical driving behavior data (vehicle speed, acceleration, braking frequency, steering angle, etc.) and employs the K-means clustering algorithm to classify drivers into three categories: aggressive, moderate, and cautious. ① Aggressive characteristics: average vehicle speed ≥ 90% of the road speed limit, frequency of rapid acceleration (acceleration ≥ 2m / s²) ≥ 5 times / hour, frequency of rapid braking (deceleration ≤ -2m / s²) ≥ 4 times / hour; ② Moderate characteristics: average vehicle speed 70%-90% of the road speed limit, frequency of rapid acceleration / deceleration ≤ 2 times / hour, steering angle change rate ≤ 5° / s; ③ Cautious characteristics: average vehicle speed ≤ 70% of the road speed limit, following distance ≥ 1.2 times the safe distance, and 100% compliance rate in using lights during nighttime driving. The model outputs driving habit labels and key behavioral parameters.

[0129] The system tracks the emotion recognition results in real time within the first hour and records the continuous effective duration of the match between the emotion and the second preset emotion state. The rate of change of emotion intensity is calculated by linear fitting method. A rate of change of emotion intensity greater than 0 indicates an increase in intensity, while a rate of change of emotion intensity less than or equal to 0 indicates stability or a decrease. A weighted adjustment coefficient is constructed based on duration and rate of change to amplify the risk contribution of long-term or escalating emotions. The weighted adjustment coefficient increases by 0.03 for every 1 second increase in duration. The effect of increased intensity is only considered when the rate of change in emotion intensity is greater than 0; for every 0.01 increase in the rate of change in emotion intensity, the weighted adjustment coefficient increases by 0.15. To avoid excessive weight amplification, the value range of the weighted adjustment coefficient is limited to [1.0, 2.0]. The adjusted emotion intensity is the product of the average emotion intensity over the first time period and the weighted adjustment coefficient. As can be seen from the above, this embodiment constructs a unique emotional baseline based on the driver's historical data, avoiding the use of general standards to judge emotional abnormalities and improving the accuracy of emotional abnormality judgment. For drivers with habitually aggressive driving, a risk correction coefficient is added, increasing the warning intensity for high-risk groups; for drivers with stable driving, excessive intervention is reduced, balancing personalization and safety. Secondly, non-interventional warnings such as voice prompts and visual warnings are prioritized, triggering vehicle control only in high-risk situations, respecting the driver's driving control to the greatest extent while ensuring a safety net under extreme risks. This embodiment collects data anonymously with the driver's authorization, providing sample support for model optimization while protecting driver privacy.

[0130] Corresponding to the vehicle driving safety warning method based on driver emotions in the above embodiment, Figure 2 This is a structural block diagram of a vehicle driving safety warning system based on driver emotions, provided as an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2 The vehicle driving safety warning system 20 based on driver emotions includes: a feature extraction module 21, a feature fusion module 22, an emotion recognition module 23, a first matching module 24, a second matching module 25, and a warning strategy module 26.

[0131] The feature extraction module 21 is used to extract features from multi-source data to obtain multimodal feature data. The multimodal feature data includes vehicle driving features, facial features, and voice features. Feature fusion module 22 is used to fuse multimodal feature data based on an attention mechanism to obtain a target fused feature vector; The emotion recognition module 23 is used to input the target fusion feature vector into the emotion recognition model to obtain the emotion recognition result, which includes emotion type and emotion intensity. The first matching module 24 is used to generate an early warning message in response to the matching of the emotion recognition result with the first preset emotion state; The second matching module 25 is used to calculate the driving safety risk value based on the emotional intensity and the confidence level of the multimodal feature data in response to the driver's emotion recognition result matching the second preset emotional state in the first time. The early warning strategy module 26 is used to determine the corresponding early warning strategy based on the driving safety risk value.

[0132] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the feature extraction module 21, feature fusion module 22, emotion recognition module 23, first matching module 24, second matching module 25, and early warning strategy module 26 are shown.

[0133] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0134] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0135] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0136] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in any embodiment of the vehicle driving safety warning method based on driver emotions provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.

[0137] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0138] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0139] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0143] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0144] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for vehicle driving safety early warning based on driver emotions, characterized in that, include: Feature extraction is performed on multi-source data to obtain multimodal feature data; The multimodal feature data includes: vehicle driving features, facial features, and voice features; The multimodal feature data are fused based on an attention mechanism to obtain a target fused feature vector; The target fusion feature vector is input into the emotion recognition model to obtain the emotion recognition result, which includes emotion type and emotion intensity. In response to the emotion recognition result matching with a first preset emotion state, an early warning message is generated; In response to the driver's emotion recognition result matching the second preset emotion state in the first time, the driving safety risk value is calculated based on the emotion intensity and the confidence level of the multimodal feature data. The corresponding early warning strategy is determined based on the driving safety risk value.

2. The vehicle driving safety early warning method based on driver emotions according to claim 1, characterized in that, The process of fusing the multimodal feature data based on an attention mechanism to obtain a target fused feature vector includes: The vehicle driving features, facial features, and voice features are respectively input into independent temporal convolutional networks to extract enhanced temporal features for each modality; The enhanced temporal features of each modality are encoded using a multi-head self-attention mechanism to obtain the context-aware feature vectors of each modality. The context-aware feature vectors of each modality are concatenated to obtain a feature vector; The feature vectors are then passed through a fully connected layer and a Softmax function to generate a cross-modal attention matrix; The feature vectors are weighted and fused based on the cross-modal attention matrix to obtain a preliminary fused vector; Based on the contextual information composed of the vehicle driving state and driving environment information, the initial fusion vector is weighted through an attention gating mechanism to generate the target fusion feature vector.

3. The vehicle driving safety early warning method based on driver emotions according to claim 2, characterized in that, The step of weighted fusing the feature vectors based on the cross-modal attention matrix to obtain a preliminary fused vector includes: The vehicle driving features are transformed into a query vector through a linear transformation; the facial features and voice features are transformed into a key vector and a value vector, respectively, through a linear transformation. A dynamic mask is constructed, which is used to adjust the weights of each component of facial features and speech features in the key vector according to the vehicle driving status and driving environment information. Calculate the attention score matrix based on the query vector, key vector, and the dimensions of the key vector; The attention score matrix is ​​corrected using the dynamic mask to obtain the corrected attention score matrix. The value vectors are weighted and fused based on the corrected attention score matrix to obtain weighted multimodal features; The weighted multimodal features and vehicle driving features are concatenated and their dimensions reduced to obtain the preliminary fusion vector.

4. The vehicle driving safety early warning method based on driver emotions according to claim 3, characterized in that, The construction of the dynamic mask includes: In response to a vehicle speed greater than or equal to a first preset threshold, a first sub-mask is generated. The first sub-mask is used to assign a first enhancement coefficient to the key vector elements in the facial features corresponding to the eye region, and to assign a first weakening coefficient to all key vector elements corresponding to the speech features. In response to a steering wheel angle greater than or equal to a second preset threshold, a second sub-mask is generated. The second sub-mask is used to assign a second enhancement coefficient to the key vector elements in the facial features corresponding to the mouth region. In response to the ambient light intensity being less than a third preset threshold, a third sub-mask is generated, which is used to assign a third enhancement coefficient to all key vector elements corresponding to the facial features. In response to an ambient noise intensity greater than a fourth preset threshold, a fourth sub-mask is generated, which is used to assign a fourth weakening coefficient to all key vector elements corresponding to the speech features. The first, second, third, and fourth sub-masks are merged to generate the final dynamic mask matrix.

5. The vehicle driving safety early warning method based on driver emotions according to claim 3, characterized in that, The weighted fusion of the value vector based on the corrected attention score matrix yields weighted multimodal features, including: Calculate the attention contribution rate for each channel in the corrected attention score matrix, where the attention contribution rate is the ratio of the attention score of each channel to the total score; Channels with attention contribution rates less than or equal to a preset contribution rate threshold are identified as low contribution channels and removed to obtain a second attention matrix; The weighted multimodal features are calculated based on the second attention matrix.

6. The vehicle driving safety early warning method based on driver emotions according to claim 1, characterized in that, After generating an early warning message in response to the emotion recognition result matching the first preset emotion state, the system also includes: Acquire at least one confirmation signal from the driver's head movements, voice input, or gestures; If the driver makes a confirmation signal that matches the preset confirmation action in the second time, the reliability of the confirmation is calculated based on the standardization of the confirmation signal, environmental interference factors and signal quality. If the confirmation confidence level is greater than or equal to the first confidence level threshold, then the warning level is reduced. If the driver fails to provide a valid confirmation signal within the second time, the upgraded driving safety risk value is calculated based on the duration of the emotional intensity and the confidence level of the multimodal feature data, wherein the risk attenuation model is adjusted based on the historical trend of the emotional intensity and the environmental risk index. Based on the risk range of the upgraded driving safety risk value, the corresponding upgrade warning strategy is triggered.

7. A vehicle driving safety early warning method based on driver emotions according to claim 6, characterized in that, If the driver issues a confirmation signal matching the preset confirmation action within the second time period, the confirmation credibility is calculated based on the standardization of the confirmation signal, environmental interference factors, and signal quality. In response to the confirmation credibility being greater than or equal to a first credibility threshold, a warning level reduction process is executed, including: The pre-trained action recognition model is used to extract the feature vector of the driver's head movements, including nodding frequency, head shaking amplitude, head rotation angular velocity, and action duration. The similarity between the feature vector and the preset confirmation action template is calculated to obtain the action matching degree; The credibility of action confirmation is calculated based on the action matching degree and environmental interference factors. If the confidence level of the action confirmation is greater than or equal to the first confidence threshold, the current warning level shall be immediately reduced by one level. In response to the action confirmation confidence level being between the first confidence level threshold and the second confidence level threshold, the warning notification interval is extended; The warning is cancelled when the credibility of the action confirmation is greater than or equal to the second credibility threshold for N consecutive monitoring periods. The first confidence threshold and the second confidence threshold are set individually based on the driver's historical behavior data.

8. A vehicle driving safety early warning method based on driver emotions according to claim 6, characterized in that, The upgraded driving safety risk value is located within a certain risk range, triggering a corresponding upgrade warning strategy, including: When the upgraded driving safety risk value is in the first-level risk range, the intensity of the warning prompt is enhanced, and the prompt volume and vibration frequency are increased based on the first step length. When the upgraded driving safety risk value is in the level 2 risk range, the vehicle intervention system is activated to limit the vehicle's maximum acceleration; When the upgraded driving safety risk value is in the level three risk range, the vehicle intervention system is activated to limit the vehicle's maximum speed and send a confirmation request to the driver.

9. A vehicle driving safety early warning method based on driver emotions according to claim 8, characterized in that, Also includes: Acquire vehicle external environment data, including traffic density, weather conditions, and road type; Calculate the environmental risk index based on the environmental data; The threshold for adjusting the risk range is based on the aforementioned environmental risk index; Based on the environmental risk index, the classification threshold and feature fusion weights of the emotion recognition model are adjusted.

10. A vehicle driving safety early warning system based on driver emotions, characterized in that, include: The feature extraction module is used to extract features from multi-source data to obtain multimodal feature data; The multimodal feature data includes: vehicle driving features, facial features, and voice features; The feature fusion module is used to fuse the multimodal feature data based on an attention mechanism to obtain a target fused feature vector; An emotion recognition module is used to input the target fused feature vector into an emotion recognition model to obtain an emotion recognition result, which includes emotion type and emotion intensity. The first matching module is used to generate an early warning message in response to the matching of the emotion recognition result with the first preset emotion state; The second matching module is used to calculate the driving safety risk value based on the emotional recognition result of the driver in the first time and the second preset emotional state in response to the driver's emotional intensity and the confidence level of the multimodal feature data. The early warning strategy module is used to determine the corresponding early warning strategy based on the driving safety risk value.