Driving safety analysis method and related device based on speech emotion recognition

By obtaining a multi-level emotion feature extraction and safety risk prediction model for driving voice data streams, the problem of difficulty in capturing driver emotional fluctuations in existing technologies is solved, early prediction and precise intervention of driving safety risks are achieved, and the incidence of traffic accidents is reduced.

CN120340541BActive Publication Date: 2025-09-05GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510839215.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-05
Estimated Expiration
2045-06-23

Smart Images

  • Figure CN120340541B_ABST
    Figure CN120340541B_ABST
Patent Text Reader

Abstract

The present invention provides a driving safety analysis method and related device based on speech emotion recognition. The method obtains a driving speech data stream collected by a target vehicle in a driving environment, extracts emotional features from the driving speech data stream, generates a multi-level emotional feature set, and inputs the multi-level emotional feature set into a safety risk prediction model to generate a driver's safety behavior score curve within a preset time period. Based on risk nodes in the safety behavior score curve that exceed a preset threshold, driving intervention prompt information corresponding to the risk nodes is generated and sent to the target vehicle's onboard terminal for real-time display. The present invention can reduce hardware deployment costs, effectively suppress interference with analysis results caused by fluctuations in speech quality in complex driving environments, and improve the accuracy of driving safety warnings and reduce the incidence of traffic accidents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a driving safety analysis method based on speech emotion recognition and a related device. Background Art

[0002] Driving safety analysis aims to improve driving safety by monitoring driver status and vehicle behavior to identify potential risks. Existing driving safety analysis typically uses on-board sensors to collect operational data such as steering angle and braking frequency, or biosensors to monitor physiological indicators such as the driver's heart rate and eye movement. These methods then generate safety warnings based on threshold-based judgment rules. However, indirect inference of the driver's emotional state in existing methods is subject to lags. This is especially true when emotional fluctuations haven't yet triggered overt operational anomalies. Traditional sensors are unable to capture dangerous emotional characteristics such as anger and anxiety implicit in speech. Furthermore, static warning mechanisms based on fixed thresholds struggle to adapt to the dynamic correlation between emotion and operational behavior in complex road environments, resulting in high false alarm rates and weak risk attribution. Furthermore, using independent physiological sensors to monitor emotional state presents challenges such as high hardware deployment costs and susceptibility to environmental interference. Traditional speech emotion recognition technology only classifies emotions based on discrete keywords and cannot parse the semantic association between emotional intensity fluctuations in continuous speech segments and vehicle interaction commands. This results in significant delays in warning responses for scenarios such as distracted driving and sudden road rage.

[0003] It should be noted that the above background introduction of the present invention is only to introduce the technical field to which the present invention relates. If the prior art does not disclose the above technical features and technical problems, it cannot be used as a basis for judging the creativity of the present invention. Summary of the Invention

[0004] The present invention provides a driving safety analysis method based on speech emotion recognition and a related device.

[0005] In a first aspect, an embodiment of the present invention provides a driving safety analysis method based on speech emotion recognition, the method comprising: obtaining a driving speech data stream collected by a target vehicle in a driving environment, the driving speech data stream comprising continuous speech segments of the driver within a preset time period and interactive instructions related to vehicle control; performing emotional feature extraction on the driving speech data stream to generate a multi-level emotional feature set, wherein the multi-level emotional feature set comprises real-time emotional fluctuation trajectories corresponding to the continuous speech segments and semantic emotional intensity parameters corresponding to the interactive instructions; inputting the multi-level emotional feature set into a safety risk prediction model to generate a safety behavior score curve of the driver within the preset time period, the safety behavior score curve being used to characterize the correlation between the driver's emotional stability and driving operation risks at different time nodes; generating driving intervention prompt information corresponding to the risk node in the safety behavior score curve that exceeds a preset threshold, and sending the driving intervention prompt information to the on-board terminal of the target vehicle for real-time display.

[0006] In a second aspect, an embodiment of the present invention provides a driving safety analysis device, comprising: a memory storing a computer program; and a processor for loading the computer program to implement the driving safety analysis method based on speech emotion recognition as described above.

[0007] The driving safety analysis method based on speech emotion recognition, provided by the present invention, extracts a multi-level emotional feature set consisting of real-time emotion fluctuation trajectories and semantic emotion intensity parameters from continuous speech segments and vehicle control interaction commands collected from a target vehicle's driving environment. Based on a safety risk prediction model, it generates a safety behavior scoring curve and generates driving intervention prompts based on risk nodes. This method dynamically associates the driver's emotional fluctuation characteristics with vehicle control behavior. By leveraging the emotional continuity patterns implicit in the speech data and the real-time intent analysis of semantic commands, it enables early prediction and precise intervention of driving safety risks, thus overcoming the shortcomings of traditional systems such as high false alarm rates and difficulty in tracing risks. By cross-validating multi-level features based on pitch change patterns and speech rate fluctuation frequency in continuous speech segments with semantic emotion intensity in interaction commands, the method effectively distinguishes between normal driver emotional expressions and dangerous emotional tendencies. Combining a dual-dimensional detection mechanism based on preset thresholds and risk nodes in the safety behavior scoring curve, the method can identify vehicle control deviation trends caused by anger, anxiety, or fatigue at the early stages of emotional fluctuations. Furthermore, the method dynamically associates the current driving environment parameters to generate differentiated intervention prompts, significantly improving the relevance of warning messages and driver response efficiency. In addition, based on the real-time processing and closed-loop feedback mechanism of voice data streams, emotional risk analysis can be achieved with only on-board voice acquisition equipment. While reducing hardware deployment costs, through the generation of standardized voice data sequences and multi-level feature fusion optimization, the interference of voice quality fluctuations on analysis results in complex driving environments is effectively suppressed, ultimately achieving the effect of improving the accuracy of driving safety warnings and reducing the incidence of traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0009] Figure 1 This is a flowchart of a driving safety analysis method based on speech emotion recognition provided by an embodiment of the present invention.

[0010] Figure 2 Schematic diagram of the composition of a driving safety analysis device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0012] See also Figure 1 , Figure 1 A flowchart of a driving safety analysis method based on speech emotion recognition provided by an embodiment of the present invention. The driving safety analysis method based on speech emotion recognition can be performed by a driving safety analysis device. The driving safety analysis method based on speech emotion recognition may include the following steps:

[0013] Step S100: Acquire a driving voice data stream collected by a target vehicle in a driving environment. The driving voice data stream includes continuous voice segments of the driver within a preset time period and interactive instructions related to vehicle control.

[0014] Driving voice data streams are collected while the target vehicle is in motion. These streams contain continuous driver speech clips within a preset time period. These clips may include daily conversations, emotional expressions, and interactive commands related to vehicle operation, such as "turn on the air conditioning" and "close the windows." The preset time period is a predefined period of time, defined based on actual needs, that limits the timeframe for collecting driving voice data.

[0015] In practice, the target vehicle can collect driving voice data streams using microphones installed inside the vehicle. For example, in a smart car, the microphone array inside the vehicle can collect the driver's voice signals from all directions and convert them into digital signals for storage and subsequent processing.

[0016] As an embodiment, in step S100, obtaining the driving voice data stream collected by the target vehicle in the driving environment can be implemented as follows:

[0017] Step S110: Segment the driving voice data stream to obtain multiple voice data units of equal length, and perform environmental noise filtering on each voice data unit to generate multiple denoised voice segments; wherein the segmentation processing includes endpoint detection based on voice energy mutation and context division based on semantic coherence; environmental noise filtering includes eliminating vehicle engine noise through frequency domain filtering, and separating the driver's voice from the passenger's voice through voiceprint feature matching.

[0018] Segmentation is the process of dividing the collected driving voice data stream into multiple voice data units of equal duration according to preset rules. Endpoint detection based on voice energy mutations determines the start and end points of speech by detecting the locations of sudden energy changes in the voice signal. For example, when the driver begins to speak, the energy of the voice signal suddenly increases. By detecting this energy mutation, the start and end points of the speech can be found. Context segmentation based on semantic coherence divides the semantically coherent parts into a voice data unit based on the semantic logical relationship of the speech content. For example, if a driver repeatedly says "turn on the air conditioner and adjust the temperature to 25 degrees," these two sentences are semantically coherent and can be divided into a single voice data unit.

[0019] Ambient noise filtering is used to remove interfering noise from the driver's voice data stream, improving voice data quality. Frequency domain filtering is a signal processing method that converts voice signals into the frequency domain. It then removes noise within a certain frequency range, such as that generated by the vehicle engine, based on the characteristics of different frequency components. Voiceprint feature matching utilizes each person's unique voiceprint to distinguish between the driver's voice and the passenger's voice. Because each person's voiceprint is as unique as a fingerprint, the driver's voice can be isolated by comparing the collected voice signal with the pre-registered driver voiceprint.

[0020] For example, when a car is driving, the voice data stream collected by the in-car microphone contains the roar of the engine, the conversations of the passengers, and the driver's instructions. First, through endpoint detection based on speech energy mutations and context segmentation based on semantic coherence, the voice data stream is segmented into multiple voice data units of equal duration. Then, frequency domain filtering is used to remove engine noise. Finally, voiceprint feature matching is used to isolate the driver's voice, ultimately generating multiple denoised voice segments.

[0021] Step S120: splicing multiple denoised speech segments in chronological order to generate a standardized speech data sequence corresponding to the driving speech data stream.

[0022] The standardized speech data sequence is obtained by reassembling multiple denoised speech segments after segmentation and denoising, in the order in which they appear in the original driving speech data stream. This is done to obtain a complete, ordered, and noise-free speech data sequence for subsequent processing such as emotion feature extraction.

[0023] For example, after processing in step S110, three denoised speech segments are obtained: "Turn on," "Air conditioning," and "Temperature adjusted to 25 degrees." These segments are obtained in the order in which the driver spoke. In step S120, these three denoised speech segments are concatenated in chronological order to obtain a standardized speech data sequence: "Turn on the air conditioning and adjust the temperature to 25 degrees."

[0024] Step S200: Extract emotional features from the driving voice data stream to generate a multi-level emotional feature set, wherein the multi-level emotional feature set includes real-time emotional fluctuation trajectories corresponding to continuous voice segments and semantic emotional intensity parameters corresponding to interactive instructions.

[0025] Emotional feature extraction is the process of extracting characteristic information that reflects the driver's emotional state from driving speech data streams. A multi-level emotional feature set encompasses multiple emotional features. Real-time emotional fluctuation trajectories refer to how the driver's emotions change over time during a continuous speech segment. These trajectories can be derived by analyzing speech characteristics such as pitch, speech rate, and amplitude. Semantic emotional intensity parameters are associated with interactive commands and are used to measure the emotional intensity of these commands. For example, a command like "Turn on the air conditioner now!" is likely to contain stronger emotional intensity than "Turn on the air conditioner."

[0026] For example, in a driving voice data stream, the driver initially calmly says "Open the window," then becomes agitated by the rising temperature inside the car and shouts, "Turn down the temperature!" By extracting emotional features, we can capture the driver's real-time emotional fluctuations throughout this speech, from calm to agitated, and also identify the higher semantic emotion intensity parameter corresponding to the interactive command "Turn down the temperature!"

[0027] As an embodiment, in step S200, the emotional feature extraction of the driving voice data stream to generate a multi-level emotional feature set can be implemented as follows:

[0028] Step S210: Input the standardized speech data sequence into the pre-trained emotion recognition network, and extract the pitch change pattern, speech rate fluctuation frequency, and speech amplitude distribution parameters of each denoised speech segment through the acoustic feature extraction layer in the emotion recognition network.

[0029] The pre-trained emotion recognition network is a neural network model trained with a large amount of speech data, which can be used to recognize emotion information in speech. The acoustic feature extraction layer is a component of the emotion recognition network, and its main function is to extract emotion-related acoustic features from the input standardized speech data sequence. The pitch change pattern refers to the change of pitch at different time points in speech. For example, the increase or decrease of pitch may reflect the driver's emotional changes. The pitch may increase when excited and remain relatively stable when calm. The speech rate fluctuation frequency refers to the change frequency of the speech rate. An increase in the speech rate may indicate that the driver is more nervous or excited, while a decrease in the speech rate may indicate that the driver is more relaxed or thinking. The speech amplitude distribution parameter refers to the distribution of the amplitudes of the speech signal. A larger amplitude may indicate that the driver's emotion is stronger, while a smaller amplitude may indicate a calmer emotion.

[0030] For example, when the standardized speech data sequence "Turn on the air conditioner and set the temperature to 25 degrees" is input into the pre-trained emotion recognition network, the acoustic feature extraction layer can analyze that the pitch is relatively stable, the speech rate is moderate, and the amplitude is normal when saying "Turn on"; while when saying "Set the temperature to 25 degrees", the pitch may slightly increase, the speech rate slightly accelerates, and the amplitude also slightly increases, thus obtaining the corresponding pitch change pattern, speech rate fluctuation frequency, and speech amplitude distribution parameter.

[0031] Step S220: Identify the command keywords, emotional modifiers, and modal particles included in the denoised speech segment through the semantic parsing layer in the emotion recognition network, and generate a semantic emotion label corresponding to each denoised speech segment.

[0032] The semantic parsing layer is the part in the emotion recognition network for processing speech semantic information. Command keywords refer to the keyword vocabulary related to vehicle control or other operations in the denoised speech segment, such as "Turn on", "Turn off", "Increase", "Decrease", etc. Emotional modifiers are words used to modify emotions, such as "Very", "Especially", "A little", etc., which can enhance or weaken the degree of emotion expression. Modal particles are words that play a role in expressing the mood at the end of a sentence, such as "Ah", "Ya", "Ne", etc. Different modal particles can convey different emotions. For example, "Ya" in "Hurry up and turn it on!" can better reflect the urgent emotion than "Hurry up and turn it on". The semantic emotion label is a label representing the emotion type assigned to the denoised speech segment according to its semantic information, such as "Calm", "Angry", "Anxious", etc.

[0033] For example, for the denoised speech segment "Immediately open the window, ah!", the semantic parsing layer can identify the command keyword "Open", the emotional modifier "Immediately", and the modal particle "Ah", and generate the corresponding semantic emotion label as "Urgent" based on this information.

[0034] Step S230: Analyze the emotion transition features between adjacent denoised speech segments through the context association layer in the emotion recognition network, and generate a real-time emotion fluctuation trajectory based on the emotion transition features.

[0035] The contextual association layer is the part of the emotion recognition network that analyzes speech context. Emotional transition features between adjacent denoised speech segments indicate the change in emotion from one denoised segment to the next, such as the transition from calm to anger. Real-time emotion fluctuation trajectories are generated based on these emotional transition features, visually demonstrating the dynamic changes in a driver's emotions over time.

[0036] For example, consider two adjacent denoised speech clips: the first, "The weather is nice today," conveys a calm and pleasant emotion, while the second, "What's that car driving in front of me?", expresses anger. The contextual association layer can analyze the emotional transition between the two clips, identifying a sudden shift from calm to anger. Based on this transition, it generates a real-time emotion fluctuation trajectory, clearly showing the sharp rise in emotion on the trajectory graph.

[0037] Step S240: performing feature fusion on the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion tags, and real-time emotion fluctuation trajectory to generate a multi-level emotion feature set.

[0038] Feature fusion is the process of integrating different types of emotional features to create a more comprehensive and accurate emotional representation. By fusing features such as pitch variation patterns, speech rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion labels, and real-time emotion fluctuation trajectories, a multi-level emotional feature set is generated. This set contains information reflecting the driver's emotional state from multiple perspectives, both acoustically and semantically.

[0039] For example, in the previous example, we obtained the pitch variation pattern, speaking rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion labels, and real-time emotion fluctuation trajectory for each denoised speech segment. Through feature fusion, this information is integrated to form a multi-level emotion feature set that more comprehensively reflects the driver's emotional state throughout the speech process.

[0040] As an implementation method, step S240, performing feature fusion on the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion tags, and real-time emotion fluctuation trajectory to generate a multi-level emotion feature set, can be implemented as follows:

[0041] Step S241: perform time axis normalization on the pitch change pattern to generate pitch change pattern parameters consistent with the time granularity of the standardized speech data sequence, and extract the pitch mutation time nodes corresponding to the abnormal pitch segments exceeding the preset pitch threshold in the pitch change pattern parameters.

[0042] Timeline normalization unifies pitch variation patterns along the timeline, aligning them with the time granularity of the standardized speech data sequence. Time granularity refers to the precision with which data is divided in time, for example, in seconds. The preset pitch threshold is a pre-set pitch limit; pitch segments exceeding this threshold are considered abnormal. The pitch mutation time point is the point at which the abnormal pitch segment begins to appear.

[0043] For example, a standardized speech data sequence is divided into time units per second, while the original pitch change pattern may be recorded at a finer or coarser time granularity. Through time axis normalization, the pitch change pattern is adjusted to units per second, resulting in pitch change pattern parameters that align with the time granularity of the standardized speech data sequence. Assuming the preset pitch threshold is 80 decibels, and a pitch segment is found in the pitch change pattern parameters to reach 90 decibels, the start time of this abnormal pitch segment exceeding 80 decibels is the pitch mutation time node.

[0044] Step S242: Perform frequency domain segmented sampling on the speech rate fluctuation frequency to generate speech rate fluctuation frequency distribution parameters based on a fixed time window, and perform local time alignment on the speech rate fluctuation frequency distribution parameters based on the pitch mutation time node to generate speech rate fluctuation associated parameters synchronized with the pitch change pattern parameters.

[0045] Frequency domain segmented sampling is the process of segmenting the speech rate fluctuation frequency in the frequency domain and sampling each frequency segment. A fixed time window refers to a time period of fixed length set in time for analyzing the speech rate fluctuation frequency. Through frequency domain segmented sampling, the speech rate fluctuation frequency distribution parameters based on the fixed time window can be obtained, which reflects the distribution of speech rate fluctuation frequency in different time periods. Local time alignment refers to the temporal adjustment of the speech rate fluctuation frequency distribution parameters around the time node of the pitch mutation, so that they are synchronized with the pitch change pattern parameters.

[0046] For example, we set a fixed time window of 10 seconds and perform frequency-domain segmented sampling of the speaking rate fluctuation frequency to obtain the speaking rate fluctuation frequency distribution parameters within each 10-second time period. Assuming a pitch mutation time node occurs at the 20th second, we perform local time alignment on the speaking rate fluctuation frequency distribution parameters near this time node to synchronize them with the pitch change pattern parameters, generating the speaking rate fluctuation correlation parameters.

[0047] Step S243: Dynamic range compression is performed on the speech amplitude distribution parameters to generate compressed speech amplitude smoothing parameters, and the speech amplitude smoothing parameters are superimposed on the speech rate fluctuation correlation parameters in the time domain to generate superimposed speech energy comprehensive parameters.

[0048] Dynamic range compression is a method for processing speech amplitude distribution parameters. It reduces the dynamic range of speech amplitude, making excessively large or small amplitude values ​​more balanced, thereby generating compressed speech amplitude smoothing parameters. Time domain superposition is the process of superimposing speech amplitude smoothing parameters and speech rate fluctuation-related parameters in the time domain. This superposition comprehensively considers the impact of speech amplitude and speech rate fluctuations on speech energy, generating a superimposed speech energy composite parameter.

[0049] For example, the original speech amplitude distribution parameters may have very large amplitudes in some time periods and very small amplitudes in others. Dynamic range compression adjusts these amplitudes to produce smoother speech amplitude smoothing parameters. These smoothing parameters are then temporally superimposed with the previously generated speech rate fluctuation correlation parameters to produce a composite speech energy parameter, which more comprehensively reflects the energy characteristics of the speech.

[0050] Step S244: converting the semantic emotion label into an emotion category coding vector, and performing context expansion on the emotion category coding vector based on the emotion change direction of adjacent time nodes in the real-time emotion fluctuation trajectory to generate an expanded emotion category coding sequence containing historical emotion dependencies.

[0051] The emotion category encoding vector converts semantic emotion labels into a vector representation, with each vector element corresponding to an emotion category. This vector representation facilitates subsequent calculations and processing. Contextual expansion further expands the emotion category encoding vector based on the direction of emotion change at adjacent time nodes in the real-time emotion fluctuation trajectory, in order to include historical emotion dependencies.

[0052] For example, semantic emotion labels such as "calm," "angry," and "anxiety" are converted into emotion category encoding vectors, such as the vector [1, 0, 0] for "calm," the vector [0, 1, 0] for "angry," and the vector [0, 0, 1] for "anxiety." If, based on the real-time emotion fluctuation trajectory, a transition from "calm" to "angry" is detected, the historical information of this emotion change is taken into account when generating the extended emotion category encoding sequence, allowing the encoding sequence to better reflect the dynamic changes in emotion.

[0053] Step S245: divide the extended emotion category coding sequence into time slices according to the pitch mutation time nodes to generate multiple emotion coding subsequences, and perform feature splicing on each emotion coding subsequence and the speech energy comprehensive parameter of the corresponding time interval to generate a preliminary fusion feature segment.

[0054] Time slicing divides the extended emotion category coding sequence into multiple subsequences of emotion coding according to the time nodes of pitch mutations. Feature concatenation combines each emotion coding subsequence with the comprehensive speech energy parameters of the corresponding time interval. This concatenation allows for a preliminary fusion of acoustic and semantic features, generating a preliminary fused feature segment.

[0055] For example, based on the previously determined pitch mutation time nodes, the extended emotion category coding sequence is divided into three emotion coding subsequences, each corresponding to a different time period. Each emotion coding subsequence is then concatenated with the speech energy comprehensive parameter within that time period to generate three preliminary fused feature segments.

[0056] Step S246: performing redundant feature filtering on the preliminary fused feature segments, retaining feature dimensions that are strongly correlated with emotional turning points in the real-time emotional fluctuation trajectory, and generating an optimized fused feature subset.

[0057] Redundant feature filtering is the process of removing features that are repeated or have little contribution to emotional expression from the initial fused feature fragment. Emotional turning points are the points in the real-time emotional fluctuation trajectory where significant emotional changes occur. By retaining feature dimensions that are strongly correlated with emotional turning points, the optimized fused feature subset can more accurately reflect the driver's emotional changes.

[0058] For example, the initial fusion feature fragment may contain some features that do not have much impact on emotional expression. These features are removed through redundant feature filtering, and only feature dimensions closely related to emotional turning points are retained, such as the feature dimensions related to the sudden increase in pitch and the sudden increase in speaking speed at the turning point from calm to angry emotions, thereby generating an optimized fusion feature subset.

[0059] Step S247: The optimized fusion feature subset is associated across segments in chronological order to generate a continuous emotion feature sequence covering a preset time period, and the integrity of the continuous emotion feature sequence is verified to ensure a balanced distribution of feature contributions of pitch change patterns, speaking rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion labels, and real-time emotion fluctuation trajectories.

[0060] Cross-segment association is the process of connecting and associating optimized fused feature subsets in chronological order, allowing features from different time periods to form a continuous sequence. Integrity verification checks the generated continuous emotion feature sequence to ensure that the contribution of each feature is evenly distributed, preventing some features from having too large or too small an impact.

[0061] For example, the previously generated optimized fusion feature subsets are correlated across segments in chronological order to generate a continuous sequence of emotional features covering a preset time period. This sequence is then verified for integrity, examining the balanced contribution of features such as pitch variation patterns, speech rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion labels, and real-time emotion fluctuation trajectories within the sequence. If unbalanced, appropriate adjustments are made.

[0062] Step S248: Compare the verified continuous emotion feature sequence with the driver's personalized emotion baseline model, adjust the weights of abnormal emotion features in the continuous emotion feature sequence that deviate from the baseline range, and generate a multi-level emotion feature set including standardized emotion intensity and personalized correction weights.

[0063] The personalized emotion baseline model is a pre-established benchmark model based on the driver's individual voice and emotion characteristics. It reflects the driver's emotional state under normal circumstances. By comparing a verified continuous emotion feature sequence with the personalized emotion baseline model, abnormal emotion features that deviate from the baseline range can be identified. Adjusting the weights of abnormal emotion features refers to adjusting the weights of these abnormal features in the multi-level emotion feature set, so that the multi-level emotion feature set reflects standardized emotion intensity while taking into account the driver's individual characteristics.

[0064] For example, a driver's personalized emotional baseline model shows that their voice is normally steady and their speech rate is moderate. However, within a continuous emotional feature sequence, there are periods of time when their voice pitch suddenly rises and their speech rate increases, deviating from the baseline range. By adjusting the weights of these abnormal emotional features, the multi-level emotional feature set more accurately reflects the driver's emotional state while also taking into account their personalized emotional characteristics.

[0065] As an embodiment, the emotion recognition network is obtained by jointly training the acoustic feature extraction layer, the semantic parsing layer, and the context association layer. The joint training process includes:

[0066] Step S201: Obtain a labeled training speech dataset, where each speech sample in the training speech dataset is labeled with a tone feature label, a semantic emotion label, and a context emotion continuity label.

[0067] The labeled training speech dataset is a set of speech data used to train emotion recognition networks. Each speech sample is annotated with emotion-related labels. The pitch feature label describes the tonal characteristics of a speech sample, such as pitch and pitch variation. The semantic emotion label assigns emotion labels based on the semantic content of the speech sample, such as "happy," "sad," and "angry." The contextual emotion continuity label describes the emotional continuity between adjacent speech samples and reflects the transition of emotions between contexts.

[0068] For example, in a training speech dataset, there is a speech sample "I'm so happy today!" Its pitch feature label may be labeled as "high-pitched and stable", and its semantic emotion label may be labeled as "happy". If the speech before this speech sample also expresses positive emotions, then the context emotion continuity label can be labeled as "emotion continuously positive".

[0069] Step S202: Input the speech sample into the initial emotion recognition network, generate pitch prediction features through the acoustic feature extraction layer, generate semantic emotion prediction labels through the semantic analysis layer, and generate emotion continuity prediction values ​​through the context association layer.

[0070] The initial emotion recognition network is a neural network model in its initial state before training begins. When a speech sample is input into the initial emotion recognition network, the acoustic feature extraction layer analyzes the acoustic features of the speech sample to generate pitch prediction features, such as the pitch variation pattern and pitch height of the speech. The semantic parsing layer processes the semantic content of the speech sample to generate a semantic emotion prediction label, that is, to predict the type of emotion expressed by the speech sample. The context association layer considers the context of the speech sample to generate an emotion continuity prediction value, which is used to predict the emotional continuity between adjacent speech samples.

[0071] For example, when the speech sample "That's terrible!" is input into the initial emotion recognition network, the acoustic feature extraction layer may generate a pitch prediction feature of "low pitch and a downward trend", the semantic analysis layer generates a semantic emotion prediction label of "anger", and the context association layer generates an emotion continuity prediction value based on context information, indicating that the emotion of the speech sample is continuously negative compared to the emotion of the previous sample.

[0072] Step S203: Calculate the first loss function between the tone prediction feature and the tone feature label, calculate the second loss function between the semantic emotion prediction label and the semantic emotion label, and calculate the third loss function between the emotion continuity prediction value and the context emotion continuity label.

[0073] Loss functions measure the difference between the model's predictions and the true labels. The first loss function measures the difference between the pitch prediction features and the pitch feature labels, reflecting the prediction accuracy of the acoustic feature extraction layer. The second loss function measures the difference between the semantic emotion prediction labels and the semantic emotion labels, reflecting the prediction accuracy of the semantic parsing layer. The third loss function measures the difference between the emotion continuity prediction values ​​and the contextual emotion continuity labels, reflecting the prediction accuracy of the contextual association layer.

[0074] For example, if the pitch prediction feature is "high and stable" while the pitch feature label is "low and rising," the first loss function can be used to determine the difference between the two. Similarly, if the semantic emotion prediction label is "happy" while the semantic emotion label is "angry," the second loss function can be used to measure the prediction error at the semantic parsing layer. Similar calculations are performed for emotion continuity prediction values ​​and contextual emotion continuity labels.

[0075] Step S204: performing weighted summation on the first loss function, the second loss function, and the third loss function to generate a joint training loss value, and updating the parameters of the initial emotion recognition network through the back propagation algorithm until convergence.

[0076] Weighted summation involves assigning different weights to the first, second, and third loss functions, and then summing them to obtain the joint training loss. The weights can be determined based on the importance of different layers in emotion recognition. The backpropagation algorithm is used to update neural network parameters. It adjusts the parameters of each layer in the initial emotion recognition network based on the joint training loss, gradually bringing the model's predictions closer to the true labels until convergence is reached, meaning the joint training loss no longer decreases significantly.

[0077] For example, assuming the weight of the first loss function is 0.3, the weight of the second loss function is 0.5, and the weight of the third loss function is 0.2, the three loss functions are weighted and summed according to these weights to obtain the joint training loss value. The parameters of the initial emotion recognition network are then continuously updated through the backpropagation algorithm. After multiple iterations, when the joint training loss value no longer changes significantly, the model is considered converged.

[0078] As an implementation method, the above-mentioned joint training may further include:

[0079] Step S205: adding interfering speech samples containing vehicle environmental noise to the training speech data set, and setting noise suppression weights corresponding to the interfering speech samples.

[0080] Interfering speech samples are those containing vehicle ambient noise, which can interfere with the training of the emotion recognition network. Adding interfering speech samples to the training speech dataset allows the model to learn how to accurately recognize emotions in the presence of noise. The noise suppression weight is a value set for interfering speech samples. It controls the degree of processing applied to interfering speech samples during training. A larger weight indicates a greater emphasis on noise suppression.

[0081] For example, the training speech dataset originally contained pure speech samples. Now, some speech samples containing interference such as vehicle engine noise and wind noise are added, and the noise suppression weight for these interfering speech samples is set to 0.6, indicating that the training process focuses on processing these noises.

[0082] Step S206: When extracting features from the interference speech sample through the acoustic feature extraction layer, the contribution of the frequency band features corresponding to the vehicle engine noise in the interference speech sample is reduced based on the noise suppression weight.

[0083] When the acoustic feature extraction layer extracts features from interfering speech samples, it processes the frequency band features corresponding to vehicle engine noise based on the noise suppression weights. Reducing the contribution of these frequency band features can reduce the interference of engine noise on model training, allowing the model to focus more on the emotional characteristics of the speech.

[0084] For example, vehicle engine noise is mainly concentrated in a certain frequency band. When extracting features from interfering speech samples, the contribution of the frequency band features in the entire feature extraction process is reduced according to the set noise suppression weight, so that the model will not produce incorrect emotion recognition results due to the influence of engine noise.

[0085] Step S207: When parsing the interfering speech sample through the semantic parsing layer, the keyword retrieval strength related to the driver's speech is increased based on the noise suppression weight to cover the semantic ambiguity area caused by the noise.

[0086] When parsing interfering speech samples at the semantic parsing layer, the presence of noise may cause ambiguity in the speech semantics. By increasing the search intensity of keywords related to the driver's speech based on noise suppression weights, the semantic information in the speech can be more accurately identified, addressing the semantic ambiguity caused by noise and improving the accuracy of semantic parsing.

[0087] For example, in the presence of noise, certain key words in the speech may not be clear, resulting in semantic ambiguity. By increasing the keyword search intensity, the semantic parsing layer can more actively search for keywords related to the driver's speech, and more accurately understand the semantic content of the speech even in the presence of noise.

[0088] Step S300: Input the multi-level emotional feature set into the safety risk prediction model to generate a safety behavior score curve for the driver within a preset time period. The safety behavior score curve is used to characterize the correlation between the driver's emotional stability and driving operation risk at different time nodes.

[0089] The safety risk prediction model is used to predict a driver's driving safety risk. It analyzes the driver's emotional state based on a multi-level set of emotional features and predicts their driving risk at different time points. The safety behavior score curve is generated based on the model's predictions. It reflects the correlation between the driver's emotional stability and driving risk at different time points within a preset time period. The height of the curve indicates the magnitude of driving risk. The more unstable the emotion, the higher the curve is likely to be, and the corresponding driving risk is greater.

[0090] For example, the previously generated multi-level emotional feature set is fed into the safety risk prediction model. Based on these features, the model analyzes the driver's emotional changes over a preset time period and generates a safety behavior score curve. If the driver is emotionally agitated and has poor emotional stability at a certain time point, the corresponding score on the safety behavior score curve may be higher, indicating a higher driving risk.

[0091] As an embodiment, step S300, inputting the multi-level emotional feature set into the safety risk prediction model to generate a driver's safety behavior score curve within a preset time period, can be implemented as follows:

[0092] Step S310: Using the time series analysis module in the safety risk prediction model, the real-time emotion fluctuation trajectory is periodically decomposed to extract the driver's emotion mutation frequency and mutation amplitude within a preset time window.

[0093] The time series analysis module is the component of the safety risk prediction model that analyzes time series data. Real-time emotion fluctuation trajectories are sequence data that reflect changes in a driver's emotions over time. Periodic decomposition breaks down these real-time emotion fluctuation trajectories into distinct periodic components to better analyze patterns in emotion. Emotional mutation frequency refers to the number of sudden changes in a driver's emotions within a preset time window, while emotional mutation amplitude refers to the change in emotional intensity with each sudden change.

[0094] For example, within a preset time window of 10 minutes, the real-time emotional fluctuation trajectory was periodically decomposed through the time series analysis module, and it was found that the driver's emotions underwent three obvious mutations within these 10 minutes. The change in emotional intensity during each mutation was 20%, 30%, and 25%, respectively. Therefore, the frequency of emotional mutation was 3 times / 10 minutes, and the mutation amplitudes were 20%, 30%, and 25%, respectively.

[0095] Step S320: Using the risk association mapping module in the safety risk prediction model, the semantic emotion intensity parameter is matched with the predefined driving operation risk level to determine the vehicle control deviation probability corresponding to different emotion intensities.

[0096] The risk association mapping module is the component of the safety risk prediction model that establishes a correlation between semantic emotion intensity parameters and driving operation risk levels. Predefined driving operation risk levels, such as low risk, medium risk, and high risk, are pre-defined based on extensive experiments and data analysis. The vehicle control deviation probability refers to the likelihood of a driver deviating from vehicle control under varying levels of emotion intensity.

[0097] For example, the predefined driving operation risk levels are divided into low, medium, and high risk levels, and the semantic emotion intensity parameters are weak, medium, and strong. The risk association mapping module matches the semantic emotion intensity parameters with the driving operation risk levels and finds that when the semantic emotion intensity is weak, the corresponding probability of vehicle control deviation is 10%, which is a low risk level; when the semantic emotion intensity is medium, the corresponding probability of vehicle control deviation is 30%, which is a medium risk level; and when the semantic emotion intensity is strong, the corresponding probability of vehicle control deviation is 60%, which is a high risk level.

[0098] As an embodiment, step S320, matching the semantic emotion intensity parameter with the predefined driving operation risk level through the risk association mapping module in the safety risk prediction model to determine the vehicle control deviation probability corresponding to different emotion intensities, can be implemented as follows:

[0099] Step S321: Based on the emotion type label and intensity quantization value included in the semantic emotion intensity parameter, the driver's emotion expression is divided into emotion intensity intervals to generate an emotion intensity interval division result including multiple continuous emotion intensity levels.

[0100] The emotion type label describes the driver's emotion type, such as anger, anxiety, and calmness. The intensity quantization value is a specific numerical representation of the intensity of the emotion, for example, a value from 0 to 100 can be used to represent the intensity of the emotion. The emotion intensity interval classification divides the driver's emotional expressions into different intensity intervals based on the emotion type label and intensity quantization value, forming multiple continuous emotion intensity levels.

[0101] For example, the semantic emotion intensity parameter contains the emotion type label "anger" and an intensity quantization value of 30-100. Anger can be divided into three intensity levels: mild anger (intensity quantization value 30-50), moderate anger (intensity quantization value 51-70), and severe anger (intensity quantization value 71-100), thus obtaining the emotion intensity interval classification result.

[0102] Step S322: Based on the emotion intensity threshold range corresponding to each risk level in the predefined driving operation risk levels, each emotion intensity level in the emotion intensity interval division result is matched with the risk level layer by layer to generate an emotion intensity-risk level mapping relationship, where the risk level includes the probability of emergency braking false triggering due to anger, the frequency of lane departure due to anxiety, and the duration of throttle response delay due to fatigue.

[0103] Each predefined driving risk level corresponds to an emotional intensity threshold range. For example, a low risk level might correspond to an emotional intensity threshold range of 0-30, a medium risk level might correspond to an emotional intensity threshold range of 31-60, and a high risk level might correspond to an emotional intensity threshold range of 61-100. By matching each emotional intensity level in the emotional intensity interval classification results with these risk levels layer by layer, a mapping relationship between emotional intensity and risk level can be established.

[0104] For example, considering the three levels of anger intensity described above, mild anger (intensity values ​​of 30-50) might correspond to a low risk level, moderate anger (intensity values ​​of 51-70) might correspond to a medium risk level, and severe anger (intensity values ​​of 71-100) might correspond to a high risk level. Furthermore, different risk levels correspond to specific risk indicators, such as the probability of emergency braking misactivation due to anger, the frequency of lane departures due to anxiety, and the duration of throttle response delay due to fatigue.

[0105] Step S323: Based on the emotion intensity-risk level mapping relationship, the occurrence frequency of vehicle control deviation events in the historical driving operation data set corresponding to each emotion intensity level is extracted, and an initial vehicle control deviation probability is generated according to the occurrence frequency.

[0106] The historical driving operation dataset contains a large amount of historical driving operation data from drivers, recording the occurrence of vehicle control deviation events under different emotional states. Based on the mapping between emotion intensity and risk level, the frequency of vehicle control deviation events corresponding to each emotion intensity level is extracted from the historical driving operation dataset. For example, the frequency of vehicle control deviation events under mild anger is 10%, under moderate anger, 30%, and under severe anger, 60%. An initial vehicle control deviation probability is then generated based on these occurrence frequencies.

[0107] For example, for mild anger, based on the corresponding vehicle control deviation event frequency of 10%, the initial vehicle control deviation probability is generated as 10%; for moderate anger, based on the occurrence frequency of 30%, the initial vehicle control deviation probability is generated as 30%; for severe anger, based on the occurrence frequency of 60%, the initial vehicle control deviation probability is generated as 60%.

[0108] Step S324: Dynamically correct the initial vehicle control deviation probability. The dynamic environmental correction includes adjusting the weight distribution of the initial vehicle control deviation probability in different environments based on the road complexity parameters and traffic flow density parameters in the current driving environment to generate a vehicle control deviation probability after environmental adaptive correction.

[0109] Dynamic environmental correction is the process of adjusting the initial vehicle control deviation probability to account for the impact of the current driving environment on the probability of vehicle control deviation. The road complexity parameter describes the complexity of road conditions, such as the number of curves and slope variations. The traffic density parameter refers to the number of vehicles passing through a certain road section per unit time. By combining these two parameters, the weight distribution of the initial vehicle control deviation probability is adjusted under different environments, making the vehicle control deviation probability more consistent with the actual driving environment.

[0110] For example, in an environment with high road complexity and high traffic density, even if the driver's emotional intensity is in a state of mild anger, the initial vehicle control deviation probability is 10%, but due to the influence of the environment, this probability may need to be adjusted to 20% to generate a vehicle control deviation probability after environmental adaptive correction.

[0111] Step S325: The vehicle control deviation probability after environmental adaptive correction is correlated with the vehicle control status data collected in real time for verification. The correlation verification includes detecting whether the number of sudden changes in the steering wheel angle and the abnormal fluctuations in the brake pedal pressure are consistent with the predicted trend of the vehicle control deviation probability, and calibrating the confidence of the vehicle control deviation probability based on the verification results.

[0112] Correlation verification is the process of comparing and verifying the vehicle control deviation probability, corrected through environmental adaptive testing, with real-time vehicle handling status data. The number of sudden steering angle changes and abnormal fluctuations in brake pedal pressure are important indicators of vehicle handling abnormalities. By testing whether these indicators align with the predicted trend of the vehicle control deviation probability, the accuracy of the vehicle control deviation probability can be assessed. Confidence calibration adjusts the vehicle control deviation probability based on the verification results to make it more reliable.

[0113] For example, the vehicle control deviation probability prediction after environmental adaptive correction is more likely to have vehicle control deviation within a certain time period, while the real-time collected vehicle control status data shows that the number of sudden changes in the steering wheel steering angle increases and the brake pedal depression force fluctuates abnormally, indicating that the predicted trend is consistent with the actual situation, and the confidence level of the vehicle control deviation probability is high, so no major adjustments are required; if the predicted trend is inconsistent with the actual situation, the vehicle control deviation probability needs to be calibrated accordingly.

[0114] Step S326: Based on the confidence-calibrated vehicle control deviation probability, generate the final vehicle control deviation probability corresponding to each emotion intensity level, and associate the final vehicle control deviation probability with each semantic emotion intensity parameter within the preset time period in chronological order to form a probability input sequence for driving the generation of the safety behavior scoring curve.

[0115] The final vehicle control deviation probability is a confidence-calibrated probability that more accurately reflects the likelihood of vehicle control deviation at different levels of emotion intensity. The final vehicle control deviation probability is chronologically associated with each semantic emotion intensity parameter within a preset time period to form a probability input sequence. This sequence serves as input data for generating a safety behavior scoring curve for the safety risk prediction model.

[0116] For example, after confidence calibration, the probability of a vehicle control deviation corresponding to mild anger is 15%, moderate anger is 35%, and severe anger is 65%. These probabilities are chronologically associated with the semantic emotion intensity parameters within a preset time period to form a probability input sequence, which is used to drive the generation of the safety behavior scoring curve.

[0117] Step S330: The dynamic score generation module in the safety risk prediction model is used to calculate the driver's safety score values ​​at multiple consecutive time points, combining the emotion mutation frequency, mutation amplitude, and vehicle control deviation probability.

[0118] The dynamic score generation module is the component of the safety risk prediction model used to calculate the driver's safety score. The frequency and magnitude of emotional fluctuations, as well as the probability of vehicle control deviation, are key factors influencing the driver's safety score. By comprehensively considering these factors, a driver's safety score can be calculated at multiple consecutive time points. These scores reflect the driver's driving safety at different points in time.

[0119] For example, at a certain point in time, the driver's mood changes suddenly twice per minute, the amplitude of the change is 25%, and the probability of vehicle control deviation is 30%. The dynamic score generation module uses a preset algorithm to combine these factors and calculates a safety score of 70 points at that point in time.

[0120] As an embodiment, step S330, by using the dynamic score generation module in the safety risk prediction model, combined with the emotion mutation frequency, mutation amplitude, and vehicle control deviation probability, calculates the driver's safety score value at multiple consecutive time points, which can be implemented as follows:

[0121] Step S331: determining the time interval of the driver's emotion fluctuations between adjacent time nodes based on the emotion mutation frequency, and generating a first influencing factor negatively correlated with driving concentration based on the emotion fluctuation time interval.

[0122] The emotional fluctuation interval refers to the length of time between driver emotions. The higher the frequency of emotional changes, the shorter the emotional fluctuation interval. The first impact factor is negatively correlated with driving concentration. The shorter the emotional fluctuation interval, the larger the first impact factor, indicating lower driving concentration.

[0123] For example, if the frequency of emotional changes is 3 times per minute, then the interval between emotional fluctuations at adjacent time nodes is approximately 20 seconds. According to the preset rules, when the interval between emotional fluctuations is 20 seconds, the first impact factor generated is 0.8, indicating that the driver's driving concentration is low.

[0124] Step S332: extracting the driver's emotion intensity difference value within a preset time window according to the mutation amplitude, and comparing the emotion intensity difference value with a predefined emotion stability threshold to generate a second influencing factor reflecting the amplitude of emotion fluctuation.

[0125] The emotional intensity difference value refers to the change in emotional intensity when a driver's emotions suddenly change within a preset time window. The predefined emotional stability threshold is a pre-set limit for emotional intensity difference, used to determine whether the driver's emotions are stable. By comparing the emotional intensity difference value with the emotional stability threshold, a second impact factor reflecting the magnitude of emotional fluctuations can be generated. A larger emotional intensity difference value indicates a larger second impact factor, indicating greater emotional fluctuations.

[0126] For example, within a preset 5-minute window, the driver's emotional intensity fluctuates by 30%, and the predefined emotional stability threshold is 20%. Because the 30% emotional intensity difference exceeds the 20% threshold, the generated second impact factor is 0.9, indicating a large emotional fluctuation.

[0127] Step S333: Based on the vehicle control deviation probability, the number of vehicle control delay events caused by the driver's emotional fluctuations within the historical time node is identified, and a third influencing factor representing the cumulative degree of operating errors is generated according to the number.

[0128] Vehicle control delay events occur when drivers experience delays in vehicle control due to emotional fluctuations. The vehicle control deviation probability can be used to identify the number of vehicle control delay events caused by emotional fluctuations from historical driving data. A third impact factor is generated based on the number of these events to indicate the cumulative degree of operational errors. A higher number of events indicates a higher third impact factor value.

[0129] For example, based on the probability of vehicle control deviation, historical driving data identified five instances of vehicle control delays due to driver emotional fluctuations over the past 10 minutes. Based on pre-set rules, the resulting third impact factor was 0.7, indicating a high level of cumulative operational errors.

[0130] Step S334: Adjust the weight distribution of the first influencing factor, the second influencing factor, and the third influencing factor to obtain a dynamic weight distribution result, wherein the weight distribution adjustment is dynamically corrected based on the degree of deviation between the driver's emotional mutation pattern and the actual driving trajectory of the vehicle within a preset time window.

[0131] Weight adjustment assigns different weights to the first, second, and third influencing factors to comprehensively consider their impact on the safety score. Dynamic adjustment means that the weighting is not fixed but rather adjusted based on the degree of deviation between the driver's emotional fluctuation pattern and the vehicle's actual driving trajectory within a preset time window. If the degree of deviation between the driver's emotional fluctuation pattern and the vehicle's actual driving trajectory is large, indicating a significant impact of emotions on driving performance, the weights of the first and second influencing factors may need to be increased.

[0132] For example, within a preset 15-minute window, the driver's emotional abrupt change pattern deviated significantly from the vehicle's actual driving trajectory. Originally, the weight of the first influencing factor was 0.3, the weight of the second influencing factor was 0.3, and the weight of the third influencing factor was 0.4. After dynamic correction, the weight of the first influencing factor was adjusted to 0.4, the weight of the second influencing factor was adjusted to 0.4, and the weight of the third influencing factor was adjusted to 0.2, resulting in a dynamic weight allocation result.

[0133] Step S335: Perform weighted fusion on the first influencing factor, the second influencing factor, and the third influencing factor according to the dynamic weight allocation result to generate an initial safety score sequence, and superimpose a time decay factor on each score value in the initial safety score sequence to obtain a smoothed intermediate safety score sequence. The time decay factor is used to reduce the influence weight of the score value of the historical time node on the current time node.

[0134] Weighted fusion generates an initial safety score sequence by weighting the first, second, and third influencing factors according to the dynamic weight distribution. The time decay factor is a factor that gradually decreases over time. It is used to reduce the influence of the score values ​​at historical time points on the current time point, making the safety score sequence more focused on the current driving state.

[0135] For example, the first impact factor is 0.8, with a weight of 0.4; the second impact factor is 0.9, with a weight of 0.4; and the third impact factor is 0.7, with a weight of 0.2. After weighted fusion, the initial safety score is 0.8 × 0.4 + 0.9 × 0.4 + 0.7 × 0.2 = 0.82. Then, a time decay factor (assuming a time decay factor of 0.9) is added to this initial safety score, resulting in a smoothed intermediate safety score of 0.82 × 0.9 = 0.738.

[0136] Step S336: The intermediate safety score sequence is correlated and calibrated with the driver's real-time physiological indicator data to generate a calibrated safety score value, where the physiological indicator data includes heart rate change trends and hand grip stability parameters. The correlation calibration is used to correct the score deviation caused by voice data collection errors.

[0137] Real-time physiological indicator data reflects the driver's current physiological state, such as heart rate trends and grip stability parameters. Correlation calibration compares and adjusts the intermediate safety score sequence with these physiological indicator data to correct for score deviations caused by voice data collection errors. If the physiological indicator data indicates a driver's condition inconsistent with the intermediate safety score sequence, the safety score may need to be calibrated accordingly.

[0138] For example, the intermediate safety score sequence shows a driver's safety score of 70, but real-time physiological indicator data shows a significantly elevated heart rate and unstable grip, indicating that the driver's actual condition may be more dangerous than the score sequence indicates. Through correlation calibration, the safety score is adjusted to 60, generating a calibrated safety score.

[0139] Step S337: Map the calibrated safety score value to each time node within the preset time period in chronological order to generate a safety behavior score curve, and perform differential calculation on the score values ​​of adjacent time nodes in the safety behavior score curve to generate a score gradient parameter indicating the rate of change of driving risk to assist in the identification of risk nodes.

[0140] By mapping the calibrated safety score values ​​to each time point within a preset time period in chronological order, a safety behavior score curve can be generated. The score gradient parameter is calculated by taking the difference between the score values ​​at adjacent time points in the safety behavior score curve and can indicate the rate of change in driving risk. A larger score gradient parameter indicates a faster change in driving risk, and the corresponding time point may be a risky one.

[0141] For example, the calibrated safety score values ​​at time points within a preset time period are 70, 65, and 60, respectively. Through differential calculation, the score gradient parameters for adjacent time points are -5 points / time unit and -5 points / time unit, respectively. These score gradient parameters can help identify time points where driving risk changes rapidly, assisting in determining risk points.

[0142] Step S340: Connect the safety score values ​​in chronological order to generate a safety behavior score curve, and smooth the safety behavior score curve to eliminate short-term noise interference.

[0143] In this step, the safety score values ​​calculated at multiple consecutive time points are linked together in chronological order to form a safety behavior score curve. Smoothing is performed to remove short-term noise interference from the safety behavior score curve, making the curve smoother and more representative of the true correlation between driver emotional stability and driving risk.

[0144] For example, the safety behavior scoring curve can be processed through smoothing methods such as moving average to remove fluctuations in the scoring value caused by accidental factors, making the curve smoother and facilitating subsequent analysis and judgment.

[0145] In one embodiment, the safety risk prediction model is obtained by joint training with multi-dimensional driving data, where the multi-dimensional driving data includes historical driving voice data, vehicle operation log data, and actual accident record data. The joint training with multi-dimensional driving data includes:

[0146] Step S301: Input historical driving voice data into the emotion recognition network to generate a historical emotion feature set, and convert the steering wheel angle, braking frequency and lane deviation in the vehicle operation log data into a control risk index.

[0147] Historical driving voice data is collected from past driving experiences. By feeding this data into the emotion recognition network, a set of historical emotion features is generated, reflecting the driver's emotional state during these past driving experiences. Vehicle operation log data records various vehicle operation information during driving, such as steering angle, braking frequency, and lane deviation. This information can be converted into operation risk indicators. For example, the magnitude of operation risk can be assessed based on the magnitude of changes in steering angle, braking frequency, and lane deviation.

[0148] For example, historical driving speech data from the past month is fed into an emotion recognition network to generate a set of historical emotion features. Simultaneously, steering wheel angle, braking frequency, and lane deviation are extracted from vehicle operation log data. Steering wheel angle changes exceeding a preset threshold, excessive braking frequency, and excessive lane deviation are converted into indicators of high control risk.

[0149] Step S302: Through the feature fusion module in the security risk prediction model, the historical emotion feature set and the manipulation risk index are time-aligned and spatially mapped to generate a joint input feature vector.

[0150] The feature fusion module is the component of the security risk prediction model that integrates different types of features. Temporal alignment aligns the historical sentiment feature set with the manipulation risk indicator over time, enabling correlation between data from the same point in time. Spatial mapping maps different types of features into the same feature space for subsequent processing. Through temporal alignment and spatial mapping, the historical sentiment feature set and the manipulation risk indicator are fused together to generate a joint input feature vector.

[0151] For example, the historical sentiment feature set records the intensity of sentiment at each point in time, while the manipulation risk indicator records the magnitude of manipulation risk at each point in time. By temporally aligning these, the sentiment intensity and manipulation risk magnitude at the same point in time can be correlated. Then, through spatial mapping, the sentiment intensity and manipulation risk magnitude are mapped into a two-dimensional feature space to generate a joint input feature vector.

[0152] Step S303: The risk prediction module in the safety risk prediction model predicts the probability of an accident in the future time window based on the joint input feature vector, and calculates the prediction error loss according to the actual accident time point in the actual accident record data.

[0153] The risk prediction module is the component of the safety risk prediction model that predicts accident probability. Based on the combined input feature vectors, it uses a pre-defined algorithm and model to predict the probability of an accident within a future time window. Actual accident records record the actual time points of past accidents. The model's prediction accuracy is evaluated by comparing the predicted accident probability with actual accident records and calculating the prediction error loss.

[0154] For example, based on the joint input feature vector, the risk prediction module predicts a 20% probability of an accident occurring within the next 10 minutes. However, actual accident records show that no accidents occurred within the next 10 minutes. In this case, the prediction error loss can be calculated to measure the model's prediction bias.

[0155] Step S304: Iteratively optimize the safety risk prediction model based on the prediction error loss until the prediction accuracy of the accident probability reaches a preset convergence condition.

[0156] Iterative optimization is the process of continuously adjusting the parameters of the safety risk prediction model based on the prediction error loss, thereby continuously improving the model's prediction accuracy. The preset convergence condition is a pre-set accuracy limit. When the predicted accuracy of the accident probability reaches this limit, the model is considered converged and the iterative optimization process ends.

[0157] For example, by continuously adjusting the parameters of the safety risk prediction model, the prediction error loss is gradually reduced. When the prediction accuracy of the accident probability reaches the preset convergence condition, such as the prediction error loss is less than 0.05, the iterative optimization stops, and the model's prediction accuracy meets the requirements.

[0158] As an implementation method, the multi-dimensional driving data joint training may further include:

[0159] Step S305: Introducing an attention weight allocation mechanism into the feature fusion module to adjust the contribution ratio of the historical sentiment feature set and the manipulation risk indicator in the joint input feature vector;

[0160] The attention weight allocation mechanism adjusts the contribution of different features to the joint input feature vector. By introducing this mechanism, different weights can be assigned to the historical emotion feature set and manipulation risk indicator based on their importance to accident probability prediction, allowing the joint input feature vector to more reasonably reflect the impact of various factors on driving safety.

[0161] For example, if it is found that the historical emotional feature set has a greater impact on the prediction of the probability of an accident, then the attention weight distribution mechanism can be used to increase the weight of the historical emotional feature set in the joint input feature vector and reduce the weight of the manipulation risk indicator.

[0162] Among them, the attention weight allocation mechanism includes:

[0163] Step S3051: determining the subset of emotion features in the historical emotion feature set that overlaps with the manipulation risk indicator in time through time correlation analysis;

[0164] Temporal correlation analysis is a method used to analyze the temporal relationship between historical sentiment feature sets and manipulation risk indicators. This analysis identifies the temporal overlap between historical sentiment feature sets and manipulation risk indicators, forming a subset of sentiment features. This subset, which includes sentiment features that change synchronously with the manipulation risk indicators, may be more meaningful in predicting accident probability.

[0165] For example, the historical sentiment feature set records daily changes in sentiment intensity over the past month, while the manipulation risk indicator records the magnitude of manipulation risk each day during the same time period. A temporal correlation analysis reveals that within a week, the trends in sentiment intensity and manipulation risk are relatively consistent. The sentiment features corresponding to that week constitute the sentiment feature subset.

[0166] Step S3052: Calculate the covariance matrix between the emotion feature subset and the manipulation risk indicator, and generate the emotion feature weights and the manipulation indicator weights based on the eigenvalue distribution of the covariance matrix;

[0167] The covariance matrix is ​​used to measure the covariance relationship between two variables. By calculating the covariance matrix between a subset of sentiment features and a manipulation risk indicator, we can understand the correlation between them. The eigenvalue distribution of the covariance matrix reflects the contribution of each variable to the overall change. Based on the eigenvalue distribution, we can generate sentiment feature weights and manipulation indicator weights. A larger weight indicates a greater contribution of the feature to the joint input feature vector.

[0168] For example, the covariance matrix between the sentiment feature subset and the manipulation risk indicator is calculated, and the eigenvalue distribution is obtained by performing eigenvalue decomposition on the covariance matrix. Based on the eigenvalue distribution, a weight of 0.6 is assigned to the sentiment feature subset and a weight of 0.4 is assigned to the manipulation risk indicator.

[0169] Step S3053: Perform weighted fusion on the emotion feature subset and the manipulation risk index according to the emotion feature weight and the manipulation index weight to generate a joint input feature vector.

[0170] Weighted fusion is the process of combining a subset of emotional features and a manipulation risk indicator, weighted by the weights of the emotional features and the manipulation risk indicator. This process combines different features in appropriate proportions to generate a combined input feature vector that more accurately reflects the impact of various factors on driving safety.

[0171] For example, the values ​​of the sentiment feature subset are [0.2, 0.3, 0.4], with a weight of 0.6; the values ​​of the manipulation risk indicator are [0.1, 0.2, 0.3], with a weight of 0.4. After weighted fusion, the resulting joint input feature vector is [0.2×0.6+0.1×0.4, 0.3×0.6+0.2×0.4, 0.4×0.6+0.3×0.4]=[0.16, 0.26, 0.36].

[0172] As an implementation manner, after generating the driving intervention prompt information in step S3053, the method provided by the embodiment of the present invention may further include:

[0173] Step S3054: monitoring the driver's response to the driving intervention prompt in real time, and collecting an updated driving voice data stream after the response;

[0174] Real-time monitoring involves the continuous observation and recording of driver actions. When a driver intervention prompt is generated, the driver may respond accordingly, such as adjusting driving behavior or replying to voice commands. By monitoring these responses in real time and collecting the updated driving voice data stream following these responses, we can understand the driver's reaction to the prompts, providing data support for subsequent analysis and processing.

[0175] For example, the driving intervention prompt information prompts the driver to slow down, monitors in real time whether the driver steps on the brake pedal to slow down the vehicle, and collects the voice emitted by the driver after making the response operation to form an updated driving voice data stream.

[0176] Step S3055: extracting emotional features from the updated driving voice data stream to generate an updated safety behavior score curve;

[0177] The process for extracting emotional features from the updated driving voice data stream is similar to the previously described process. By extracting emotional features from the updated driving voice data stream, a multi-level emotional feature set is generated. This feature set is then input into the safety risk prediction model to generate an updated safety behavior score curve. The updated safety behavior score curve reflects the driver's emotional state and changes in driving safety risk after responding to driving intervention prompts.

[0178] For example, the updated driving voice data stream is processed according to the previous steps to extract emotional features. This is then input into the safety risk prediction model to generate an updated safety behavior score curve. If the driver responds positively to the prompt information and becomes more emotionally stable, the updated safety behavior score curve may decrease, indicating a reduced driving safety risk.

[0179] Step S3056: If the updated safety behavior score curve does not return to the safety threshold range within the preset time, the active intervention protocol of the vehicle control system is triggered, and the active intervention protocol includes at least one of automatically reducing the driving speed, starting the emergency braking preparation mode, and sending an assistance request to the traffic management center.

[0180] The safety threshold range is a pre-set range of safety score values. When the safety behavior score curve falls within this range, it indicates that the driving safety risk is at an acceptable level. If the updated safety behavior score curve fails to return to the safety threshold range within the preset time, it indicates that the driver's emotional state and driving safety risk remain high, and the vehicle control system's active intervention protocol needs to be triggered. The active intervention protocol includes various measures, such as automatically reducing driving speed to reduce the possibility of accidents, activating the emergency brake standby mode for faster response in emergency situations, and sending an assistance request to the traffic management center to obtain external support and assistance.

[0181] For example, if the preset time is 5 minutes and the safety threshold range is 70-100 points, if the updated safety behavior score curve remains below 70 points for 5 minutes, the vehicle control system's active intervention protocol will be triggered, automatically reducing the driving speed and sending an assistance request to the traffic management center.

[0182] Step S400: Based on the risk nodes exceeding a preset threshold in the safety behavior score curve, driving intervention prompt information corresponding to the risk nodes is generated, and the driving intervention prompt information is sent to the on-board terminal of the target vehicle for real-time display.

[0183] The preset threshold is a pre-set safety score limit. Any safety score node exceeding this limit is considered a risk node. Driving intervention prompts are generated based on the risk node to remind the driver to pay attention to driving safety and take appropriate measures. These driving intervention prompts are sent to the target vehicle's onboard terminal for real-time display, allowing drivers to promptly understand their driving risk situation and make appropriate adjustments.

[0184] For example, if the preset threshold is 80 points and a node in the safety behavior scoring curve scores 90 points, exceeding the preset threshold, this node is considered a risk node. Based on the emotional state and driving operation risk corresponding to this risk node, a driving intervention prompt message is generated, such as "You are currently quite emotional. Please remain calm and drive safely." This prompt message is sent to the in-vehicle terminal's display screen for real-time display.

[0185] As an embodiment, step S400, based on the risk nodes exceeding a preset threshold in the safety behavior score curve, generates driving intervention prompt information corresponding to the risk nodes, which can be implemented as follows:

[0186] Step S410: performing a gradient test on the safety behavior score curve, identifying consecutive time intervals in which the rate of decrease of the safety score value exceeds a rate threshold within a preset time interval, and marking the consecutive time intervals as high-risk periods;

[0187] Gradient detection is the process of calculating and analyzing the slope of the safety behavior score curve. Gradient detection identifies the rate of decline of the safety score within a preset time interval. The rate threshold is a pre-set limit for the rate of decline. When the rate of decline of the safety score exceeds this limit, the corresponding continuous time interval is marked as a high-risk period. High-risk periods indicate that the driver's driving safety risk has increased dramatically during this period, requiring special attention.

[0188] For example, if the preset time interval is 10 minutes and the rate threshold is set at a drop of 5 points per minute, a gradient test of the safety behavior score curve reveals that from the 3rd to 6th minute, the safety score drops from 85 to 70, with a rate of decline exceeding 5 points per minute. Therefore, this 3- to 6-minute interval is marked as a high-risk period. This indicates that during this period, the driver's emotions or driving behavior may have fluctuated significantly, leading to a significant increase in driving safety risks.

[0189] Step S420: extracting the dominant emotion type from the multi-level emotion feature set corresponding to the high-risk period, and selecting a matching prompt content template from a preset intervention strategy library based on the dominant emotion type;

[0190] The dominant emotion type is the emotion that dominates the multi-level emotional signature set corresponding to high-risk periods, reflecting the driver's primary emotional state during those periods. The preset intervention strategy library is a pre-established database containing a variety of prompt content templates tailored to different emotion types. Selecting a matching prompt content template from the intervention strategy library based on the dominant emotion type makes the generated driving intervention prompts more targeted and effective.

[0191] For example, analyzing the multi-level emotional signature set corresponding to high-risk periods reveals that the dominant emotion is "anger." Within the pre-set intervention strategy library, a prompt template for this emotion might include the following: "You seem a little angry right now. Please calm down first to avoid emotional agitation that could affect driving safety." By selecting this matching prompt template, guidance and reminders can be provided directly targeting the driver's current anger.

[0192] Step S430: generating supplementary warning information in a prompt content template based on the vehicle control deviation probability and the current vehicle driving parameters, wherein the supplementary warning information includes recommended adjustment parameters related to the steering wheel angle and the accelerator pedal depth;

[0193] The vehicle control deviation probability reflects the likelihood of a driver's control deviation during vehicle control operations based on their current emotional state. Current vehicle driving parameters include steering wheel angle, accelerator pedal depth, and vehicle speed. This information is used to generate supplementary warning information within the prompt content template, providing drivers with more specific operational recommendations, helping them adjust their driving behavior and reduce driving risks.

[0194] For example, if the vehicle control deviation probability indicates a high likelihood of steering wheel angle deviation due to anger, and current vehicle driving parameters indicate excessive steering angle, a supplementary warning message might be, "Your steering wheel angle is too large. Please reduce it appropriately to maintain stable driving." Alternatively, if the accelerator pedal is pressed too deeply and the vehicle control deviation probability indicates a risk of overacceleration, a supplementary warning message might be, "Your accelerator pedal is pressed too deeply. Please relax and control your speed."

[0195] Step S440: Combine the prompt content template and the supplementary warning information to generate driving intervention prompt information, and send the driving intervention prompt information to different display areas of the vehicle terminal according to preset priority.

[0196] Combining the prompt content template with supplementary warning information creates a complete driver intervention prompt. The preset priority is a pre-set order based on the importance of the prompt information. Different display areas can be different locations on the vehicle terminal display, such as the main screen, secondary screen, and head-up display area. Sorting the prompts according to the preset priority and sending them to different display areas of the vehicle terminal ensures that the driver can see the most important prompts promptly and clearly.

[0197] For example, the prompt content template is "You seem a little angry now, please calm down first to avoid affecting driving safety due to emotional excitement." The supplementary warning information is "Your steering wheel angle is too large, please reduce the angle appropriately to keep the vehicle driving smoothly." After combining them, the driving intervention prompt information "You seem a little angry now, please calm down first to avoid affecting driving safety due to emotional excitement. Your steering wheel angle is too large, please reduce the angle appropriately to keep the vehicle driving smoothly." Assuming that the preset priority of this prompt information is high, it will be sent to the main screen of the vehicle terminal for display so that the driver can see it and take corresponding measures at the first time.

[0198] See also Figure 2 , Figure 2 This is a schematic diagram of the structure of a driving safety analysis device provided in an embodiment of the present invention. The driving safety analysis device is, for example, an on-board terminal on a vehicle, or a server communicatively connected to the on-board terminal. The driving safety analysis device includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 may be connected via a bus or other means. The processor 101 (also known as the central processing unit (CPU)) is the computing and control core of the driving safety analysis device, capable of parsing various instructions within the driving safety analysis device and processing various data within the driving safety analysis device. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi or a mobile communication interface), and may be used to send and receive data under the control of the processor 101. The communication interface 102 may also be used for data transmission and interaction within the driving safety analysis device. The memory 103 is a storage device within the driving safety analysis device, used to store programs and data. It is understood that the memory 103 herein may include both the built-in memory of the driving safety analysis device and the extended memory supported by the driving safety analysis device. The memory 103 provides a storage space for storing the operating system of the driving safety analysis device, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc., and the present invention is not limited to this.

[0199] In one embodiment, the processor 101 executes the driving safety analysis method based on speech emotion recognition provided in the above embodiment of the present invention by running the computer program in the memory 103.

Claims

1. A driving safety analysis method based on speech emotion recognition, characterized in that: The method comprises: Acquire a driving voice data stream collected by a target vehicle in a driving environment, wherein the driving voice data stream includes continuous voice segments of the driver within a preset time period and interactive instructions related to vehicle control; Extracting emotional features from the driving voice data stream to generate a multi-level emotional feature set, wherein the multi-level emotional feature set includes a real-time emotional fluctuation trajectory corresponding to the continuous voice segment and a semantic emotional intensity parameter corresponding to the interactive instruction; Inputting the multi-level emotional feature set into a safety risk prediction model to generate a safety behavior scoring curve for the driver within the preset time period, wherein the safety behavior scoring curve is used to characterize the correlation between the driver's emotional stability and driving operation risk at different time points; generating driving intervention prompt information corresponding to a risk node exceeding a preset threshold in the safety behavior score curve, and sending the driving intervention prompt information to an onboard terminal of the target vehicle for real-time display; The step of inputting the multi-level emotional feature set into the safety risk prediction model to generate the driver's safety behavior score curve within the preset time period includes: The time series analysis module in the safety risk prediction model is used to periodically decompose the real-time emotion fluctuation trajectory and extract the frequency and amplitude of the driver's emotion mutation within a preset time window; Matching the semantic emotion intensity parameter with a predefined driving operation risk level through a risk association mapping module in the safety risk prediction model to determine the vehicle control deviation probability corresponding to different emotion intensities; Calculating the driver's safety score at multiple consecutive time points by combining the emotion mutation frequency, the mutation amplitude, and the vehicle control deviation probability through a dynamic score generation module in the safety risk prediction model; The safety score values ​​are connected in chronological order to generate the safety behavior score curve, and the safety behavior score curve is smoothed to eliminate short-term noise interference.

2. The method according to claim 1, characterized in that The step of obtaining a driving voice data stream collected by the target vehicle in a driving environment includes: Segmenting the driving voice data stream to obtain multiple voice data units of equal duration, and filtering each voice data unit for ambient noise to generate multiple denoised voice segments; wherein the segmentation includes endpoint detection based on voice energy mutations and context segmentation based on semantic coherence; the ambient noise filtering includes eliminating vehicle engine noise through frequency domain filtering and separating the driver's voice from the passenger's voice through voiceprint feature matching; The multiple denoised speech segments are spliced ​​in chronological order to generate a standardized speech data sequence corresponding to the driving speech data stream.

3. The method according to claim 2, characterized in that The step of extracting emotion features from the driving voice data stream to generate a multi-level emotion feature set includes: Inputting the standardized speech data sequence into a pre-trained emotion recognition network, and extracting the pitch change pattern, speech rate fluctuation frequency, and speech amplitude distribution parameters of each denoised speech segment through the acoustic feature extraction layer in the emotion recognition network; Identifying instruction keywords, emotion modifiers, and modal particles contained in the denoised speech segments through the semantic parsing layer in the emotion recognition network, and generating semantic emotion tags corresponding to each of the denoised speech segments; Analyzing the emotion transition features between adjacent denoised speech segments through a context association layer in the emotion recognition network, and generating the real-time emotion fluctuation trajectory based on the emotion transition features; The pitch change pattern, the speaking rate fluctuation frequency, the speech amplitude distribution parameter, the semantic emotion tag and the real-time emotion fluctuation trajectory are subjected to feature fusion to generate the multi-level emotion feature set.

4. The method according to claim 1, wherein The step of generating driving intervention prompt information corresponding to a risk node exceeding a preset threshold in the safety behavior scoring curve includes: Performing a gradient test on the safety behavior score curve to identify consecutive time intervals in which the rate of decrease of the safety score value exceeds a rate threshold within a preset time interval, and marking the consecutive time intervals as high-risk periods; Extracting the dominant emotion type from the multi-level emotion feature set corresponding to the high-risk period, and selecting a matching prompt content template from a preset intervention strategy library based on the dominant emotion type; generating supplementary warning information in the prompt content template based on the vehicle control deviation probability and current vehicle driving parameters, the supplementary warning information including suggested adjustment parameters related to the steering wheel angle and the accelerator pedal depth; The prompt content template is combined with the supplementary warning information to generate the driving intervention prompt information, and the driving intervention prompt information is sent to different display areas of the vehicle terminal according to preset priority.

5. The method according to claim 3, characterized in that The emotion recognition network is obtained by jointly training the acoustic feature extraction layer, the semantic parsing layer, and the context association layer. The joint training includes: Obtaining a labeled training speech dataset, wherein each speech sample in the training speech dataset is labeled with a tone feature label, a semantic emotion label, and a contextual emotion continuity label; Inputting the speech sample into the initial emotion recognition network, generating a tone prediction feature through the acoustic feature extraction layer, generating a semantic emotion prediction label through the semantic parsing layer, and generating an emotion continuity prediction value through the context association layer; Calculating a first loss function between the tone prediction feature and the tone feature label, calculating a second loss function between the semantic emotion prediction label and the semantic emotion label, and calculating a third loss function between the emotion continuity prediction value and the context emotion continuity label; A weighted sum is performed on the first loss function, the second loss function, and the third loss function to generate a joint training loss value, and the parameters of the initial emotion recognition network are updated through a back propagation algorithm until convergence.

6. The method according to claim 5, characterized in that The joint training also includes: Adding an interfering speech sample containing vehicle environmental noise to the training speech data set, and setting a noise suppression weight corresponding to the interfering speech sample; When extracting features from the interference speech sample through the acoustic feature extraction layer, reducing the contribution of frequency band features corresponding to vehicle engine noise in the interference speech sample based on the noise suppression weight; When the interfering speech sample is parsed by the semantic parsing layer, the keyword retrieval strength related to the driver's speech is increased based on the noise suppression weight to cover the semantic ambiguity area caused by noise.

7. The method according to claim 1, characterized in that The safety risk prediction model is obtained by joint training of multi-dimensional driving data, wherein the multi-dimensional driving data includes historical driving voice data, vehicle operation log data, and actual accident record data; The multi-dimensional driving data joint training includes: Inputting the historical driving voice data into an emotion recognition network to generate a historical emotion feature set, and converting the steering wheel angle, braking frequency, and lane deviation in the vehicle operation log data into a control risk indicator; The historical sentiment feature set and the manipulation risk indicator are temporally aligned and spatially mapped by a feature fusion module in the security risk prediction model to generate a joint input feature vector; Predicting the probability of an accident occurring in a future time window based on the combined input feature vector by a risk prediction module in the safety risk prediction model, and calculating a prediction error loss according to the actual accident time point in the actual accident record data; The safety risk prediction model is iteratively optimized based on the prediction error loss until the prediction accuracy of the accident probability reaches a preset convergence condition.

8. The method according to claim 7, characterized in that The multi-dimensional driving data joint training also includes: Introducing an attention weight allocation mechanism into the feature fusion module to adjust the contribution ratio of the historical sentiment feature set and the manipulation risk indicator in the joint input feature vector; The attention weight allocation mechanism includes: Determining, through time correlation analysis, a subset of emotional features in the historical emotional feature set that overlaps with the manipulation risk indicator in time; Calculating a covariance matrix between the subset of emotional features and the manipulation risk indicator, and generating emotional feature weights and manipulation indicator weights based on an eigenvalue distribution of the covariance matrix; performing weighted fusion of the emotion feature subset and the manipulation risk index according to the emotion feature weight and the manipulation index weight to generate the joint input feature vector; After generating the driving intervention prompt information, the method further includes: monitoring the driver's response to the driving intervention prompt information in real time, and collecting an updated driving voice data stream after the response operation; Extracting emotional features from the updated driving voice data stream to generate an updated safety behavior scoring curve; If the updated safety behavior score curve does not return to the safety threshold range within the preset time, the active intervention protocol of the vehicle control system is triggered, and the active intervention protocol includes at least one of automatically reducing the driving speed, starting the emergency braking preparation mode, and sending an assistance request to the traffic management center.

9. A driving safety analysis device, characterized in that: include: a memory storing a computer program; A processor, configured to load the computer program to implement the driving safety analysis method based on speech emotion recognition as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Driving behavior risk detection method, system and device and storage medium

    CN119559620A