Driving safety analysis method based on voice emotion recognition and related device

By obtaining the multi-level emotional feature set and safety risk prediction model of driving voice data flow, the problem that sensors cannot capture driver emotional fluctuations in traditional methods is solved, and early prediction and precise intervention of driving safety risks is achieved, false alarm rates and hardware costs are reduced, and driving safety is improved.

CN120340541AActive Publication Date: 2025-07-18GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS

Patent Information

Application Number
CN202510839215.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-18
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

In the existing driving safety analysis methods, traditional sensors cannot capture driver emotional fluctuations, resulting in high warning false alarm rate, weak risk traceability, and high hardware deployment costs; traditional voice emotion recognition technology cannot analyze the semantic association between emotional intensity fluctuations and vehicle interaction instructions in continuous voice clips, resulting in delayed warning response.

Method used

By obtaining driving voice data streams, multi-level emotional characteristics collections, including real-time emotional fluctuations trajectory and semantic emotional intensity parameters, inputting safety risk prediction models to generate safety behavior score curves, combining risk nodes to generate driving intervention prompt information, and using the implicit emotional continuity change laws and semantic instructions to realize early prediction and precise intervention of driving safety risks.

Benefits of technology

Effectively distinguish between drivers' normal emotional expression and dangerous emotional tendencies, reduce hardware deployment costs, improve driving safety warning accuracy, reduce traffic accidents, and achieve precise intervention in complex driving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340541A_ABST
    Figure CN120340541A_ABST
Patent Text Reader

Abstract

The invention provides a driving safety analysis method based on voice emotion recognition and a related device, and the method comprises the steps: obtaining a driving voice data flow collected by a target vehicle in a driving environment, carrying out the emotion feature extraction of the driving voice data flow, and generating a multi-level emotion feature set; inputting the multi-level emotion feature set into a safety risk prediction model, generating a safety behavior scoring curve of the driver in a preset time period, and generating driving intervention prompt information corresponding to risk nodes according to the risk nodes exceeding a preset threshold in the safety behavior scoring curve, and the driving intervention prompt information is sent to the vehicle-mounted terminal of the target vehicle for real-time display. According to the invention, hardware deployment cost can be reduced, interference of voice quality fluctuation on an analysis result in a complex driving environment is ensured to be effectively suppressed, and effects of improving driving safety early warning accuracy and reducing traffic accident rate are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular, to a driving safety analysis method and related device based on speech emotion recognition. Background Art

[0002] Driving safety analysis aims to improve driving safety by monitoring driver states and vehicle behaviors to identify potential risks. Existing driving safety analysis usually uses in-vehicle sensors to collect control data such as steering wheel steering angles and braking frequencies, or monitors physiological indicators such as driver heart rates and eye movement trajectories through biosensors, and then generates safety warnings based on threshold judgment rules. However, in existing methods, the indirect inference of the driver's emotional state has a lag. Especially when the driver's emotional fluctuations have not caused obvious operation abnormalities, traditional sensors cannot capture dangerous emotional characteristics such as anger and anxiety hidden in speech, and the static warning mechanism based on fixed thresholds is difficult to adapt to the dynamic relationship between emotions and control behaviors in complex road environments, resulting in a high false alarm rate of warnings and weak risk tracing ability; in addition, when using independent physiological sensors to monitor emotional states, there are problems of high hardware deployment costs and susceptibility to environmental interference, while traditional speech emotion recognition technologies only classify emotions for discrete keywords and cannot analyze the semantic relationship between the emotional intensity fluctuations in continuous speech segments and vehicle interaction instructions, resulting in a significant delay in warning responses for scenarios such as distracted driving and sudden road rage.

[0003] It should be noted that the above background introduction of the present invention is only to introduce the technical field involved in the present invention. Without the premise that the above technical features and technical problems are disclosed in the prior art, it cannot be used as a basis for judging the creativity of the present invention. Summary of the Invention

[0004] The present invention provides a driving safety analysis method and related device based on speech emotion recognition.

[0005] In a first aspect, an embodiment of the present invention provides a driving safety analysis method based on voice emotion recognition. The method includes: obtaining a driving voice data stream collected by a target vehicle in a driving environment, where the driving voice data stream includes continuous voice segments of a driver within a preset time period and interaction instructions related to vehicle control; extracting emotion features from the driving voice data stream to generate a multi-level emotion feature set, where the multi-level emotion feature set includes a real-time emotion fluctuation trajectory corresponding to the continuous voice segments and a semantic emotion intensity parameter corresponding to the interaction instructions; inputting the multi-level emotion feature set into a safety risk prediction model to generate a safety behavior score curve of the driver within the preset time period, where the safety behavior score curve is used to characterize the correlation between the driver's emotional stability and driving operation risk at different time nodes; generating driving intervention prompt information corresponding to the risk nodes according to the risk nodes exceeding a preset threshold in the safety behavior score curve, and sending the driving intervention prompt information to an in-vehicle terminal of the target vehicle for real-time display.

[0006] In a second aspect, an embodiment of the present invention provides a driving safety analysis device, including: a memory in which a computer program is stored; a processor configured to load the computer program to implement the above-described driving safety analysis method based on voice emotion recognition.

[0007] The driving safety analysis method based on voice emotion recognition provided by the present invention obtains continuous voice segments and vehicle control interaction instructions in the driving voice data stream collected in the driving environment of the target vehicle, extracts a multi-level emotion feature set including real-time emotion fluctuation trajectories and semantic emotion intensity parameters, generates a safety behavior score curve based on a safety risk prediction model and combines risk nodes to generate driving intervention prompt information. It can dynamically associate the emotion fluctuation characteristics of the driver with the vehicle control behavior, and utilize the implicit emotion continuity change law in the voice data and the real-time intention analysis of semantic instructions to realize the early prediction and precise intervention of driving safety risks, thereby overcoming the defects of high false alarm rates and difficult risk tracing in traditional systems. Through the cross-validation of multi-level features such as the pitch change pattern in continuous voice segments, the speech rate fluctuation frequency, and the semantic emotion intensity in interaction instructions, it can effectively distinguish the normal emotion expressions of drivers from dangerous emotion tendencies. Combining the two-dimensional detection mechanism of preset thresholds and risk nodes in the safety behavior score curve, it can identify the vehicle control deviation trend caused by anger, anxiety, or fatigue at the initial stage of emotion fluctuation, and generate differentiated intervention prompt content by dynamically associating the current driving environment parameters, significantly improving the pertinence of warning information and the driver response efficiency. In addition, based on the real-time processing and closed-loop feedback mechanism of the voice data stream, only an in-vehicle voice collection device is required to realize emotion risk analysis. While reducing the hardware deployment cost, through the generation of standardized voice data sequences and the fusion and optimization of multi-level features, it effectively suppresses the interference of voice quality fluctuations in complex driving environments on the analysis results, and finally achieves the effects of improving the accuracy of driving safety warnings and reducing the incidence of traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0009] Figure 1 is a flowchart of a driving safety analysis method based on voice emotion recognition provided by an embodiment of the present invention.

[0010] Figure 2 is a schematic diagram of the composition of a driving safety analysis device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0012] Please refer to Figure 1 , Figure 1 which is a flowchart of a driving safety analysis method based on voice emotion recognition provided by an embodiment of the present invention. This driving safety analysis method based on voice emotion recognition can be executed by a driving safety analysis device, and this driving safety analysis method based on voice emotion recognition may include the following steps: Step S100: Obtain a driving voice data stream collected by a target vehicle in a driving environment. The driving voice data stream includes continuous voice segments of a driver within a preset time period and interaction instructions related to vehicle control.

[0013] The driving voice data stream refers to the voice data collected during the driving process of the target vehicle, which includes continuous voice segments uttered by the driver within a preset time period. These voice segments may be the driver's daily communication words, emotional expressions, etc., and also include interaction instructions related to vehicle control, such as instructions like "turn on the air conditioner" and "close the window". The preset time period is a time range preset according to actual needs, used to define the time limit for collecting driving voice data.

[0014] In actual operation, the target vehicle can collect the driving voice data stream through a microphone device installed in the vehicle. For example, in an intelligent vehicle, the microphone array in the vehicle can collect the voice signals emitted by the driver in all directions and convert them into digital signals for storage and subsequent processing.

[0015] As an implementation manner, in step S100, obtaining the driving voice data stream collected by the target vehicle in the driving environment can be implemented as the following steps: Step S110: Perform segmentation processing on the driving voice data stream to obtain multiple voice data units of equal duration, and perform environmental noise filtering on each voice data unit to generate multiple denoised voice segments; wherein, the segmentation processing includes endpoint detection based on voice energy mutation and context division based on semantic coherence; the environmental noise filtering includes eliminating vehicle engine noise through frequency domain filtering and separating the driver's voice from the passenger's voice through voiceprint feature matching.

[0016] Segmentation processing is the process of dividing the collected driving voice data stream into multiple voice data units of equal duration according to preset rules. Endpoint detection based on voice energy mutation refers to determining the start and end points of voice by detecting the positions where the energy in the voice signal suddenly changes. For example, when the driver starts speaking, the energy of the voice signal will suddenly increase, and the start endpoint of the voice can be found by detecting this energy mutation. Context segmentation based on semantic coherence divides parts with coherent semantics into one voice data unit according to the semantic logical relationship of the voice content. For example, when the driver continuously says "Turn on the air conditioner and set the temperature to 25 degrees", these two sentences have coherent semantics and can be divided into one voice data unit.

[0017] Ambient noise filtering is to remove the interfering noise in the driving voice data stream and improve the quality of the voice data. Frequency domain filtering is a signal processing method that can convert the voice signal into the frequency domain and then remove the noise in a certain frequency range generated by the vehicle engine and other sources according to the characteristics of different frequency components. Voiceprint feature matching uses the unique voiceprint features of each person to distinguish the driver's voice from the passenger's voice. Since everyone's voiceprint is unique like a fingerprint, by comparing the collected voice signal with the pre-registered driver's voiceprint features, the driver's voice can be separated.

[0018] For example, during the driving of a car, the voice data stream collected by the in-vehicle microphone contains the roar of the engine, the conversations of passengers, and the driver's instructions. First, through endpoint detection based on voice energy mutation and context segmentation based on semantic coherence, the voice data stream is segmented into multiple voice data units of equal duration. Then, frequency domain filtering is used to remove the engine noise, and voiceprint feature matching is used to separate the driver's voice, finally obtaining multiple denoised voice segments.

[0019] Step S120: Concatenate the multiple denoised voice segments in chronological order to generate a standardized voice data sequence corresponding to the driving voice data stream.

[0020] The standardized voice data sequence is a voice data sequence obtained by re-concatenating multiple denoised voice segments after segmentation and denoising processing in the chronological order in which they appear in the original driving voice data stream. The purpose of doing this is to obtain a complete, ordered voice data sequence with noise interference removed for subsequent processing such as emotion feature extraction.

[0021] For example, after the processing in step S110, three denoised speech segments are obtained, namely "turn on", "air conditioner", and "set the temperature to 25 degrees". These three segments are obtained in the order of the driver's speech. In step S120, these three denoised speech segments are concatenated in chronological order to obtain a standardized speech data sequence like "turn on the air conditioner and set the temperature to 25 degrees".

[0022] Step S200: Extract emotional features from the driving speech data stream to generate a multi-level emotional feature set. Among them, the multi-level emotional feature set includes a real-time emotional fluctuation trajectory corresponding to continuous speech segments and a semantic emotional intensity parameter corresponding to an interaction instruction.

[0023] Emotional feature extraction is a process of extracting feature information from the driving speech data stream that can reflect the driver's emotional state. The multi-level emotional feature set is a set containing various emotional features. The real-time emotional fluctuation trajectory refers to the change of the driver's emotion over time during continuous speech segments, which can be obtained by analyzing features such as the pitch, speech rate, and amplitude of the speech. The semantic emotional intensity parameter is related to the interaction instruction and is used to measure the emotional intensity contained in the interaction instruction. For example, an instruction like "turn on the air conditioner immediately!" may contain a stronger emotional intensity than "turn on the air conditioner".

[0024] For example, in a driving speech data stream, the driver initially says "open the window" in a calm tone, and then becomes a bit irritable due to the rising temperature in the car and shouts "turn the temperature down!". Through emotional feature extraction, the real-time emotional fluctuation trajectory of the driver during this speech process, from calm to irritable, can be obtained, and at the same time, a relatively high semantic emotional intensity parameter corresponding to the interaction instruction "turn the temperature down!" can also be obtained.

[0025] As an implementation, in step S200, extracting emotional features from the driving speech data stream to generate a multi-level emotional feature set can be implemented as the following steps: Step S210: Input the standardized speech data sequence into a pre-trained emotion recognition network, and extract the pitch change pattern, speech rate fluctuation frequency, and speech amplitude distribution parameters of each denoised speech segment through the acoustic feature extraction layer in the emotion recognition network.

[0026] The pre-trained emotion recognition network is a neural network model trained with a large amount of speech data. It can be used to recognize emotional information in speech. The acoustic feature extraction layer is a component of the emotion recognition network. Its main function is to extract acoustic features related to emotions from the input standardized speech data sequence. The pitch change pattern refers to the change of the pitch of the speech at different time points. For example, the increase or decrease of the pitch may reflect the driver's emotional changes. The pitch may increase when excited, and the pitch is relatively stable when calm. The speech rate fluctuation frequency refers to the frequency of change of the speech rate. The faster the speech rate, the more nervous or excited the driver is. The slower the speech rate, the more relaxed or thoughtful the driver is. The speech amplitude distribution parameter refers to the distribution of the amplitude of the speech signal. A larger amplitude may indicate that the driver is more emotional, and a smaller amplitude may indicate that the emotion is calmer.

[0027] For example, by inputting the standardized speech data sequence of "turn on the air conditioner and set the temperature to 25 degrees" into the pre-trained emotion recognition network, the acoustic feature extraction layer can analyze that when saying "turn on", the pitch is relatively stable, the speaking speed is moderate, and the amplitude is normal; while when saying "adjust the temperature to 25 degrees", the pitch may rise slightly, the speaking speed is slightly faster, and the amplitude is slightly increased, thereby obtaining the corresponding pitch change pattern, speaking speed fluctuation frequency and speech amplitude distribution parameters.

[0028] Step S220: Identify the instruction keywords, emotion modifiers, and modal particles contained in the denoised speech segments through the semantic parsing layer in the emotion recognition network, and generate semantic emotion tags corresponding to each denoised speech segment.

[0029] The semantic parsing layer is the part of the emotion recognition network that processes speech semantic information. Command keywords refer to key words related to vehicle control or other operations in the denoised speech segment, such as "open", "close", "increase", "lower", etc. Emotion modifiers are words used to modify emotions, such as "very", "especially", "a little", etc., which can enhance or weaken the expression of emotions. Modal particles are words that express the mood at the end of a sentence, such as "ah", "ah", "ne", etc. Different modal particles can convey different emotions. For example, "ah" in "open it quickly!" can better reflect the urgency of the emotion than "open it quickly". Semantic emotion labels are labels that represent the type of emotions assigned to the denoised speech segment based on its semantic information, such as "calm", "angry", "anxious", etc.

[0030] For example, for the denoised speech segment "Open the window immediately!", the semantic parsing layer can identify the command keyword "open", the emotional modifier "immediately", and the modal particle "ah", and generate the corresponding semantic emotion label "urgent" based on this information.

[0031] Step S230: Analyze the emotional transition features between adjacent denoised speech segments through the context association layer in the emotion recognition network, and generate a real-time emotional fluctuation trajectory based on the emotional transition features.

[0032] The context association layer is the part of the emotion recognition network used to analyze the speech context relationship. The emotional transition features between adjacent denoised speech segments refer to the change situation from the emotion expressed by one denoised speech segment to the emotion expressed by the adjacent next denoised speech segment. For example, the transition from a calm emotion to an angry emotion. The real-time emotional fluctuation trajectory is generated based on these emotional transition features, which can intuitively display the dynamic change process of the driver's emotion over a period of time.

[0033] For example, there are two adjacent denoised speech segments. The first one is "The weather is nice today", expressing a calm and pleasant emotion, and the second one is "How is that car in front driving!" expressing an angry emotion. The context association layer can analyze that the emotional transition feature between these two segments is a sudden change from calm to angry, and then generate a real-time emotional fluctuation trajectory based on this transition feature. It can be clearly seen from the trajectory graph that the emotion rises sharply.

[0034] Step S240: Perform feature fusion on the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameter, semantic emotion label, and real-time emotional fluctuation trajectory to generate a multi-level emotional feature set.

[0035] Feature fusion is a process of integrating different types of emotional features to obtain a more comprehensive and accurate representation of emotional features. By fusing these features such as the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameter, semantic emotion label, and real-time emotional fluctuation trajectory, a multi-level emotional feature set can be generated. This set contains information reflecting the driver's emotional state from multiple acoustic and semantic perspectives.

[0036] For example, in the previous example, the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameter, semantic emotion label, and real-time emotional fluctuation trajectory of each denoised speech segment have been obtained. Through feature fusion, these information are integrated together to form a multi-level emotional feature set, which can more comprehensively reflect the driver's emotional state during the entire speech process.

[0037] As an implementation manner, Step S240, performing feature fusion on the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameter, semantic emotion label, and real-time emotional fluctuation trajectory to generate a multi-level emotional feature set, can be implemented as the following steps: Step S241: Perform time-axis normalization on the pitch change pattern to generate pitch change pattern parameters consistent with the time granularity of the normalized speech data sequence, and extract the pitch mutation time nodes corresponding to the abnormal pitch segments exceeding the preset pitch threshold in the pitch change pattern parameters.

[0038] Time-axis normalization is to uniformly process the pitch change pattern on the time axis to make it consistent with the time granularity of the normalized speech data sequence. The time granularity refers to the division precision of the data in time, for example, divided in seconds. The preset pitch threshold is a preset pitch boundary, and the pitch segments exceeding this threshold are considered abnormal pitch segments. The pitch mutation time node refers to the time point when the abnormal pitch segment starts to appear.

[0039] For example, the normalized speech data sequence is divided with a time granularity of one second per unit, while the original pitch change pattern may be recorded with a finer or coarser time granularity. Through time-axis normalization, the pitch change pattern is adjusted to be in units of seconds to obtain pitch change pattern parameters consistent with the time granularity of the normalized speech data sequence. Suppose the preset pitch threshold is 80 decibels. In the pitch change pattern parameters, it is found that a certain pitch reaches 90 decibels. Then the start time point corresponding to this abnormal pitch segment exceeding 80 decibels is the pitch mutation time node.

[0040] Step S242: Perform frequency-domain segmented sampling on the speech rate fluctuation frequency to generate speech rate fluctuation frequency distribution parameters based on a fixed time window, and perform local time alignment on the speech rate fluctuation frequency distribution parameters based on the pitch mutation time node to generate speech rate fluctuation correlation parameters synchronized with the pitch change pattern parameters.

[0041] Frequency-domain segmented sampling is the process of segmenting the speech rate fluctuation frequency in the frequency domain and sampling each frequency band. The fixed time window refers to a fixed-length time period set in time for analyzing the speech rate fluctuation frequency. Through frequency-domain segmented sampling, speech rate fluctuation frequency distribution parameters based on a fixed time window can be obtained, which reflect the distribution of the speech rate fluctuation frequency in different time periods. Local time alignment means adjusting the speech rate fluctuation frequency distribution parameters in the vicinity of the pitch mutation time node according to the pitch mutation time node to make it synchronized with the pitch change pattern parameters in time.

[0042] For example, set the fixed time window to 10 seconds, perform frequency-domain segmented sampling on the speech rate fluctuation frequency, and obtain the speech rate fluctuation frequency distribution parameters for each 10-second time period. Suppose the pitch mutation time node appears at the 20th second. Then perform local time alignment on the speech rate fluctuation frequency distribution parameters in the vicinity of this time node to make the speech rate fluctuation frequency distribution parameters synchronized with the pitch change pattern parameters in time, and generate the speech rate fluctuation correlation parameters.

[0043] Step S243: Perform dynamic range compression on the speech amplitude distribution parameters to generate compressed speech amplitude smoothing parameters, and perform time-domain superposition of the speech amplitude smoothing parameters and the speech rate fluctuation correlation parameters to generate a superimposed speech energy comprehensive parameter.

[0044] Dynamic range compression is a method for processing speech amplitude distribution parameters. It can narrow the dynamic range of speech amplitudes, making overly large or small amplitude values more balanced, thereby generating compressed speech amplitude smoothing parameters. Time-domain superposition is the process of superimposing the speech amplitude smoothing parameters and the speech rate fluctuation correlation parameters in the time domain. By superimposing, the effects of speech amplitude and speech rate fluctuations on speech energy can be comprehensively considered to generate a superimposed speech energy comprehensive parameter.

[0045] For example, in the original speech amplitude distribution parameters, the amplitude values in some time periods are very large, while those in some time periods are very small. Through dynamic range compression, these amplitude values are adjusted to obtain smoother speech amplitude smoothing parameters. Then, the speech amplitude smoothing parameters and the previously generated speech rate fluctuation correlation parameters are superimposed in time to obtain a superimposed speech energy comprehensive parameter, which more comprehensively reflects the energy characteristics of the speech.

[0046] Step S244: Convert the semantic emotion label into an emotion category encoding vector, and perform context expansion on the emotion category encoding vector based on the emotion change direction between adjacent time nodes in the real-time emotion fluctuation trajectory to generate an extended emotion category encoding sequence containing historical emotion dependency relationships.

[0047] The emotion category encoding vector is a representation that converts the semantic emotion label into vector form. Each vector element corresponds to an emotion category, and subsequent calculations and processing can be more conveniently performed in vector form. Context expansion refers to further expanding the emotion category encoding vector according to the emotion change direction between adjacent time nodes in the real-time emotion fluctuation trajectory, so that it contains the dependency relationships of historical emotions.

[0048] For example, the semantic emotion labels are "calm", "angry", "anxious", etc. They are respectively converted into emotion category encoding vectors. For example, "calm" corresponds to the vector [1, 0, 0], "angry" corresponds to the vector [0, 1, 0], and "anxious" corresponds to the vector [0, 0, 1]. Then, according to the real-time emotion fluctuation trajectory, it is found that the emotion changes from "calm" to "angry". When generating the extended emotion category encoding sequence, the historical information of this emotion change will be considered, making the encoding sequence better reflect the dynamic change process of emotions.

[0049] Step S245: Perform time slicing on the extended emotion category coding sequence according to the pitch mutation time nodes to generate multiple emotion coding subsequences, and splice the features of each emotion coding subsequence with the comprehensive speech energy parameters of the corresponding time interval to generate a preliminary fusion feature segment.

[0050] Time slicing is to divide the extended emotion category coding sequence in time according to the pitch mutation time nodes to obtain multiple emotion coding subsequences in different time periods. Feature splicing is the process of combining each emotion coding subsequence with the comprehensive speech energy parameters of the corresponding time interval. Through splicing, the acoustic features and semantic features can be preliminarily fused to generate a preliminary fusion feature segment.

[0051] For example, according to the pitch mutation time nodes obtained previously, the extended emotion category coding sequence is divided into three emotion coding subsequences, corresponding to different time periods respectively. Then, the features of each emotion coding subsequence are spliced with the comprehensive speech energy parameters in that time period to obtain three preliminary fusion feature segments.

[0052] Step S246: Filter redundant features from the preliminary fusion feature segments, and retain the feature dimensions that are strongly correlated with the emotion turning points in the real-time emotion fluctuation trajectory to generate an optimized fusion feature subset.

[0053] Redundant feature filtering is the process of removing the repeated or less contributing features to emotion expression in the preliminary fusion feature segments. An emotion turning point refers to the time point where the emotion changes significantly in the real-time emotion fluctuation trajectory. By retaining the feature dimensions that are strongly correlated with the emotion turning points, the optimized fusion feature subset can more accurately reflect the driver's emotion changes.

[0054] For example, the preliminary fusion feature segments may contain some features that have little impact on emotion expression. Through redundant feature filtering, these features are removed, and only the feature dimensions that are closely related to the emotion turning points are retained. For example, at the turning point where the emotion changes from calm to angry, the feature dimensions related to the sudden increase in pitch and the sudden acceleration of speech rate are retained, thus generating an optimized fusion feature subset.

[0055] Step S247: Cross-correlate the optimized fusion feature subset in time sequence to generate a continuous emotion feature sequence covering a preset time period, and verify the integrity of the continuous emotion feature sequence to ensure an even distribution of the feature contribution degrees of the pitch change pattern, speech rate fluctuation frequency, speech amplitude distribution parameters, semantic emotion labels, and real-time emotion fluctuation trajectory.

[0056] Cross - segment association is the process of connecting and associating the optimized fusion feature subsets in chronological order, enabling features in different time periods to form a continuous sequence. Integrity verification is to check the generated continuous emotional feature sequence to ensure that the contribution degrees of each feature are evenly distributed, avoiding the influence of some features being too large or too small.

[0057] For example, cross - segment association is performed on multiple previously generated optimized fusion feature subsets in chronological order to obtain a continuous emotional feature sequence covering a preset time period. Then, integrity verification is carried out on this sequence to check whether the contribution degrees of features such as pitch change patterns, speech rate fluctuation frequencies, speech amplitude distribution parameters, semantic emotion labels, and real - time emotional fluctuation trajectories in the sequence are balanced. If not, corresponding adjustments are made.

[0058] Step S248: Compare the verified continuous emotional feature sequence with the driver's personalized emotional baseline model, adjust the weights of abnormal emotional features that deviate from the baseline range in the continuous emotional feature sequence, and generate a multi - level emotional feature set including standardized emotional intensity and personalized correction weights.

[0059] The personalized emotional baseline model is a benchmark model established in advance based on the driver's personal speech and emotional characteristics, which reflects the driver's emotional state under normal circumstances. By comparing the verified continuous emotional feature sequence with the personalized emotional baseline model, abnormal emotional features that deviate from the baseline range in the continuous emotional feature sequence can be found. Adjusting the weights of abnormal emotional features means adjusting the weights of these abnormal features in the multi - level emotional feature set, so that the multi - level emotional feature set can not only reflect the standardized emotional intensity but also take into account the driver's personalized characteristics.

[0060] For example, the personalized emotional baseline model of a certain driver shows that their pitch is relatively stable and the speech rate is moderate under normal circumstances. In the continuous emotional feature sequence, it is found that there is a period of time when the pitch suddenly rises and the speech rate speeds up, and these features deviate from the baseline range. By adjusting the weights of these abnormal emotional features, the multi - level emotional feature set can more accurately reflect the driver's emotional state while also taking into account their personalized emotional characteristics.

[0061] As an implementation method, the emotion recognition network is obtained through joint training of an acoustic feature extraction layer, a semantic parsing layer, and a context association layer. The process of this joint training includes: Step S201: Obtain a labeled training speech data set, and each speech sample in the training speech data set is labeled with a pitch feature label, a semantic emotion label, and a context emotion continuity label.

[0062] The labeled training speech dataset is a set of speech data used to train an emotion recognition network, where each speech sample is labeled with emotion-related tag information. The pitch feature tag is a description of the pitch features in a speech sample, such as the high or low pitch, change trend, etc. The semantic emotion tag is an emotion type tag assigned to the speech sample according to its semantic content, such as "happy", "sad", "angry", etc. The context emotion continuity tag is a tag used to describe the emotion continuity between adjacent speech samples, which can reflect the transition of emotions between contexts.

[0063] For example, in a training speech dataset, there is a speech sample "I'm so happy today!", whose pitch feature tag may be labeled as "high and stable pitch", and the semantic emotion tag is labeled as "happy". If the speech before this speech sample also expresses positive emotions, then the context emotion continuity tag can be labeled as "emotion continuously positive".

[0064] Step S202: Input the speech sample into the initial emotion recognition network, and respectively generate pitch prediction features through the acoustic feature extraction layer, generate semantic emotion prediction tags through the semantic parsing layer, and generate emotion continuity prediction values through the context association layer.

[0065] The initial emotion recognition network is a neural network model in an initial state before the start of training. After inputting the speech sample into the initial emotion recognition network, the acoustic feature extraction layer will analyze the acoustic features of the speech sample and generate pitch prediction features, such as predicting the pitch change pattern and pitch level of the speech. The semantic parsing layer will process the semantic content of the speech sample and generate semantic emotion prediction tags, that is, predict the emotion type expressed by this speech sample. The context association layer will consider the context information of the speech sample and generate emotion continuity prediction values for predicting the emotion continuity between adjacent speech samples.

[0066] For example, input the speech sample "This is terrible!" into the initial emotion recognition network. The pitch prediction features generated by the acoustic feature extraction layer may be "low pitch and downward trend", the semantic emotion prediction tag generated by the semantic parsing layer is "angry", and the emotion continuity prediction value generated by the context association layer according to the context information indicates that the emotion of this speech sample is continuously negative with the emotion of the previous sample.

[0067] Step S203: Calculate the first loss function between the pitch prediction features and the pitch feature tags, calculate the second loss function between the semantic emotion prediction tags and the semantic emotion tags, and calculate the third loss function between the emotion continuity prediction values and the context emotion continuity tags.

[0068] The loss function is a function used to measure the difference between the model's prediction results and the true labels. The first loss function is used to measure the difference between the pitch prediction features and the pitch feature labels, which can reflect the prediction accuracy of the acoustic feature extraction layer. The second loss function is used to measure the difference between the semantic emotion prediction labels and the semantic emotion labels, reflecting the prediction accuracy of the semantic parsing layer. The third loss function is used to measure the difference between the emotion continuity prediction values and the context emotion continuity labels, reflecting the prediction accuracy of the context association layer.

[0069] For example, if the pitch prediction feature is "high and stable pitch", while the pitch feature label is "low pitch with an upward trend", the difference between the two can be obtained by calculating the first loss function. Similarly, if the semantic emotion prediction label is "happy", while the semantic emotion label is "angry", calculating the second loss function can measure the prediction error of the semantic parsing layer. Similar calculations are also performed for the emotion continuity prediction values and the context emotion continuity labels.

[0070] Step S204: Perform a weighted sum of the first loss function, the second loss function, and the third loss function to generate a joint training loss value, and update the parameters of the initial emotion recognition network through the backpropagation algorithm until convergence.

[0071] Weighted sum means assigning different weights to the first loss function, the second loss function, and the third loss function respectively, and then adding them to obtain the joint training loss value. The setting of the weights can be determined according to the importance of different layers in emotion recognition. The backpropagation algorithm is an algorithm used to update the parameters of a neural network. It adjusts the parameters of each layer in the initial emotion recognition network according to the joint training loss value, so that the model's prediction results gradually approach the true labels until the convergence condition is reached, that is, the joint training loss value no longer decreases significantly.

[0072] For example, assume that the weight of the first loss function is 0.3, the weight of the second loss function is 0.5, and the weight of the third loss function is 0.2. The three calculated loss functions are weighted and summed according to these weights to obtain the joint training loss value. Then, the parameters of the initial emotion recognition network are continuously updated through the backpropagation algorithm. After multiple iterations, when the joint training loss value no longer changes significantly, the model is considered to have converged.

[0073] As an implementation, the above joint training may further include: Step S205: Add interference speech samples containing vehicle environmental noise to the training speech dataset, and set the noise suppression weight corresponding to the interference speech samples.

[0074] An interfering speech sample refers to a speech sample that contains vehicle environmental noise, which may interfere with the training of the emotion recognition network. By adding interfering speech samples to the training speech dataset, the model can learn how to accurately recognize emotions in the presence of noise interference. The noise suppression weight is a weight value set for the interfering speech sample, which is used to control the degree of processing of the interfering speech sample during training. The larger the weight, the more the model focuses on noise suppression.

[0075] For example, originally, the training speech dataset only contained pure speech samples. Now, some speech samples containing interference such as vehicle engine noise and wind noise are added, and a noise suppression weight of 0.6 is set for these interfering speech samples, indicating that during training, attention is paid to the processing of this noise.

[0076] Step S206: When extracting features from the interfering speech sample through the acoustic feature extraction layer, based on the noise suppression weight, reduce the contribution degree of the frequency band features corresponding to the vehicle engine noise in the interfering speech sample.

[0077] When the acoustic feature extraction layer extracts features from the interfering speech sample, it will process the frequency band features corresponding to the vehicle engine noise according to the noise suppression weight. Reducing the contribution degree of these frequency band features can reduce the interference of engine noise on model training and make the model pay more attention to the emotion features in the speech.

[0078] For example, vehicle engine noise is mainly concentrated in a certain frequency band range. When extracting features from the interfering speech sample, according to the set noise suppression weight, reduce the contribution degree of the frequency band features in the entire feature extraction process, so that the model will not produce incorrect emotion recognition results due to the influence of engine noise.

[0079] Step S207: When parsing the interfering speech sample through the semantic parsing layer, based on the noise suppression weight, increase the keyword retrieval intensity related to the driver's speech to cover the semantic ambiguity area caused by noise.

[0080] When the semantic parsing layer parses the interfering speech sample, the presence of noise may cause semantic ambiguity in the speech. By increasing the keyword retrieval intensity related to the driver's speech based on the noise suppression weight, the semantic information in the speech can be more accurately recognized, covering the semantic ambiguity area caused by noise, and improving the accuracy of semantic parsing.

[0081] For example, in the presence of noise interference, some keywords in the speech may not be clearly heard, resulting in semantic ambiguity. By increasing the keyword retrieval intensity, the semantic parsing layer can more actively search for keywords related to the driver's speech and can more accurately understand the semantic content of the speech even in the presence of noise.

[0082] Step S300: Input the multi-level emotion feature set into the safety risk prediction model to generate a safety behavior score curve for the driver within a preset time period. The safety behavior score curve is used to characterize the correlation between the driver's emotional stability and driving operation risk at different time nodes.

[0083] The safety risk prediction model is a model used to predict the driving safety risk of a driver. It can analyze the driver's emotional state based on the input multi-level emotion feature set and predict the driving operation risk at different time nodes. The safety behavior score curve is a curve generated based on the prediction results of the model, which reflects the correlation between the driver's emotional stability and driving operation risk at different time nodes within the preset time period. The height of the curve can indicate the magnitude of the driving operation risk. The more unstable the emotion, the higher the curve may be, and the corresponding driving operation risk is also greater.

[0084] For example, input the previously generated multi-level emotion feature set into the safety risk prediction model. The model analyzes the driver's emotional changes within the preset time period based on these features and then generates a safety behavior score curve. If at a certain time node, the driver's emotion is relatively excited and the emotional stability is poor, then the score value corresponding to this time node on the safety behavior score curve may be higher, indicating a greater driving operation risk.

[0085] As an implementation manner, step S300, inputting the multi-level emotion feature set into the safety risk prediction model to generate a safety behavior score curve for the driver within a preset time period, can be implemented as the following steps: Step S310: Through the time series analysis module in the safety risk prediction model, perform periodic decomposition on the real-time emotion fluctuation trajectory to extract the emotion mutation frequency and mutation amplitude of the driver within a preset time window.

[0086] The time series analysis module is the part of the safety risk prediction model used to analyze time series data. The real-time emotion fluctuation trajectory is the sequence data reflecting the driver's emotion changing over time. Periodic decomposition is to decompose the real-time emotion fluctuation trajectory into different periodic components to better analyze the change law of the emotion. The emotion mutation frequency refers to the number of times the driver's emotion mutates within a preset time window, and the emotion mutation amplitude refers to the change magnitude of the emotion intensity each time the emotion mutates.

[0087] For example, within a time period with a preset time window of 10 minutes, through the time series analysis module to perform periodic decomposition on the real-time emotional fluctuation trajectory, it is found that the driver's emotion has undergone 3 obvious mutations within these 10 minutes. The change amplitudes of the emotion intensity at each mutation are 20%, 30%, and 25% respectively. Then the emotion mutation frequency is 3 times / 10 minutes, and the mutation amplitudes are 20%, 30%, and 25% respectively.

[0088] Step S320: Through the risk association mapping module in the safety risk prediction model, match the semantic emotion intensity parameter with the predefined driving operation risk levels to determine the vehicle control deviation probabilities corresponding to different emotion intensities.

[0089] The risk association mapping module is the part in the safety risk prediction model used to establish the association relationship between the semantic emotion intensity parameter and the driving operation risk levels. The predefined driving operation risk levels are different risk levels preset based on a large number of experiments and data analyses, such as low risk, medium risk, high risk, etc. The vehicle control deviation probability refers to the possibility of the driver having a deviation in the vehicle control operation under different emotion intensities.

[0090] For example, the predefined driving operation risk levels are divided into three levels: low risk, medium risk, and high risk, and the semantic emotion intensity parameters have three situations: weak, medium, and strong. The risk association mapping module matches the semantic emotion intensity parameter with the driving operation risk levels and finds that when the semantic emotion intensity is weak, the corresponding vehicle control deviation probability is 10%, belonging to the low risk level; when the semantic emotion intensity is medium, the corresponding vehicle control deviation probability is 30%, belonging to the medium risk level; when the semantic emotion intensity is strong, the corresponding vehicle control deviation probability is 60%, belonging to the high risk level.

[0091] As an implementation manner, step S320, through the risk association mapping module in the safety risk prediction model, match the semantic emotion intensity parameter with the predefined driving operation risk levels to determine the vehicle control deviation probabilities corresponding to different emotion intensities, can be implemented as the following steps: Step S321: Based on the emotion type label and intensity quantization value included in the semantic emotion intensity parameter, perform emotion intensity interval division on the driver's emotion expression to generate an emotion intensity interval division result including multiple consecutive emotion intensity levels.

[0092] The emotion type label is a description of the driver's emotion type, such as anger, anxiety, calm, etc. The intensity quantization value is a specific numerical representation of the emotion intensity. For example, a numerical value from 0 - 100 can be used to represent the emotion intensity. The emotion intensity interval division is to divide the driver's emotion expression into different intensity intervals according to the emotion type label and intensity quantization value, forming multiple consecutive emotion intensity levels.

[0093] For example, the emotion type label included in the semantic emotion intensity parameter is "anger", and the intensity quantization value is 30 - 100. The anger emotion can be divided into three intensity levels: mild anger (intensity quantization value 30 - 50), moderate anger (intensity quantization value 51 - 70), and severe anger (intensity quantization value 71 - 100), obtaining the emotion intensity interval division result.

[0094] Step S322: According to the emotion intensity threshold range corresponding to each risk level in the predefined driving operation risk level, layer by layer match each emotion intensity level in the emotion intensity interval division result with the risk level to generate an emotion intensity - risk level mapping relationship, where the risk level includes the probability of accidental triggering of emergency braking caused by anger emotion, the lane departure frequency caused by anxiety emotion, and the throttle response delay duration caused by fatigue emotion.

[0095] Each risk level in the predefined driving operation risk level corresponds to an emotion intensity threshold range. For example, the emotion intensity threshold range corresponding to the low risk level may be 0 - 30, the emotion intensity threshold range corresponding to the medium risk level may be 31 - 60, and the emotion intensity threshold range corresponding to the high risk level may be 61 - 100. By layer by layer matching each emotion intensity level in the emotion intensity interval division result with these risk levels, a mapping relationship between emotion intensity and risk level can be established.

[0096] For example, for the three intensity levels of the anger emotion divided previously, the risk level corresponding to mild anger (intensity quantization value 30 - 50) may be low risk, the risk level corresponding to moderate anger (intensity quantization value 51 - 70) may be medium risk, and the risk level corresponding to severe anger (intensity quantization value 71 - 100) may be high risk. At the same time, different risk levels correspond to different specific risk indicators, such as the probability of accidental triggering of emergency braking caused by anger emotion, the lane departure frequency caused by anxiety emotion, the throttle response delay duration caused by fatigue emotion, etc.

[0097] Step S323: Based on the emotion intensity - risk level mapping relationship, extract the occurrence frequency of vehicle control deviation events in the historical driving operation dataset corresponding to each emotion intensity level, and generate an initial vehicle control deviation probability according to the occurrence frequency.

[0098] The historical driving operation dataset is a collection containing a large amount of drivers' historical driving operation data, which records the occurrence of vehicle control deviation events in different emotional states. According to the mapping relationship between emotional intensity and risk level, the occurrence frequency of vehicle control deviation events corresponding to each emotional intensity level is extracted from the historical driving operation dataset. For example, in a state of mild anger, the occurrence frequency of vehicle control deviation events is 10%, in a state of moderate anger, the occurrence frequency is 30%, and in a state of severe anger, the occurrence frequency is 60%. Then, the initial vehicle control deviation probability is generated based on these occurrence frequencies.

[0099] For example, for the state of mild anger, based on the corresponding occurrence frequency of vehicle control deviation events of 10%, the initial vehicle control deviation probability is generated as 10%; for the state of moderate anger, based on the occurrence frequency of 30%, the initial vehicle control deviation probability is generated as 30%; for the state of severe anger, based on the occurrence frequency of 60%, the initial vehicle control deviation probability is generated as 60%.

[0100] Step S324: Dynamically correct the initial vehicle control deviation probability. Among them, the dynamic environment correction includes combining the road complexity parameter and traffic flow density parameter in the current driving environment, adjusting the weight distribution of the initial vehicle control deviation probability in different environments, and generating the vehicle control deviation probability after environment-adaptive correction.

[0101] The dynamic environment correction is a process of adjusting the initial vehicle control deviation probability considering the influence of the current driving environment on the vehicle control deviation probability. The road complexity parameter is a parameter describing the complexity of the road conditions, such as the number of road curves, slope changes, etc. The traffic flow density parameter refers to the number of vehicles passing through a certain road section per unit time. By combining these two parameters, the weight distribution of the initial vehicle control deviation probability in different environments is adjusted, making the vehicle control deviation probability more in line with the actual driving environment.

[0102] For example, in an environment with high road complexity and high traffic flow density, even if the driver's emotional intensity is in a state of mild anger and the initial vehicle control deviation probability is 10%, due to the influence of the environment, this probability may need to be adjusted to 20% to generate the vehicle control deviation probability after environment-adaptive correction.

[0103] Step S325: Correlate and verify the vehicle control deviation probability after environment-adaptive correction with the real-time collected vehicle operation state data. The correlation verification includes detecting whether the number of sudden changes in the steering wheel rotation angle and the abnormal fluctuations in the braking pedal depression force are consistent with the predicted trend of the vehicle control deviation probability, and calibrating the confidence level of the vehicle control deviation probability according to the verification result.

[0104] The associated verification is a process of comparing and verifying the vehicle control deviation probability after environment - adaptive correction with the vehicle operation state data collected in real - time. The number of sudden changes in the steering wheel angle and the abnormal fluctuations in the braking pedal force are important indicators reflecting whether the vehicle operation state is abnormal. By detecting whether these indicators are consistent with the predicted trend of the vehicle control deviation probability, the accuracy of the vehicle control deviation probability can be evaluated. The confidence calibration is to adjust the vehicle control deviation probability according to the verification results to make it more reliable.

[0105] For example, the predicted vehicle control deviation probability after environment - adaptive correction indicates a high probability of vehicle control deviation within a certain time period. The real - time collected vehicle operation state data shows an increase in the number of sudden changes in the steering wheel angle and abnormal fluctuations in the braking pedal force, indicating that the predicted trend is consistent with the actual situation, and the confidence level of the vehicle control deviation probability is high, and no major adjustment is required; if the predicted trend does not match the actual situation, the vehicle control deviation probability needs to be calibrated accordingly.

[0106] Step S326: Based on the vehicle control deviation probability after confidence calibration, generate the final vehicle control deviation probability corresponding to each emotion intensity level, and associate the final vehicle control deviation probability with each semantic emotion intensity parameter within a preset time period in chronological order to form a probability input sequence for driving the generation of the safety behavior scoring curve.

[0107] The final vehicle control deviation probability is the vehicle control deviation probability after confidence calibration, which more accurately reflects the probability of vehicle control deviation at different emotion intensity levels. Associating the final vehicle control deviation probability with each semantic emotion intensity parameter within a preset time period in chronological order forms a probability input sequence, which can be used as the input data for the safety risk prediction model to generate the safety behavior scoring curve.

[0108] For example, after confidence calibration, the final vehicle control deviation probability corresponding to mild anger is 15%, the final vehicle control deviation probability corresponding to moderate anger is 35%, and the final vehicle control deviation probability corresponding to severe anger is 65%. Associating these probabilities with the semantic emotion intensity parameters within a preset time period in chronological order forms a probability input sequence for driving the generation of the safety behavior scoring curve.

[0109] Step S330: Through the dynamic scoring generation module in the safety risk prediction model, calculate the safety score values of the driver at multiple consecutive time points by combining the emotion mutation frequency, mutation amplitude, and vehicle control deviation probability.

[0110] The dynamic scoring generation module is the part of the safety risk prediction model used to calculate the driver's safety score value. The frequency of emotional mutation, the amplitude of mutation, and the probability of vehicle control deviation are important factors affecting the driver's safety score value. By comprehensively considering these factors, the safety score values of the driver at multiple consecutive time points can be calculated, and these score values reflect the driving safety level of the driver at different time points.

[0111] For example, at a certain time point, the frequency of the driver's emotional mutation is 2 times per minute, the amplitude of mutation is 25%, and the probability of vehicle control deviation is 30%. The dynamic scoring generation module calculates the safety score value at this time point as 70 points according to the preset algorithm, combining these factors.

[0112] As an implementation method, step S330, through the dynamic scoring generation module in the safety risk prediction model, combining the frequency of emotional mutation, the amplitude of mutation, and the probability of vehicle control deviation, calculating the safety score values of the driver at multiple consecutive time points, can be implemented as the following steps: Step S331: Determine the emotional fluctuation time interval between adjacent time nodes based on the frequency of emotional mutation, and generate a first influence factor negatively correlated with driving concentration based on the emotional fluctuation time interval.

[0113] The emotional fluctuation time interval refers to the time length of the emotional fluctuation of the driver between adjacent time nodes. The higher the frequency of emotional mutation, the shorter the emotional fluctuation time interval. The first influence factor is a factor negatively correlated with driving concentration, that is, the shorter the emotional fluctuation time interval, the larger the value of the first influence factor, indicating that the driving concentration of the driver is lower.

[0114] For example, if the frequency of emotional mutation is 3 times per minute, then the emotional fluctuation time interval between adjacent time nodes is about 20 seconds. According to the preset rules, when the emotional fluctuation time interval is 20 seconds, the generated first influence factor is 0.8, indicating that the driving concentration of the driver is lower.

[0115] Step S332: Extract the emotional intensity difference value of the driver within the preset time window according to the amplitude of mutation, and compare the emotional intensity difference value with the predefined emotional stability threshold to generate a second influence factor reflecting the amplitude of emotional fluctuation.

[0116] The emotional intensity difference value refers to the change magnitude of the emotional intensity when the driver's emotion mutates within the preset time window. The predefined emotional stability threshold is a preset emotional intensity difference limit used to judge whether the driver's emotion is stable. By comparing the emotional intensity difference value with the emotional stability threshold, a second influence factor reflecting the amplitude of emotional fluctuation can be generated. The larger the emotional intensity difference value, the larger the value of the second influence factor, indicating that the amplitude of emotional fluctuation is larger.

[0117] For example, within a preset time window of 5 minutes, the sudden change amplitude of the driver's emotion is 30%, and the predefined emotion stability threshold is 20%. Since the difference value of emotion intensity 30% is greater than the threshold 20%, the generated second influence factor is 0.9, indicating a relatively large amplitude of emotion fluctuation.

[0118] Step S333: Identify the number of vehicle control delay events caused by the driver's emotion fluctuation at historical time nodes based on the vehicle control deviation probability, and generate a third influence factor representing the cumulative degree of operation errors according to the number.

[0119] A vehicle control delay event refers to an event where the driver has a delay in vehicle control operations when experiencing emotion fluctuation. The number of vehicle control delay events caused by emotion fluctuation can be identified from historical driving data through the vehicle control deviation probability. The third influence factor is generated based on the number of these events and is used to represent the cumulative degree of operation errors. The larger the number of events, the larger the value of the third influence factor.

[0120] For example, according to the vehicle control deviation probability, it is identified from historical driving data that within the past 10 minutes, there are 5 vehicle control delay events caused by the driver's emotion fluctuation. According to the preset rules, the generated third influence factor is 0.7, indicating a relatively high cumulative degree of operation errors.

[0121] Step S334: Perform weight assignment adjustment on the first influence factor, the second influence factor, and the third influence factor to obtain a dynamic weight assignment result, where the weight assignment adjustment is dynamically corrected based on the deviation degree between the driver's emotion mutation pattern and the actual driving trajectory of the vehicle within the preset time window.

[0122] The weight assignment adjustment is to assign different weights to the first influence factor, the second influence factor, and the third influence factor respectively to comprehensively consider their influence on the safety score value. Dynamic correction means that the weight assignment is not fixed but is adjusted according to the deviation degree between the driver's emotion mutation pattern and the actual driving trajectory of the vehicle within the preset time window. If the deviation degree between the emotion mutation pattern and the actual driving trajectory of the vehicle is relatively large, it indicates that the emotion has a relatively large impact on driving operations, and it may be necessary to increase the weights of the first influence factor and the second influence factor.

[0123] For example, within a preset time window of 15 minutes, it is found that the deviation degree between the driver's emotion mutation pattern and the actual driving trajectory of the vehicle is relatively large. Originally, the weight of the first influence factor is 0.3, the weight of the second influence factor is 0.3, and the weight of the third influence factor is 0.4. After dynamic correction, the weight of the first influence factor is adjusted to 0.4, the weight of the second influence factor is adjusted to 0.4, and the weight of the third influence factor is adjusted to 0.2, obtaining a dynamic weight assignment result.

[0124] Step S335: Weightedly fuse the first influencing factor, the second influencing factor, and the third influencing factor according to the dynamic weight allocation result to generate an initial safety score sequence, and superimpose a time decay factor on each score value in the initial safety score sequence to obtain a smoothed intermediate safety score sequence. The time decay factor is used to reduce the influence weight of the score value at the historical time node on the current time node.

[0125] Weighted fusion is to perform weighted summation on the first influencing factor, the second influencing factor, and the third influencing factor according to the dynamic weight allocation result to generate an initial safety score sequence. The time decay factor is a factor that gradually decreases over time, and it is used to reduce the influence weight of the score value at the historical time node on the current time node, making the safety score sequence pay more attention to the current driving state.

[0126] For example, the first influencing factor is 0.8, and the weight is 0.4; the second influencing factor is 0.9, and the weight is 0.4; the third influencing factor is 0.7, and the weight is 0.2. After weighted fusion, the initial safety score value is 0.8×0.4 + 0.9×0.4 + 0.7×0.2 = 0.82. Then, superimpose the time decay factor on this initial safety score value. Assuming the time decay factor is 0.9, the smoothed intermediate safety score value is 0.82×0.9 = 0.738.

[0127] Step S336: Correlate and calibrate the intermediate safety score sequence with the driver's real-time physiological index data to generate a calibrated safety score value. The physiological index data includes the heart rate change trend and the handgrip stability parameter. The correlation and calibration are used to correct the scoring deviation caused by the voice data acquisition error.

[0128] The real-time physiological index data is an index reflecting the driver's current physiological state, such as the heart rate change trend and the handgrip stability parameter. The correlation and calibration are to compare and adjust the intermediate safety score sequence with these physiological index data to correct the scoring deviation caused by the voice data acquisition error. If the physiological index data shows that the driver's state is inconsistent with the intermediate safety score sequence, it may be necessary to calibrate the safety score value accordingly.

[0129] For example, the intermediate safety score sequence shows that the driver's safety score value is 70 points, but the real-time physiological index data shows that the driver's heart rate has increased significantly and the handgrip is unstable, indicating that the driver's actual state may be more dangerous than that shown in the score sequence. Through correlation and calibration, the safety score value is adjusted to 60 points to generate a calibrated safety score value.

[0130] Step S337: Map the calibrated safety score values to each time node within a preset time period in chronological order to generate a safety behavior score curve, and perform a difference calculation on the score values of adjacent time nodes in the safety behavior score curve to generate a score gradient parameter for indicating the change rate of driving risk to assist in the identification of risk nodes.

[0131] By mapping the calibrated safety score values to each time node within a preset time period in chronological order, a safety behavior score curve can be obtained. The score gradient parameter is obtained by performing a difference calculation on the score values of adjacent time nodes in the safety behavior score curve, and it can indicate the change rate of driving risk. The larger the score gradient parameter, the faster the change in driving risk, and the corresponding time node may be a risk node.

[0132] For example, the calibrated safety score values at time nodes within a preset time period are 70 points, 65 points, and 60 points in sequence. Through difference calculation, the score gradient parameters of adjacent time nodes are -5 points per time unit and -5 points per time unit respectively. These score gradient parameters can help identify the time nodes with rapid changes in driving risk and assist in determining risk nodes.

[0133] Step S340: Connect the safety score values in chronological order to generate a safety behavior score curve, and perform smoothing processing on the safety behavior score curve to eliminate short-term noise interference.

[0134] In this step, the safety score values at multiple consecutive time points calculated previously are connected in chronological order to form a safety behavior score curve. The smoothing process is to remove short-term noise interference in the safety behavior score curve, make the curve smoother, and better reflect the true correlation between the driver's emotional stability and driving operation risk.

[0135] For example, through smoothing processing methods such as moving average, the safety behavior score curve is processed to remove the fluctuations in score values caused by accidental factors, make the curve smoother, and facilitate subsequent analysis and judgment.

[0136] As an implementation method, the safety risk prediction model is obtained through joint training with multi-dimensional driving data, and the multi-dimensional driving data includes historical driving voice data, vehicle operation log data, and actual accident record data; among them, the joint training of the multi-dimensional driving data includes: Step S301: Input the historical driving voice data into an emotion recognition network to generate a set of historical emotion features, and convert the steering wheel rotation angle, braking frequency, and lane deviation in the vehicle operation log data into operation risk indicators.

[0137] Historical driving voice data is the voice data collected during the driver's past driving. Inputting it into the emotion recognition network can generate a set of historical emotion features, which reflect the driver's emotional state during the historical driving. The vehicle operation log data records various operation information during the vehicle's driving, such as the steering angle of the steering wheel, the braking frequency, and the lane deviation amount, etc. Convert these information into operation risk indicators. For example, the magnitude of the operation risk can be evaluated according to the change range of the steering angle of the steering wheel, the level of the braking frequency, and the size of the lane deviation amount.

[0138] For example, input the historical driving voice data of the past month into the emotion recognition network to obtain a set of historical emotion features. At the same time, extract the steering angle of the steering wheel, the braking frequency, and the lane deviation amount from the vehicle operation log data. Convert the situations where the change range of the steering angle of the steering wheel exceeds the preset threshold, the braking frequency is too high, and the lane deviation amount is too large into high operation risk indicators respectively.

[0139] Step S302: Through the feature fusion module in the safety risk prediction model, perform time alignment and spatial mapping on the set of historical emotion features and the operation risk indicators to generate a joint input feature vector.

[0140] The feature fusion module is the part in the safety risk prediction model used to fuse different types of features. Time alignment is to correspond the set of historical emotion features and the operation risk indicators in time, so that the data at the same time point can be correlated with each other. Spatial mapping is to map different types of features into the same feature space for subsequent processing. Through time alignment and spatial mapping, the set of historical emotion features and the operation risk indicators are fused together to generate a joint input feature vector.

[0141] For example, the set of historical emotion features records the emotion intensity at each time point, and the operation risk indicators record the magnitude of the operation risk at each time point. Align them in time so that the emotion intensity and the magnitude of the operation risk at the same time point can correspond. Then through spatial mapping, map the emotion intensity and the magnitude of the operation risk into a two-dimensional feature space to generate a joint input feature vector.

[0142] Step S303: Through the risk prediction module in the safety risk prediction model, predict the accident occurrence probability within the future time window based on the joint input feature vector, and calculate the prediction error loss according to the real accident time points in the actual accident record data.

[0143] The risk prediction module is the part of the safety risk prediction model that predicts the probability of an accident. Based on the combined input feature vector, it uses a preset algorithm and model to predict the probability of an accident within a future time window. The actual accident record data records the real time points of past accidents. By comparing the predicted accident probability with the actual accident record data, the prediction error loss is calculated to evaluate the prediction accuracy of the model.

[0144] For example, based on the combined input feature vector, the risk prediction module predicts that the probability of an accident within the next 10 minutes is 20%. However, the actual accident record data shows that no accident occurred within the next 10 minutes. Then, the prediction error loss can be calculated to measure the prediction deviation of the model.

[0145] Step S304: Iteratively optimize the safety risk prediction model based on the prediction error loss until the prediction accuracy of the accident probability reaches the preset convergence condition.

[0146] Iterative optimization is a process of continuously adjusting the parameters of the safety risk prediction model according to the prediction error loss, so as to continuously improve the prediction accuracy of the model. The preset convergence condition is a preset accuracy limit. When the prediction accuracy of the accident probability reaches this limit, the model is considered to have converged and the iterative optimization process ends.

[0147] For example, by continuously adjusting the parameters of the safety risk prediction model, the prediction error loss is gradually reduced. When the prediction accuracy of the accident probability reaches the preset convergence condition, such as the prediction error loss is less than 0.05, stop the iterative optimization. At this time, the prediction accuracy of the model meets the requirements.

[0148] As an implementation, the above multi-dimensional driving data joint training may further include: Step S305: Introduce an attention weight allocation mechanism in the feature fusion module to adjust the contribution ratio of the historical emotion feature set and the manipulation risk index in the combined input feature vector; The attention weight allocation mechanism is a mechanism for adjusting the contribution ratio of different features in the combined input feature vector. By introducing this mechanism, different weights can be assigned to the historical emotion feature set and the manipulation risk index according to the importance of different features for predicting the accident probability, so that the combined input feature vector can more reasonably reflect the influence of various factors on driving safety.

[0149] For example, if it is found that the historical emotion feature set has a greater impact on the prediction of the accident probability, then the attention weight allocation mechanism can be used to increase the weight of the historical emotion feature set in the combined input feature vector and reduce the weight of the manipulation risk index.

[0150] Among them, the attention weight allocation mechanism includes: Step S3051: Determine the subset of emotional features in the historical emotional feature set that overlaps with the manipulation risk indicator in terms of time through time correlation analysis; Time correlation analysis is a method used to analyze the correlation between the historical emotional feature set and the manipulation risk indicator in terms of time. Through this analysis, the overlapping part of the historical emotional feature set and the manipulation risk indicator in terms of time can be found, forming a subset of emotional features. This subset contains emotional features that change synchronously with the manipulation risk indicator in terms of time and may be more meaningful for predicting the probability of an accident.

[0151] For example, the historical emotional feature set records the daily emotional intensity changes in the past month, and the manipulation risk indicator records the daily manipulation risk levels during the same period. Through time correlation analysis, it is found that within one week, the change trends of the emotional intensity and the manipulation risk level are relatively consistent. Then, the emotional features corresponding to this week form the subset of emotional features.

[0152] Step S3052: Calculate the covariance matrix between the subset of emotional features and the manipulation risk indicator, and generate the emotional feature weights and manipulation indicator weights based on the eigenvalue distribution of the covariance matrix; The covariance matrix is a matrix used to measure the covariance relationship between two variables. By calculating the covariance matrix between the subset of emotional features and the manipulation risk indicator, the correlation between them can be understood. The eigenvalue distribution of the covariance matrix reflects the contribution degree of each variable to the overall change. Based on the eigenvalue distribution, the emotional feature weights and manipulation indicator weights can be generated. The larger the weight, the greater the contribution of the feature to the combined input feature vector.

[0153] For example, the covariance matrix between the subset of emotional features and the manipulation risk indicator is calculated. By performing eigenvalue decomposition on the covariance matrix, the eigenvalue distribution is obtained. According to the eigenvalue distribution, the weight of the subset of emotional features is assigned as 0.6, and the weight of the manipulation risk indicator is assigned as 0.4.

[0154] Step S3053: Perform weighted fusion on the subset of emotional features and the manipulation risk indicator according to the emotional feature weights and manipulation indicator weights to generate a combined input feature vector.

[0155] Weighted fusion is a process of weighted summation of the subset of emotional features and the manipulation risk indicator according to the emotional feature weights and manipulation indicator weights. Through weighted fusion, different features can be combined in an appropriate proportion to generate a combined input feature vector, making the combined input feature vector more accurately reflect the influence of various factors on driving safety.

[0156] For example, the value of the emotional feature subset is [0.2, 0.3, 0.4], and the weight is 0.6; the value of the manipulation risk index is [0.1, 0.2, 0.3], and the weight is 0.4. After weighted fusion, the generated joint input feature vector is [0.2×0.6 + 0.1×0.4, 0.3×0.6 + 0.2×0.4, 0.4×0.6 + 0.3×0.4] = [0.16, 0.26, 0.36].

[0157] As an implementation, after generating the driving intervention prompt information in step S3053, the method provided by the embodiments of the present invention may further include: Step S3054: Monitor in real time the response operation of the driver to the driving intervention prompt information, and collect the updated driving voice data stream after the response operation; Real-time monitoring refers to continuously observing and recording the operations of the driver. After the driving intervention prompt information is generated, the driver may make corresponding response operations, such as adjusting driving behavior, replying to voice commands, etc. By monitoring these response operations in real time and collecting the updated driving voice data stream after the response operation, the reaction of the driver to the prompt information can be understood, providing data support for subsequent analysis and processing.

[0158] For example, the driving intervention prompt information prompts the driver to reduce the speed. Monitor in real time whether the driver steps on the brake pedal to reduce the speed, and collect the voice emitted by the driver after making the response operation to form an updated driving voice data stream.

[0159] Step S3055: Extract emotional features from the updated driving voice data stream to generate an updated safety behavior scoring curve; The process of extracting emotional features from the updated driving voice data stream is similar to the process of extracting emotional features from the driving voice data stream described above. By extracting the emotional features in the updated driving voice data stream, a multi-level emotional feature set is generated, and then it is input into the safety risk prediction model to generate an updated safety behavior scoring curve. The updated safety behavior scoring curve can reflect the changes in the emotional state of the driver and the driving safety risk after responding to the driving intervention prompt information.

[0160] For example, process the updated driving voice data stream according to the previous steps, extract emotional features, input them into the safety risk prediction model, and obtain an updated safety behavior scoring curve. If the driver makes a positive response to the prompt information and the emotion becomes more stable, then the updated safety behavior scoring curve may drop, indicating a reduction in driving safety risk.

[0161] Step S3056: If the updated safety behavior score curve does not return to within the safety threshold range within the preset time, trigger the active intervention protocol of the vehicle control system. The active intervention protocol includes at least one of automatically reducing the driving speed, activating the emergency braking preparation mode, and sending an assistance request to the traffic management center.

[0162] The safety threshold range is a pre-set range of safety score values. When the safety behavior score curve is within this range, it indicates that the driving safety risk is at an acceptable level. If the updated safety behavior score curve fails to return to within the safety threshold range within the preset time, it means that the driver's emotional state and driving safety risk are still relatively high, and it is necessary to trigger the active intervention protocol of the vehicle control system. The active intervention protocol includes various measures. For example, automatically reducing the driving speed can reduce the likelihood of accidents, activating the emergency braking preparation mode can respond more quickly in an emergency, and sending an assistance request to the traffic management center can obtain external support and help.

[0163] For example, the preset time is 5 minutes, and the safety threshold range is 70 - 100 points. If the updated safety behavior score curve remains below 70 points within 5 minutes, then trigger the active intervention protocol of the vehicle control system, automatically reduce the driving speed, and send an assistance request to the traffic management center.

[0164] Step S400: Generate driving intervention prompt information corresponding to the risk nodes exceeding the preset threshold in the safety behavior score curve, and send the driving intervention prompt information to the in-vehicle terminal of the target vehicle for real-time display.

[0165] The preset threshold is a pre-set safety score limit. Safety score nodes exceeding this limit are considered risk nodes. The driving intervention prompt information is generated based on the situation of the risk nodes and is used to remind the driver to pay attention to driving safety and take corresponding measures. Sending the driving intervention prompt information to the in-vehicle terminal of the target vehicle for real-time display allows the driver to timely understand their driving risk situation and make corresponding adjustments.

[0166] For example, the preset threshold is 80 points, and there is a node in the safety behavior score curve with a score value of 90 points, which exceeds the preset threshold. This node is a risk node. According to the emotional state and driving operation risk situation corresponding to this risk node, generate driving intervention prompt information such as "You are currently in a relatively excited mood. Please stay calm and pay attention to safe driving", and send this prompt information to the display screen of the in-vehicle terminal for real-time display.

[0167] As an implementation method, step S400, generating driving intervention prompt information corresponding to the risk nodes exceeding the preset threshold in the safety behavior score curve, can be implemented as the following steps: Step S410: Perform gradient detection on the safety behavior score curve, identify consecutive time intervals during which the decline rate of the safety score value exceeds the rate threshold within a preset time interval, and mark the consecutive time intervals as high-risk periods; Gradient detection is a process of calculating and analyzing the slope of the safety behavior score curve. Through gradient detection, the decline rate of the safety score value within a preset time interval can be identified. The rate threshold is a preset decline rate limit. When the decline rate of the safety score value exceeds this limit, the corresponding consecutive time interval is marked as a high-risk period. A high-risk period indicates that the driving safety risk of the driver increases sharply during this time period and requires key attention.

[0168] For example, the preset time interval is 10 minutes and the rate threshold is set to a decline of 5 points per minute. When performing gradient detection on the safety behavior score curve, it is found that within the consecutive time interval from the 3rd minute to the 6th minute, the safety score value drops from 85 points to 70 points, and the decline rate reaches more than 5 points per minute. Then, the consecutive time interval from the 3rd minute to the 6th minute is marked as a high-risk period. This means that during this time period, the driver's mood or driving operations may have fluctuated greatly, resulting in a significant increase in driving safety risk.

[0169] Step S420: Extract the dominant emotion type from the multi-level emotion feature set corresponding to the high-risk period, and select a matching prompt content template from the preset intervention strategy library based on the dominant emotion type; The dominant emotion type refers to the emotion type that occupies a dominant position in the multi-level emotion feature set corresponding to the high-risk period, and it can reflect the driver's main emotional state during this period. The preset intervention strategy library is a pre-established database containing various prompt content templates for different emotion types. Selecting a matching prompt content template from the intervention strategy library according to the dominant emotion type can make the generated driving intervention prompt information more targeted and effective.

[0170] For example, after analyzing the multi-level emotion feature set corresponding to the high-risk period, it is found that the dominant emotion type is "anger". In the preset intervention strategy library, there may be such a prompt content template for the "anger" emotion: "You seem to be a bit angry now. Please calm down first to avoid affecting driving safety due to emotional excitement." By selecting this matching prompt content template, it can directly guide and remind the driver of the current anger emotion.

[0171] Step S430: Generate supplementary warning information in the prompt content template according to the vehicle control deviation probability and the current vehicle driving parameters. The supplementary warning information includes recommended adjustment parameters related to the steering wheel angle and the depth of the accelerator pedal; The vehicle control deviation probability reflects the likelihood of a driver making a deviation during vehicle control operations in the current emotional state, and the current vehicle driving parameters include information such as steering wheel angle, accelerator pedal depth, vehicle speed, etc. Generating supplementary warning information in the prompt content template based on this information can provide drivers with more specific operation suggestions, helping them adjust their driving behavior and reduce driving risks.

[0172] For example, the vehicle control deviation probability shows that there is a high likelihood of a steering wheel angle deviation due to anger, and the current vehicle driving parameters show that the steering wheel angle is too large. At this time, the supplementary warning information can be "Your steering wheel angle is too large. Please appropriately reduce the angle to keep the vehicle driving smoothly." Or if the accelerator pedal depth is too large and the vehicle control deviation probability indicates that there may be a risk of excessive acceleration due to this situation, the supplementary warning information can be "You are pressing the accelerator pedal too deeply. Please relax appropriately to control the vehicle speed." Step S440: Combine the prompt content template with the supplementary warning information to generate a driving intervention prompt message, and sort the driving intervention prompt messages according to a preset priority and send them to different display areas of the in-vehicle terminal.

[0173] Combining the prompt content template with the supplementary warning information forms a complete driving intervention prompt message. The preset priority is the order preset according to the importance of the prompt message. Different display areas can be different positions on the in-vehicle terminal display screen, such as the main screen, the secondary screen, the head-up display area, etc. Sorting and sending to different display areas of the in-vehicle terminal according to the preset priority can ensure that the driver can see the most important prompt message in a timely and clear manner.

[0174] For example, the prompt content template is "You seem to be a bit angry now. Please calm down first to avoid affecting driving safety due to emotional excitement." The supplementary warning information is "Your steering wheel angle is too large. Please appropriately reduce the angle to keep the vehicle driving smoothly." After combining them, the driving intervention prompt message is "You seem to be a bit angry now. Please calm down first to avoid affecting driving safety due to emotional excitement. Your steering wheel angle is too large. Please appropriately reduce the angle to keep the vehicle driving smoothly." Assuming that the preset priority of this prompt message is relatively high, it is sent to the main screen of the in-vehicle terminal for display so that the driver can see and take corresponding measures in the first place.

[0175] Please refer to Figure 2 , Figure 2Schematic structural diagram of a driving safety analysis device provided by an embodiment of the present invention. The driving safety analysis device is, for example, an in-vehicle terminal on a vehicle or a server communicatively connected to the in-vehicle terminal. The driving safety analysis device at least includes a processor 101, a communication interface 102, and a memory 103. Among them, the processor 101, the communication interface 102, and the memory 103 can be connected through a bus or other means. Among them, the processor 101 (or central processing unit (CPU)) is the computing core and control core of the driving safety analysis device, which can parse various instructions in the driving safety analysis device and process various data of the driving safety analysis device. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for the transmission and interaction of internal data of the driving safety analysis device. The memory 103 (Memory) is a memory device in the driving safety analysis device, used to store programs and data. It can be understood that the memory 103 here can include both the built-in memory of the driving safety analysis device and, of course, the extended memory supported by the driving safety analysis device. The memory 103 provides a storage space, and the storage space stores the operating system of the driving safety analysis device, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present invention does not limit this.

[0176] In one embodiment, the processor 101 executes the driving safety analysis method based on voice emotion recognition provided above in the embodiments of the present invention by running a computer program in the memory 103.

Claims

1. A driving safety analysis method based on speech emotion recognition, characterized in that, The method includes: Obtaining a driving voice data stream collected by a target vehicle in a driving environment, where the driving voice data stream includes continuous voice segments of a driver within a preset time period and interaction instructions related to vehicle control; Performing emotion feature extraction on the driving voice data stream to generate a multi-level emotion feature set, where the multi-level emotion feature set includes a real-time emotion fluctuation trajectory corresponding to the continuous voice segments and a semantic emotion intensity parameter corresponding to the interaction instructions; Inputting the multi-level emotion feature set into a safety risk prediction model to generate a safety behavior score curve of the driver within the preset time period, where the safety behavior score curve is used to characterize the correlation between the emotion stability of the driver and the driving operation risk at different time nodes; Generating driving intervention prompt information corresponding to the risk nodes according to the risk nodes exceeding a preset threshold in the safety behavior score curve, and sending the driving intervention prompt information to an in-vehicle terminal of the target vehicle for real-time display.

2. The method according to claim 1, wherein The obtaining of the driving voice data stream collected by the target vehicle in the driving environment includes: Performing segmentation processing on the driving voice data stream to obtain a plurality of voice data units of equal duration, and performing environmental noise filtering on each voice data unit to generate a plurality of denoised voice segments; where the segmentation processing includes endpoint detection based on voice energy mutation and context division based on semantic coherence; the environmental noise filtering includes eliminating vehicle engine noise through frequency domain filtering and separating the driver's voice from the passenger's voice through voiceprint feature matching; Splicing the plurality of denoised voice segments in chronological order to generate a standardized voice data sequence corresponding to the driving voice data stream.

3. The method according to claim 2, characterized in that, The performing of emotion feature extraction on the driving voice data stream to generate a multi-level emotion feature set includes: Inputting the standardized voice data sequence into a pre-trained emotion recognition network, and extracting the pitch change pattern, speech rate fluctuation frequency, and speech amplitude distribution parameters of each denoised voice segment through an acoustic feature extraction layer in the emotion recognition network; Identifying instruction keywords, emotional modifiers, and modal particles included in the denoised voice segments through a semantic parsing layer in the emotion recognition network, and generating a semantic emotion label corresponding to each denoised voice segment; Analyzing the emotion transition features between adjacent denoised voice segments through a context association layer in the emotion recognition network, and generating the real-time emotion fluctuation trajectory based on the emotion transition features; Performing feature fusion on the pitch change pattern, the speech rate fluctuation frequency, the speech amplitude distribution parameter, the semantic emotion label, and the real-time emotion fluctuation trajectory to generate the multi-level emotion feature set.

4. The method according to claim 3, characterized in that, The inputting of the multi-level emotion feature set into the safety risk prediction model to generate the safety behavior score curve of the driver within the preset time period includes: Through the time series analysis module in the security risk prediction model, perform periodic decomposition on the real-time emotional fluctuation trajectory, and extract the emotional mutation frequency and mutation amplitude of the driver within a preset time window; Through the risk association mapping module in the security risk prediction model, match the semantic emotional intensity parameter with the predefined driving operation risk level, and determine the vehicle control deviation probability corresponding to different emotional intensities; Through the dynamic scoring generation module in the security risk prediction model, combine the emotional mutation frequency, the mutation amplitude, and the vehicle control deviation probability, and calculate the security score values of the driver at multiple consecutive time points; Connect the security score values in chronological order to generate the security behavior score curve, and smooth the security behavior score curve to eliminate short-term noise interference.

5. The method according to claim 4, wherein Generate driving intervention prompt information corresponding to the risk nodes exceeding the preset threshold in the security behavior score curve, including: Perform gradient detection on the security behavior score curve, identify the continuous time interval where the decline rate of the security score value exceeds the rate threshold within a preset time interval, and mark the continuous time interval as a high-risk period; Extract the dominant emotional type in the multi-level emotional feature set corresponding to the high-risk period, and select a matching prompt content template from the preset intervention strategy library based on the dominant emotional type; Generate supplementary warning information in the prompt content template according to the vehicle control deviation probability and the current vehicle driving parameters, and the supplementary warning information includes recommended adjustment parameters related to the steering wheel angle and the throttle pedal depth; Combine the prompt content template and the supplementary warning information to generate the driving intervention prompt information, and sort the driving intervention prompt information according to the preset priority and send it to different display areas of the in-vehicle terminal.

6. The method according to claim 3, wherein The emotion recognition network is obtained through joint training of an acoustic feature extraction layer, a semantic parsing layer, and a context association layer, and the joint training includes: Obtain a labeled training speech data set, and each speech sample in the training speech data set is labeled with a pitch feature label, a semantic emotion label, and a context emotion continuity label; Input the speech sample into the initial emotion recognition network, generate pitch prediction features through the acoustic feature extraction layer, generate semantic emotion prediction labels through the semantic parsing layer, and generate emotion continuity prediction values through the context association layer; Calculate the first loss function between the pitch prediction feature and the pitch feature label, calculate the second loss function between the semantic emotion prediction label and the semantic emotion label, and calculate the third loss function between the emotion continuity prediction value and the context emotion continuity label; Perform weighted summation on the first loss function, the second loss function, and the third loss function to generate a joint training loss value, and update the parameters of the initial emotion recognition network through the backpropagation algorithm until convergence.

7. The method according to claim 6, wherein The joint training further includes: Add interference voice samples containing vehicle environmental noise to the training voice dataset, and set the noise suppression weight corresponding to the interference voice samples; When extracting features from the interference voice samples through the acoustic feature extraction layer, reduce the contribution degree of the frequency band features corresponding to the vehicle engine noise in the interference voice samples based on the noise suppression weight; When parsing the interference voice samples through the semantic parsing layer, increase the keyword retrieval intensity related to the driver's voice based on the noise suppression weight to cover the semantic ambiguity area caused by the noise.

8. The method according to claim 4, wherein The safety risk prediction model is obtained through joint training with multi-dimensional driving data, and the multi-dimensional driving data includes historical driving voice data, vehicle operation log data, and actual accident record data; The joint training of the multi-dimensional driving data includes: Input the historical driving voice data into the emotion recognition network to generate a set of historical emotion features, and convert the steering wheel steering angle, braking frequency, and lane deviation amount in the vehicle operation log data into operation risk indicators; Through the feature fusion module in the safety risk prediction model, perform time alignment and spatial mapping on the set of historical emotion features and the operation risk indicators to generate a joint input feature vector; Through the risk prediction module in the safety risk prediction model, predict the accident occurrence probability within a future time window based on the joint input feature vector, and calculate the prediction error loss according to the real accident time point in the actual accident record data; Iteratively optimize the safety risk prediction model based on the prediction error loss until the prediction accuracy of the accident occurrence probability reaches a preset convergence condition.

9. The method according to claim 8, characterized in that The joint training of the multi-dimensional driving data further includes: Introduce an attention weight allocation mechanism in the feature fusion module to adjust the contribution ratio of the set of historical emotion features and the operation risk indicators in the joint input feature vector; Among them, the attention weight allocation mechanism includes: Determine the emotion feature subset of the time overlapping part between the set of historical emotion features and the operation risk indicators through time correlation analysis; Calculate the covariance matrix between the emotion feature subset and the operation risk indicators, and generate emotion feature weights and operation index weights based on the eigenvalue distribution of the covariance matrix; Perform weighted fusion on the emotion feature subset and the operation risk indicators according to the emotion feature weights and the operation index weights to generate the joint input feature vector; After generating the driving intervention prompt information, the method further includes: Real-time monitor the driver's response operations to the driving intervention prompt information, and collect the updated driving voice data stream after the response operations; Extract emotion features from the updated driving voice data stream to generate an updated safety behavior scoring curve; If the updated safety behavior scoring curve does not return to within the safety threshold range within a preset time, an active intervention protocol of the vehicle control system is triggered, and the active intervention protocol includes at least one of automatically reducing the driving speed, starting an emergency braking preparation mode, and sending an assistance request to a traffic management center.

10. A driving safety analysis device, characterized in that, Comprising: a memory in which a computer program is stored; a processor for loading the computer program to implement the driving safety analysis method based on voice emotion recognition according to any one of claims 1-9.

Citation Information

Patent Citations

  • Road rage detection method and equipment

    CN113415286A

  • Vehicle-mounted voice interaction method, device, equipment, program product and automobile

    CN119181356A

  • Vehicle control method and device, computer equipment and storage medium

    CN119459733A

  • Driving behavior risk detection method, system and device and storage medium

    CN119559620A

  • Facial emotion recognition-based driving behavior analysis method and related device

    CN119741683A

Cited By

  • Vehicle visual sentry optimization method and device, storage medium and program product

    CN121354064A