Multimodal driver anger emotion regulation method based on hearing and smell

By employing a multimodal approach to driver anger regulation using auditory and olfactory senses, and leveraging speech feature extraction and olfactory stimulation, this method addresses the issues of low accuracy and poor applicability in existing driver emotion recognition technologies, achieving efficient emotion regulation and safe driving.

CN116985741BActive Publication Date: 2026-04-14CHONGQING UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, driver emotion recognition methods based on facial expressions and physiological signals suffer from low recognition accuracy and poor applicability, especially when environmental factors change, and wearable devices are highly invasive.

Method used

A multimodal driver anger regulation method based on hearing and smell is adopted. It extracts global acoustic features and local spectral features from the driver's voice signal, combines convolutional neural network and multi-head attention mechanism for emotion classification, and regulates emotions by playing specific audio and releasing odors.

Benefits of technology

It improves the accuracy and applicability of emotion recognition, enhances the stability and reliability of emotion regulation, and evaluates the effect of emotion regulation through driving risk theory and physiological data, thus achieving effective regulation of driver emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116985741B_ABST
    Figure CN116985741B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal driver anger emotion regulation method based on hearing and olfaction, and comprises the following steps: performing global acoustic feature extraction on a driver voice signal to obtain global feature information; performing local spectrum feature extraction on the driver voice signal to obtain local feature information; fusing the global feature information and the local feature information, and performing emotion classification on the fused feature information to obtain an emotion classification result; preparing an audio and an odor; and regulating an angry emotion in the emotion classification result by playing the audio and releasing the odor. The application has high emotion recognition accuracy, wide application range and good emotion regulation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of road traffic driving, specifically to a multimodal driver anger regulation method based on hearing and smell. Background Technology

[0002] The road traffic system is a complex system composed of people, vehicles, roads, and the environment. One of the common factors leading to traffic accidents is driver anger, which negatively impacts driver behavior in traffic, causing distraction, increasing aggressive and dangerous behavior, and leading to traffic violations and vehicle damage. Detecting a driver's angry emotional state and responding to it to regulate their emotions is crucial for improving driving safety.

[0003] Currently, driver emotion recognition primarily relies on analyzing facial expressions and physiological signals to regulate anger. Facial expression-based emotion recognition methods consider the visual information of human emotional expression, using high-quality cameras to capture the driver's facial features, resulting in relatively intuitive emotion recognition results. However, the recognition process is limited by environmental factors such as light intensity and background changes, leading to low accuracy. Physiological signal-based emotion recognition is more objective and can more realistically reflect the driver's emotional state, but the detection process requires wearable devices, which is highly invasive to the driver, consumes significant resources, and has poor applicability.

[0004] Therefore, a multimodal driver anger regulation method based on hearing and smell is needed to solve the above problems. Summary of the Invention

[0005] In view of this, the purpose of this invention is to overcome the shortcomings of the prior art and provide a multimodal driver anger regulation method based on hearing and smell, which has high accuracy in emotion recognition, wide applicability, and good emotion regulation effect.

[0006] The present invention provides a multimodal driver anger regulation method based on hearing and smell, comprising:

[0007] Global acoustic feature extraction is performed on the driver's speech signal to obtain global feature information;

[0008] Local spectral features are extracted from the driver's speech signal to obtain local feature information;

[0009] Global and local feature information are fused together, and the fused feature information is used to classify emotions to obtain the emotion classification result.

[0010] Audio production and scent preparation;

[0011] The anger emotion in the emotion classification results was modulated by playing audio and releasing scents.

[0012] Furthermore, global acoustic feature extraction is performed on the driver's speech signal, specifically including:

[0013] The speech signal is segmented into frames to obtain the time-domain feature parameters of each speech frame; the time-domain feature parameters include the fundamental frequency and the root mean square energy value.

[0014] Spectral analysis is performed on the speech signal to obtain its frequency domain characteristic parameters, including Mel-frequency cepstral coefficients.

[0015] Calculate the mean value of Mel frequency cepstral coefficients ;

[0016] Calculate the mean value of the cepstral coefficients at the Mel frequency. Fundamental frequency mean with the same dimension and root mean square energy value ;

[0017] Mean value of Mel frequency cepstral coefficients , fundamental frequency mean and root mean square energy value After standardization, the normalized Mel frequency cepstral coefficient characteristics are obtained. Fundamental frequency characteristics and root mean square energy characteristics ;

[0018] The standardized Mel frequency cepstral coefficient characteristics Fundamental frequency characteristics and root mean square energy characteristics Perform splicing, then through a layer containing A fully connected layer with 100 neurons maps high-dimensional feature vectors to a low-dimensional feature space, and finally outputs a global feature representation vector. .

[0019] Furthermore, local spectral feature extraction is performed on the driver's speech signal, specifically including:

[0020] Audio spectrum processing is performed on the speech signal to obtain the Mel spectrum diagram;

[0021] Taking the logarithm of the Mel spectrum yields the logarithmic Mel spectrum;

[0022] Use a convolutional neural network to extract time-frequency feature information from a log-Mel spectrogram;

[0023] Global adaptive average pooling is performed on the time-frequency feature information along the time axis to obtain the feature vector representing the time step. ;

[0024] Based on the multi-head attention mechanism, the feature vector The process is performed to obtain the local feature representation vector. .

[0025] Furthermore, global and local feature information are fused, and the fused feature information is used for emotion classification, specifically including:

[0026] Global features and local features are concatenated to obtain concatenated feature information;

[0027] Two fully connected layers are used to reduce the dimensionality of the concatenated feature information, and the sentiment category is predicted in the form of probability using a normalized exponential function.

[0028] Further, audio production is carried out, specifically including:

[0029] The voice recordings are made by the driver's friends or family members, using a gentle tone and containing words of reminder, praise, and compliments.

[0030] Furthermore, the preparation of the odor specifically includes:

[0031] Based on emotional valence and arousal, several different odors were selected as olfactory modulation materials;

[0032] From a number of different odors, odors with positive potency and low arousal were selected as target odors;

[0033] Mix the target odor with a colorless and odorless diluent according to... Configuration Concentration of aromatherapy.

[0034] Furthermore, it also includes using an emotion regulation success scale to measure the effectiveness of anger regulation.

[0035] Furthermore, it also includes: using horizontal and vertical risk values ​​to characterize the driver's overall driving performance and analyzing the effect of anger regulation, specifically including:

[0036] Calculate the horizontal risk value :

[0037] ;

[0038] in, The material stiffness of the object in a vehicle collision. It represents the equivalent mass of the vehicle. Indicates the lateral speed of the vehicle. This is the shortest distance between the vehicle's center of gravity and the lateral obstacle. The gradient descent coefficients of the potential risk field. Indicates the shortest distance from the road boundary to the center line of the lane;

[0039] Calculate longitudinal risk value :

[0040] ;

[0041] in, , Represents objects The risk field that radiates to the surrounding road environment; , as well as Both represent risk coefficients; and All of these are road-related factors; and Representing vehicles With vehicles Driver risk factors; and Representing objects With objects The quality; Represents objects With objects Vector distance between them; and Representing objects With objects longitudinal velocity;

[0042] The horizontal and vertical risk values ​​are normalized to determine the comprehensive risk coefficient. :

[0043] ;

[0044] If emotions are regulated, the overall risk factor If the size decreases, the emotion regulation effect is better; otherwise, the emotion regulation effect is poor.

[0045] Furthermore, it also includes: evaluating the effectiveness of anger regulation based on physiological data, specifically including:

[0046] Collect EEG data and calculate the left-right brain asymmetry values ​​at various frequencies. :

[0047] ;

[0048] in, and These represent the average power of two electrode channels corresponding to the left and right hemispheres of the brain, respectively.

[0049] If there is an asymmetry between the left and right hemispheres after emotional regulation If the size increases, the emotion regulation effect is better; otherwise, the emotion regulation effect is poor.

[0050] Collect electrocardiogram data and calculate average heart rate :

[0051] ;

[0052] in, Indicates time period Heart rate; It is a mean function;

[0053] If the average heart rate decreases after emotion regulation, the emotion regulation effect is better; otherwise, the emotion regulation effect is worse.

[0054] Furthermore, it also includes: conducting correlation analysis on subjective and objective evaluation data to evaluate the effectiveness of anger regulation, specifically including:

[0055] Collect several sets of subjective and objective evaluation statistics; the subjective and objective evaluation statistics include subjective evaluation data and objective state data; wherein, the objective state data includes driver physiological data and vehicle driving state data;

[0056] A difference analysis was performed on the subjective and objective evaluation statistics to obtain the test level value; the subjective and objective evaluation statistics with test level values ​​greater than the set threshold were used as the target statistics.

[0057] Correlation analysis was performed on the subjective evaluation data and objective state data in the target statistical data to obtain the correlation level between the subjective and objective data. If the correlation level Subjective evaluation data is then used to evaluate the moderating effect on anger.

[0058] The beneficial effects of this invention are as follows: This invention discloses a multimodal driver anger regulation method based on auditory and olfactory senses. By constructing a multi-feature fusion speech emotion recognition network, it complementarily fuses the acoustic features of the entire speech segment with the local representation features extracted by deep learning, thereby improving the accuracy of emotion recognition. Based on the parameterized regulation of single-modal stimuli (auditory and olfactory), it explores multimodal regulation of driver emotions, improving the stability and reliability of emotion regulation. Based on research on driver emotion regulation, and according to the driver's lateral and longitudinal control capabilities, it proposes a technical scheme for evaluating the effect of driver emotion regulation based on driving risk theory; combining various measurement methods and algorithms (physiological, behavioral, and subjective scales), it further analyzes the effect of driver emotion regulation by combining subjective and objective data. Attached Figure Description

[0059] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0060] Figure 1 This is a schematic diagram of the multi-feature fusion speech emotion recognition network framework of the present invention;

[0061] Figure 2 This is a schematic diagram of the MFCC coefficient processing procedure of the present invention;

[0062] Figure 3 This is a schematic diagram of the time-domain feature processing of the present invention;

[0063] Figure 4 This is a schematic diagram of the global acoustic feature generation of the present invention;

[0064] Figure 5 This is a schematic diagram of the convolutional neural network module of the present invention;

[0065] Figure 6 This is a schematic diagram of the time-frequency feature information serialization of the present invention;

[0066] Figure 7 This is a schematic diagram illustrating the principle of the multi-head self-attention module of the present invention;

[0067] Figure 8 This is a schematic diagram of the decision-making module principle of the present invention;

[0068] Figure 9 This is a diagram showing the distribution of brain electrodes in the international 10-20 system of this invention;

[0069] Figure 10 This is a radar chart showing the normalized data effect of the present invention. Detailed Implementation

[0070] The present invention will be further described below with reference to the accompanying drawings, as shown in the figures:

[0071] The present invention provides a multimodal driver anger regulation method based on hearing and smell, comprising:

[0072] Global acoustic feature extraction is performed on the driver's speech signal to obtain global feature information;

[0073] Local spectral features are extracted from the driver's speech signal to obtain local feature information;

[0074] Global and local feature information are fused together, and the fused feature information is used to classify emotions to obtain the emotion classification result.

[0075] Audio production and scent preparation;

[0076] The anger emotion in the emotion classification results was modulated by playing audio and releasing scents.

[0077] This invention proposes a parallel multi-feature fusion speech emotion recognition network that complementarily fuses global acoustic features and local deep learning representation features. The overall framework of this emotion recognition network model is as follows: Figure 1 As shown. The emotion recognition network model mainly consists of a global acoustic feature extraction module, a spectral local feature extraction module, and a decision module. The global feature extraction module extracts the fundamental frequency of the speech ( Global feature information such as root mean square energy value and MFCC coefficients (Mel frequency cepstral coefficients) are concatenated into a global feature vector. The local feature extraction module is a temporal convolutional network (CNN_TCN_SA) based on a self-attention mechanism, used to mine the time-frequency features of the log-Melbourne spectrogram and generate a local feature representation vector. Finally, the features from the two levels are fused, and a decision module is used to classify speech emotions.

[0078] In this embodiment, global acoustic feature extraction of the driver's voice signal is performed, specifically including:

[0079] Before extracting global features, the speech signal needs to be segmented into frames because speech is... The time range is the most stable, therefore a frame length of 40ms and a frame shift of 10ms are used to extract the fundamental frequency of each speech frame using a sliding time window method. Root mean square energy value Isochronous time domain characteristic parameters.

[0080] To extract frequency-domain related speech emotion features, a 40ms Hamming window was superimposed on each original speech frame, and a 30-dimensional Mel frequency cepstral coefficient was generated as a frequency domain feature parameter by using a 100ms Fourier transform window and following the MFCC coefficient extraction process.

[0081] MFCC coefficients represent the spectral envelope energy of audio, providing a compact vector representation of the sound wave's amplitude spectrum. This is a state-of-the-art method for formalizing speech features. (Extracted...) It is a two-dimensional floating-point time series, where the horizontal axis represents the time axis, and... The vertical axis represents the number of frames in the entire speech segment, while the horizontal axis represents the frequency domain, indicating frequency domain feature parameters. The mean value of each frequency domain feature parameter is calculated along the time axis. As an acoustic feature at the entire discourse level, such as Figure 2 As shown.

[0082] baseband Root mean square energy value This reflects the overall changes in pitch and volume of speech over time. Because the frame length of the speech segments is relatively small, both of these feature sequences are quite long. To correspond with the feature dimension of MFCC, the n-frame speech is divided into 30 long speech frames. The fundamental frequency and mean energy of each long segment are calculated, generating two 30-dimensional feature vectors. , Approximate representation of global acoustic features, such as Figure 3 As shown.

[0083] Because the extracted speech feature units differ, they vary in data dimensions. To ensure that the different speech features are on the same order of magnitude without sacrificing the model's recognition performance, z-score normalization is applied to the three different acoustic features. The normalization formula is as follows:

[0084] ;

[0085] In the formula, This represents the raw speech feature data. This represents the mean of all data under the current speech features. This represents the standard deviation of all data under the current speech features.

[0086] Standardize global acoustic features , , The features are concatenated together, and then passed through a fully connected layer with 30 neurons to map the high-dimensional feature vector to a low-dimensional feature space, outputting the final global feature representation vector. The specific generation process is as follows: Figure 4 As shown.

[0087] In this embodiment, local spectral feature extraction of the driver's voice signal is performed, specifically including:

[0088] The speech signal is processed by audio spectrum processing to obtain the Mel spectrogram; the logarithm of the Mel spectrogram is taken to obtain the logarithmic Mel spectrogram; the Mel spectrogram is converted to a logarithmic scale by taking the logarithm of the Mel spectrogram, thereby improving the stability and discriminative ability of the features.

[0089] This invention employs a convolutional neural network (CNN) to extract time-frequency features from a log-Melbourne spectrogram. The CNN structure used is a shallow network consisting of five convolutional layers and two max-pooling layers. In the first convolutional layer, two different convolutional kernels are used in parallel: one with a longer time span and the other with a longer frequency span. The two different outputs are concatenated along the channel dimension and then fed into subsequent convolutional layers. The second and third layers are a combination of convolutional and max-pooling layers for feature dimensionality reduction; the pooling layer size is set to a specific value. In the remaining convolutional layers, the same convolution method is used to pad the output feature map, and the kernel size is set to... To ensure non-linear data transformation during the forward propagation of the network model, the Rectified Luminous Interval (ReLU) function is used as the activation function for all convolutional layers in the model, and batch normalization is fitted to all neurons in the five convolutional layers to accelerate model training. Specific CNN modules include... Figure 5 As shown.

[0090] As the number of convolutional layers in the CNN module increases, the resolution of the feature maps continuously decreases, and the speech emotion feature information is gradually transferred to the feature map channels. Due to the limitation of the total amount of speech emotion dataset, the number of convolutional layers should not be too large. Therefore, this paper adopts a lightweight CNN network. The specific network parameters are shown in Table 1 below:

[0091] Table 1

[0092]

[0093] CNN modules capture time-frequency characteristics from speech spectrograms but do not consider the temporal information of speech. To fully consider the temporal information of the output feature sequence, a temporal modeling approach from natural language processing is used to encode the speech feature sequence. For example... Figure 6 As shown, high-level speech features learned from the CNN module It is a three-dimensional array. This represents the number of channels in the feature map. and These represent the frequency and time span of the speech, respectively, corresponding to the height and width of the feature image. This data cannot be directly integrated with a time-series model; dimensionality reduction is required to generate a two-dimensional array. Therefore, global adaptive average pooling is performed along the time axis on the feature map to achieve dimensionality reduction while ensuring no loss of feature data at any given time step. The result is a feature sequence. Here, Characterizes each moment eigenvectors.

[0094] While the attention given to speech features is the same at all times, emotional features are unevenly distributed across the entire speech segment, appearing only at certain specific moments in the conversation. Given the effective focusing of local information by attention mechanisms, an attention module is used to focus on moments in the output sequence where emotional features are more pronounced. This invention uses a multi-head self-attention module; the input to the self-attention mechanism includes three encoded vectors: query, key, and value. It uses the same linear weights for the embedded feature vector at each time step. Encode, that is This means that in the input speech sequence, the speech features at each time step are compared with the speech features at all other time steps to calculate the similarity, so as to explore the dependencies between features within the speech sequence, reduce information loss, and assign greater weight to emotion-related parts.

[0095] Multi-head attention can combine information from different subspaces, learn more emotion-related features, and improve the overall recognition performance of the model. Figure 7 A schematic diagram of the attention module is shown.

[0096] Based on the aforementioned multi-head attention mechanism, the feature vector The process is performed, and the output is a local feature representation vector. .

[0097] In this embodiment, the decision module effectively complements and fuses global acoustic features and local spectral features, and then maps the fused distributed features to the emotion space through a fully connected layer to achieve emotion classification.

[0098] Local feature representation vector With global feature representation vector The final emotional representation features are then pieced together. Then, two fully connected layers are used to reduce the dimensionality of the data, and through... The function (normalized exponential function) is used to predict sentiment. The classification calculation process is as follows:

[0099] ;

[0100] ;

[0101] In the formula, , These represent the first and second parts of the decision-making module. The parameters (weights and biases) of each fully connected layer. , ; Used for splicing. Used for activation; Represents the probability of emotion classification categories; decision modules such as Figure 8 As shown.

[0102] Using the above method, the probability corresponding to the driver's emotion category can be obtained. When the probability of anger exceeds the set threshold, it can be considered that the driver is angry and it is necessary to regulate the driver's anger.

[0103] In this embodiment, when the voice emotion recognition model identifies the driver's anger, the present invention reduces the driver's anger level through auditory and olfactory multimodal modulation intervention, thereby improving driving safety.

[0104] This invention uses personalized voice as the auditory modulation material. The voice content employs an attention-based deployment strategy to divert the driver's anger, effectively shifting their attention and improving their negative emotions, thus positively impacting driving performance and safety. The voice content is recorded by the driver's friends or family. Furthermore, the voice clips use reminders, praise, and compliments, such as positive comments, to commend the driver; the delivery style is notification-based, informing the driver about the current road environment. Auditory intervention is achieved through playback via in-vehicle agent software. Table 2 lists some of the voice content used for anger modulation.

[0105] Table 2

[0106]

[0107] The scent used in this invention is selected based on emotional valence and arousal dimensions. From seven different scent types, the jasmine scent, with its positive valence and low arousal, was chosen as the olfactory modulator. A 10% concentration of jasmine aromatherapy solution was prepared by mixing jasmine essential oil and a colorless, odorless diluent at a ratio of 1:9, and released by an in-vehicle device for 10 seconds to ensure the driver could fully smell the aroma.

[0108] In this embodiment, the present invention uses an emotion regulation success scale to measure the driver's subjective emotion regulation effect and characterize the degree of emotion relief. The scale adopts a 9-point Likert scale design, where 1 point indicates that the emotion regulation plan is not successful at all, and 9 points indicates that the regulation is very successful. The intermediate regulation effect scores are evenly distributed.

[0109] This embodiment also includes: using horizontal and vertical risk values ​​to characterize the driver's overall driving performance and analyzing the effect of anger regulation.

[0110] Vehicle driving data is exported via onboard OBD or from the simulator's backend data. This invention combines driving risk field theory with horizontal and vertical risk values ​​that integrate multiple discrete driving indicators to characterize the driver's overall driving performance, indirectly reflecting the effectiveness of anger intervention.

[0111] Driving risks can be categorized into lateral risks and longitudinal risks based on the vehicle's motion state.

[0112] ① Horizontal risks

[0113] Lateral risk refers to the potential risks that occur when a vehicle moves laterally. Within the lateral region, all obstacles encountered by the target vehicle, such as static road boundaries, traffic barriers, and crash barriers, are considered as a finite scalar risk domain, which lies within the predicted motion space of the target vehicle. Based on probabilistic motion prediction, the lateral risk value is approximated by multiplying the expected collision probability between the vehicle and the obstacle by the energy generated during the collision. The calculation formula is as follows:

[0114] ;

[0115] In the formula, This represents the horizontal risk value. The material stiffness of the object in a vehicle collision. It represents the equivalent mass of the vehicle. Indicates the lateral speed of the vehicle. This is the shortest distance between the vehicle's center of gravity and a lateral obstacle, which refers to the road boundary. It represents the shortest distance from the road boundary to the center line of the lane. The gradient descent coefficients for the potential risk field are generally assumed to be... That is, the collision probability term reaches a critical value at the center of the lane. ).

[0116] As can be seen from the formula, lateral risk is the expected collision energy scaled by parameters. and collision probability term It consists of two parts; the larger the value, the higher the potential lateral risk. At the same time, the formula describes how the collision probability increases. The decrease due to the increase in distance is intuitively understandable; objects at road boundaries that are further away provide drivers with more opportunities to avoid collisions, thus reducing the risk of a collision.

[0117] ② Vertical risks

[0118] The longitudinal risk field and the lateral risk field are similar in concept and generation method. They refer to the risk of a head-on collision with other traffic elements during the longitudinal movement of a vehicle. The degree of collision hazard is also measured by the product of the collision probability and the collision energy. The longitudinal risk value is... The calculation formula is as follows:

[0119] ;

[0120] ;

[0121] In the formula, Represents objects The risk field radiating to the surrounding road environment; the magnitude of the field strength reflects the object's... The potential level of danger, , , This represents the risk factor, used to adjust for different objects. The magnitude of the risk value is related to the object's shape and type attributes; , These are all road-related factors, determined by driving environment conditions such as road adhesion coefficient and visibility. , Representing vehicles With vehicles The driver risk factor is set to 0 if the object is not a vehicle. , Representing objects With objects quality Represents objects With objects Vector distance between them , Representing objects With objects The longitudinal velocity.

[0122] This represents the sum of risks posed to the target vehicle after the risk fields of all objects are superimposed. The larger the value, the more dangerous the vehicle's longitudinal movement.

[0123] ③Comprehensive risk coefficient

[0124] To explore the potential risks arising from driver operations in dynamic environments, a driving risk assessment model coupling horizontal and vertical risk fields is adopted. Therefore, the horizontal and vertical risk data are first normalized to eliminate the influence of dimensions, and then the comprehensive risk coefficient is determined according to the following formula. :

[0125] ;

[0126] Driving risk, centered on the driver, is a real-time calculation of the scope and degree of influence of various factors. It not only characterizes the level of safety while driving but also indirectly reflects the driver's control over the vehicle. Therefore, driving risk indicators can be used to quantify the effectiveness of anger intervention, i.e., to compare the effect size of different moderating methods on expressive suppression. If, after emotion regulation, the overall risk coefficient... If the size decreases, the emotion regulation effect is better; otherwise, the emotion regulation effect is poor.

[0127] This embodiment also includes: evaluating the effect of anger regulation based on physiological data.

[0128] Electroencephalogram (EEG) data was collected and recorded using an EEG machine to evaluate the regulation of anger in drivers. The analysis of the EEG data utilized the asymmetry of frontal lobe activity, primarily analyzing the frontal lobe region of the driver's brain. The left and right brain electrodes in the EEG data were symmetrically related; the selected frontal lobe electrodes were: FP1-FP2, AF3-AF4, F3-F4, and F7-F8 (e.g., ...). Figure 9 (As shown), where odd numbers represent electrodes in the left brain region and even numbers represent electrodes in the right brain region. Based on the correspondence between the left and right brain electrodes, the left-right brain asymmetry value for each frequency is calculated using the following expression:

[0129] ;

[0130] in, and These represent the average power of the two electrode channels corresponding to the left and right hemispheres of the brain, respectively. Taking the alpha waves of electrodes F3 and F4 as an example, the asymmetry value is calculated by subtracting the natural logarithmic value of the alpha wave of F3 (left hemisphere) from the natural logarithmic value of the alpha wave of F4 (right hemisphere).

[0131] Since the alpha wave band ability is inversely proportional to the activity of the cerebral cortex, a higher score indicates relatively greater left-brain activity and a better regulatory effect; while a lower score indicates relatively greater right-brain activity and a poorer regulatory effect.

[0132] Electrocardiogram (ECG) data were acquired and recorded using a multi-channel physiological analyzer. Analysis and evaluation were primarily based on heart rate (HR), mainly real-time heart rate. Generally, a higher real-time heart rate (HR) correlates with a higher level of anger. The ECG time-domain signals before and after regulation were obtained experimentally to calculate the average heart rate. :

[0133] ;

[0134] in, Indicates time period Heart rate; It is a mean function;

[0135] If the average heart rate decreases after emotion regulation, the emotion regulation effect is better; otherwise, the emotion regulation effect is worse.

[0136] This embodiment also includes: performing correlation analysis on subjective and objective evaluation data to evaluate the effect of anger regulation.

[0137] Several sets of subjective and objective evaluation statistics were collected; these statistics included subjective evaluation data and objective state data; the objective state data included driver physiological data and vehicle driving status data; thus, descriptive statistical information of data from different dimensions was obtained. Subjective evaluation data recorded the participants' experience level of anger; however, because human judgment of personal emotions is limited by subjective consciousness, it is easy to have discrepancies between words and actions, or to misrepresent one's true feelings. Therefore, the accuracy of subjective evaluations still needs to be verified by objective data.

[0138] A difference analysis was performed on the subjective and objective evaluation statistics to obtain the significance level. Subjective and objective evaluation statistics with significance levels greater than the set threshold were used as target statistics. If the difference analysis failed (i.e., significance level p > 0.05), it indicates that the differences between different groups were not caused by artificially controlled independent variables, and the measurement error of the data itself was greater than the data differences caused by the independent variables. Therefore, the data could not be used for subsequent data analysis. Only the subjective and objective evaluation statistics that passed the difference analysis (significance level p < 0.05) were selected for correlation analysis.

[0139] Correlation analysis was performed on the subjective evaluation data and objective state data in the target statistical data to obtain the correlation level between the subjective and objective data. If the correlation level Subjective evaluation data is then used to evaluate the moderating effect on anger.

[0140] Specifically, the Pearson correlation coefficient, proposed by statistician Carl Pearson, is used, and the specific calculation method is as follows:

[0141] ;

[0142] In the formula, Refers to subjective data values ​​under a certain variable. The mean of all subjective data values ​​for a given variable. The variance of all subjective data values ​​under a certain variable; in the formula Refers to a certain type of objective data value under a certain variable (such as EEG). The mean of a certain type of objective data values ​​under a certain variable. The variance of a certain type of objective data values ​​under a certain variable.

[0143] Correlation level of subjective and objective data This value reflects the level of correlation between different data points. Under the premise of a significance level of p < 0.05, This indicates a weak correlation between the subjective and objective data. This indicates a correlation between subjective and objective data. This indicates a strong correlation between the subjective and objective data. For subjective-objective analysis results, when the correlation level is above moderate, that is... This indicates a strong correlation between objective and subjective data. Objective data can be used to verify subjective data, and subjective data is reliable and can be used to evaluate the effectiveness of anger regulation in subjects.

[0144] It should be noted that in most cases, attention should be paid not only to individual measurement results but also to the overall situation of emotion regulation. This requires integrating multiple measurement results into a single type of comprehensive score. This invention uses a percentage-based method to normalize all measurement data to 0-1. For positively correlated measures of emotion regulation, we use formula (9.1) to normalize the data. For negatively correlated measures of emotion regulation, we use formula (9.2) to normalize the data.

[0145] For indicators positively correlated with emotion regulation:

[0146] ; (9.1)

[0147] in, , , These represent the normalized data, the original data, and the maximum value in the original data, respectively.

[0148] For indicators negatively correlated with emotion regulation:

[0149] ; (9.2)

[0150] in, , , These represent the normalized data, the original data, and the minimum value in the original data, respectively.

[0151] After normalization, all data showed a positive correlation with sentiment regulation (higher values ​​are better). The effect of normalized data is like a radar. Figure 10 As shown, the comparison of the adjustment effects across different dimensions can be seen intuitively.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multimodal driver anger regulation method based on auditory and olfactory senses, characterized in that: include: Global acoustic feature extraction is performed on the driver's speech signal to obtain global feature information; Global acoustic feature extraction of the driver's speech signal, specifically including: The speech signal is segmented into frames to obtain the time-domain feature parameters of each speech frame; the time-domain feature parameters include the fundamental frequency and the root mean square energy value. Spectral analysis is performed on the speech signal to obtain its frequency domain characteristic parameters, including Mel-frequency cepstral coefficients. Calculate the mean value of Mel frequency cepstral coefficients ; Calculate the mean value of the cepstral coefficients at the Mel frequency. Fundamental frequency mean with the same dimension and root mean square energy value ; Mean value of Mel frequency cepstral coefficients , fundamental frequency mean and root mean square energy value After standardization, the normalized Mel frequency cepstral coefficient characteristics are obtained. Fundamental frequency characteristics and root mean square energy characteristics ; The standardized Mel frequency cepstral coefficient characteristics Fundamental frequency characteristics and root mean square energy characteristics Perform splicing, then through a layer containing A fully connected layer with 100 neurons maps high-dimensional feature vectors to a low-dimensional feature space, and finally outputs a global feature representation vector. ; Local spectral features are extracted from the driver's speech signal to obtain local feature information; Global and local feature information are fused together, and the fused feature information is used to classify emotions to obtain the emotion classification result. Audio production and scent preparation; The anger emotion in the emotion classification results was modulated by playing audio and releasing scents.

2. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Local spectral feature extraction of the driver's speech signal, specifically including: Audio spectrum processing is performed on the speech signal to obtain the Mel spectrum diagram; Taking the logarithm of the Mel spectrum yields the logarithmic Mel spectrum; Use a convolutional neural network to extract time-frequency feature information from a log-Mel spectrogram; Global adaptive average pooling is performed on the time-frequency feature information along the time axis to obtain the feature vector representing the time step. ; Based on the multi-head attention mechanism, the feature vector The process is performed to obtain the local feature representation vector. .

3. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Global and local feature information are fused, and the fused feature information is used for emotion classification, specifically including: Global features and local features are concatenated to obtain concatenated feature information; Two fully connected layers are used to reduce the dimensionality of the concatenated feature information, and the sentiment category is predicted in the form of probability using a normalized exponential function.

4. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Audio production specifically includes: The voice recordings are made by the driver's friends or family members, using a gentle tone and containing words of reminder, praise, and compliments.

5. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Preparation of odor, specifically including: Based on emotional valence and arousal, several different odors were selected as olfactory modulation materials; From a number of different odors, odors with positive potency and low arousal were selected as target odors; Mix the target odor with a colorless and odorless diluent according to... Configuration Concentration of aromatherapy.

6. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Also includes: The Emotion Regulation Success Scale was used to measure the effectiveness of anger regulation.

7. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Also includes: The overall driving performance of drivers is characterized by horizontal and vertical risk values, and the effect of anger regulation is analyzed, specifically including: Calculate the horizontal risk value : ; in, The material stiffness of the object in a vehicle collision. It represents the equivalent mass of the vehicle. Indicates the lateral speed of the vehicle. This is the shortest distance between the vehicle's center of gravity and the lateral obstacle. The gradient descent coefficients of the potential risk field. Indicates the shortest distance from the road boundary to the center line of the lane; Calculate longitudinal risk value : ; in, , Represents objects The risk field that radiates to the surrounding road environment; , as well as Both represent risk coefficients; and All of these are road-related factors; and Representing vehicles With vehicles Driver risk factors; and Representing objects With objects The quality; Represents objects With objects Vector distance between them; and Representing objects With objects longitudinal velocity; The horizontal and vertical risk values ​​are normalized to determine the comprehensive risk coefficient. : ; If emotions are regulated, the overall risk factor If the size decreases, the emotion regulation effect is better; otherwise, the emotion regulation effect is poor.

8. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Also includes: Based on physiological data, the effectiveness of anger regulation was evaluated, specifically including: Collect EEG data and calculate the left-right brain asymmetry values ​​at various frequencies. : ; in, and These represent the average power of two electrode channels corresponding to the left and right hemispheres of the brain, respectively. If there is an asymmetry between the left and right hemispheres after emotional regulation If the size increases, the emotion regulation effect is better; otherwise, the emotion regulation effect is poor. Collect electrocardiogram data and calculate average heart rate : ; in, Indicates time period Heart rate; It is a mean function; If the average heart rate decreases after emotion regulation, the emotion regulation effect is better; otherwise, the emotion regulation effect is worse.

9. The multimodal driver anger regulation method based on hearing and smell according to claim 1, characterized in that: Also includes: Correlation analysis was conducted on subjective and objective assessment data to evaluate the effectiveness of anger regulation, specifically including: Collect several sets of subjective and objective evaluation statistics; the subjective and objective evaluation statistics include subjective evaluation data and objective state data; wherein, the objective state data includes driver physiological data and vehicle driving state data; A difference analysis was performed on the subjective and objective evaluation statistics to obtain the test level value; the subjective and objective evaluation statistics with test level values ​​greater than the set threshold were used as the target statistics. Correlation analysis was performed on the subjective evaluation data and objective state data in the target statistical data to obtain the correlation level between the subjective and objective data. If the correlation level Subjective evaluation data is then used to evaluate the moderating effect on anger.

Citation Information

Patent Citations

  • Voice emotion recognition model and method based on joint feature representation

    CN108899051A

  • Multi-mode driver emotion assistant regulation method

    CN111329498A

  • KR20220098991A