Driver Voice Emotion Recognition Method for Intelligent Driving

By constructing a three-dimensional sound feature data set and training a Gaussian hybrid model, a multi-model fusion recognition method is formed, which solves the problem of low accuracy of single model recognition in intelligent driving scenarios and improves the accuracy of emotion recognition.

CN115227246BActive Publication Date: 2025-07-25NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210802515.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-07-25
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

In the existing intelligent driving scenarios, the speech emotion recognition method of a single model cannot effectively utilize the voice characteristics of different categories of people, resulting in low emotional recognition accuracy and serious information interference.

Method used

A three-dimensional sound feature data set is constructed, K clustering centers are found through clustering methods and Gaussian mixed model is trained to form a multi-model fusion recognition method, and sample user data of different personnel categories are used to enhance feature representation capabilities.

Benefits of technology

The accuracy of emotion recognition in intelligent driving scenarios is improved, and the accuracy of recognition is enhanced through multimodal information fusion, solving the problem of low accuracy of single model recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115227246B_ABST
    Figure CN115227246B_ABST
Patent Text Reader

Abstract

The present invention discloses a driver voice emotion recognition method for intelligent driving, including: collecting voice data with different emotions of different users in a driving scenario to construct a driver three-dimensional voice feature data set; then constructing a clustering multi-model training method based on three-dimensional voice features, obtaining sample users of different personnel categories through a clustering method based on three-dimensional voice features, and then using the data of sample users of different personnel categories to train a Gaussian mixture model to form a voice emotion recognition model for different personnel categories; after that, the user inputs the voice in a normal emotional state during initialization for initialization classification to obtain a general reference model and reference parameters; finally, during the running recognition stage, the voice of the user collected in real time is input. After the voice sample passes through the reference model, it is input into other models for multi-modal information fusion and judgment, and finally the recognition result is output. The present invention improves the accuracy of emotion recognition in the intelligent driving scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence. Background Art

[0002] With the rapid development of intelligent vehicles, the detection of drivers' emotions in dynamic driving scenarios benefits from intelligent cockpits and human-machine interaction systems. Therefore, driver emotion monitoring has become a popular research topic. General emotion recognition methods can be divided into two categories according to the detected signals: recognition based on physiological signals such as electroencephalogram, respiration, and heart rate; recognition relying on non-physiological signals, including sound signals and facial expressions. Driver emotion is an external manifestation of the driver's physiological and psychological states, which affects the driver's driving decisions and behaviors. Research shows that negative emotions such as anger, fatigue, and anxiety can reduce the driver's risk perception, easily lead to aggressive driving behaviors, and significantly increase the risk of crashes. Thus, it can be seen that driver emotion plays a crucial role in traffic safety, and accurately recognizing driver emotion is crucial for improving the driving safety and comfort of intelligent vehicles. Currently, the models used in the main voice recognition networks are relatively single. Since the voice characteristics of different categories of people vary greatly, using a single model to recognize the different emotions of different categories of people will result in the inability to fully utilize voice characteristics, and numerous information will interfere with each other, leading to unsatisfactory accuracy in emotion recognition. To sum up, improving the accuracy of speech emotion recognition has become an urgent problem to be solved in the intelligent driving scenario. Summary of the Invention

[0003] Object of the Invention: To solve the problems existing in the above-mentioned prior art, the present invention provides a method for recognizing the voice emotion of a driver for intelligent driving.

[0004] Technical Solution: The present invention provides a method for recognizing the voice emotion of a driver for intelligent driving, specifically including the following steps:

[0005] Step 1: Collect voice data of different users in Q different emotional states in the driving scenario, and construct a three-dimensional voice feature data set for each user;

[0006] Step 2: Use the clustering method to find K clustering centers {o(1), o(2),..., o(k),..., o(K)} in the constructed three-dimensional voice feature data set, where o(k) represents the k-th clustering center, and k = 1, 2,..., K; Train Gaussian mixture models with the K clustering centers to obtain K Gaussian mixture models {G1, G2,..., G k ,...., G K}, where G k represents the k-th Gaussian mixture model;

[0007] Step 3: The user inputs the voice in the normal emotional state during initialization. According to the distance between the voice and each cluster center, the optimal cluster center k corresponding to the voice input by the user in the normal emotional state during initialization is obtained. * ; The Gaussian mixture model corresponding to k * is used as the reference model. Calculate the performance metrics of

[0008] Step 4: During the driving process, the voice of the user is collected in real time. The currently collected voice is input into to obtain the probabilities of Q emotional states where q represents the q-th emotional state, q = 1, 2, …, Q, represents the probability of the q-th emotional state output; Denote the emotional state corresponding to the maximum probability as class q* ;

[0009] Step 5: Calculate the shortest distance between the currently collected voice in the q-th emotional state and all cluster centers; Denote the cluster center corresponding to the shortest distance as and the Gaussian mixture model corresponding to as to obtain the set of cluster centers corresponding to each emotional state and the set of Gaussian mixture models corresponding to each cluster center Input the currently collected voice into to obtain the probabilities of the Q emotional states output Denote the maximum probability as Denote the emotional state corresponding to the maximum probability as and calculate

[0010] Step 6: Judge the current emotional state of the user according to the performance metric and the performance metric .

[0011] Furthermore, frame and window the voice data of each user in the q-th emotional state; and calculate the energy contained in each frame according to the following formula:

[0012] E n = x(n) 2 * ω(n) 2

[0013] where E n is the energy of the n-th frame, x(n) is the frame signal, and ω(n) is the Hamming window;

[0014] Classify all frame signals of each user in the q-th emotional state according to the following formula:

[0015]

[0016] where t represents the frame category, t = 1, 2, 3, 4; E t represents the range of frame energy corresponding to the t-th frame category;

[0017] Calculate the time ratio l of each user in the q-th emotional state for the t-th frame category t :

[0018]

[0019] where Time t represents the total duration of the t-th frame category of the user in the q-th emotional state;

[0020] Calculate the short-time average frequency a1, short-time mean square frequency a2 and formant frequency a3 of each user in the q-th emotional state, and obtain the fused prosodic feature m of the user in the q-th emotional state:

[0021] m = w1·a1 + w2·a2 + w3·a3

[0022] where w1, w2 and w3 all represent relative importance;

[0023] Combine the Q emotional states, the time ratio l of the user in the Q emotional states for the t-th frame category t , and the fused prosodic feature m of the user in the q-th emotional state to form the three-dimensional voice feature dataset of the user.

[0024] Furthermore, the method for finding K clustering centers in step 2 is specifically as follows: First, take all the three-dimensional voice feature datasets as samples; randomly select K samples from all the samples as clustering centers, and then calculate the distance d(i, k) from the i-th sample to the k-th clustering center according to the following formula:

[0025]

[0026] where l iqt is the time ratio of the t-th frame category of the user corresponding to the i-th sample in the q-th emotional state, l kqt is the time ratio of the t-th frame category of the user corresponding to the k-th clustering center in the q-th emotional state; m iq is the fused prosodic feature of the user corresponding to the i-th sample in the q-th emotional state, m kq is the fused prosodic feature of the user corresponding to the k-th clustering center in the q-th emotional state, liqs The short-time average energy of the voice data of the user corresponding to the i-th sample in the q-th emotional state;

[0027] Evaluate the distance of each sample to the cluster center, and assign the sample to the cluster to which the cluster center closest to the sample belongs. Then update the cluster centers of each cluster to obtain new cluster centers, and calculate the distance between each sample and the new cluster centers until the number of iterations is greater than the preset number. Finally, obtain the set {o(1), o(2),... o(K)} of the last K cluster centers.

[0028] Furthermore, in step 2, a Gaussian mixture model is trained with K cluster centers to obtain K Gaussian mixture models corresponding to the K cluster centers, specifically:

[0029] Initialize a set of parameters using the K-means algorithm according to the following formula:

[0030] θ d ={α d , μ d , σ d}

[0031] where D is the number of Gaussian sub-models, θ d is the parameter of the d-th Gaussian sub-model, d = 1, 2,..., D, μ d is the mean of the d-th Gaussian sub-model, σ d is the variance of the d-th Gaussian sub-model, and α d is the standard deviation of the d-th Gaussian sub-model;

[0032] Calculate the probability γ d from the d-th Gaussian sub-model of θ d :

[0033]

[0034] where x q is the three-dimensional voice feature vector of the q-th emotional state, and Φ(x q |θ d ) is the Gaussian density function obtained when the input feature vector of the d-th Gaussian sub-model is x q ;

[0035] Then, according to the probability γ d , calculate the new parameter value Calculate the value of, if If it is less than the threshold, terminate the iterative calculation; otherwise, replace the old parameters with the new parameters and start the iteration until the absolute value of the difference between the parameter value obtained finally and the parameter value obtained in the previous iterative calculation is less than the preset threshold; finally, obtain the k-th Gaussian mixture model G k The expression of

[0036]

[0037] Further, in step 3, calculate the distance d between the sound of the user in the normal emotional state input at initialization and each cluster center according to the following formula 1 (∧, o(k)):

[0038]

[0039] where ∧ represents the user, l ∧1t is the time proportion of the t-th frame category in the normal emotional state when the user initializes, l o(k)1t is the time proportion of the t-th frame category in the normal emotional state of the k-th cluster center, l ∧1s is the short-time average energy of the sound data in the normal emotional state when the user initializes, m ∧1 is the fused prosody feature in the normal emotional state when the user initializes, m o(k)1 is the fused prosody feature of the k-th cluster center in the normal emotional state.

[0040] Further, step 5 is specifically as follows: Assume that the currently collected sound of the user corresponds to the q-th emotional state, and calculate the distance d from the current sound data to all cluster centers according to the following formula q (∧ c , o(k)):

[0041]

[0042] where ∧ c represents the current user, is the time proportion of the t-th frame category of the current user in the q-th emotional state, l o(k)qt is the time proportion of the t-th frame category of the k-th cluster center in the q-th emotional state, is the short-time average energy of the sound data of the current user in the q-th emotional state, is the fused prosody feature of the current user in the q-th emotional state, m o(k)q is the fused prosody feature of the k-th cluster center in the q-th emotional state;

[0043] Calculate the performance index of the Gaussian mixture model set according to the following formula

[0044]

[0045] Among them represents the distance between the voice of the current user in the q-th emotional state and the clustering center therebetween.

[0046] Furthermore, specifically in step 6:

[0047] Form a performance set from the performance indicators of each Gaussian mixture model in the Gaussian mixture model set

[0048] In and the performance set select the maximum performance indicator, denoted as OP * ; Determine the emotional state q of the current user according to the following formula end :

[0049]

[0050] Furthermore, the performance indicator is calculated according to the following formula in step 3

[0051]

[0052] Among them, represents the probability that the voice of the user in the normal emotional state during initialization enters to obtain the probability of the normal emotional state, and d * represents the distance between the voice data input by the user during initialization and the optimal clustering center k * , and d * = d 1 (∧, o(k * ))

[0053] Beneficial effects: For most machine learning algorithms, the models are relatively fixed and cannot effectively identify the emotions of different types of people in different emotions. The present invention proposes a method for personalized recognition of voice emotions with multiple models. First, a large amount of voice emotion data of people is collected to construct a three-dimensional voice emotion data set. Different sample users of different personnel categories are obtained through a clustering method based on three-dimensional voice features. Then, the Gaussian mixture model is trained using the data of different sample users of different personnel categories to form a voice emotion recognition model for different personnel categories; the subsequent fusion step of the constructed multiple models considers the problem of low recognition accuracy of a single model. By fusing multi-modal information, key information is better enhanced, and the feature representation ability is strengthened, so that the accuracy of emotion recognition in the intelligent driving scenario is better improved compared with the existing methods. ​BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0055] The accompanying drawings, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0056] As Figure 1 shown, the method of this embodiment includes the following steps:

[0057] Step 1: Collect voice data containing different emotions of different categories of users in a driving scenario, and construct a user three-dimensional voice feature dataset.

[0058] First, collect voice data of Q emotions of a large number of users of different categories in a driving scenario, and the duration of each voice data is 10 seconds; then perform frame segmentation and windowing on the voice data obtained by the sample user in the qth emotion, q = 1, 2,..., Q, and calculate the energy contained in each frame. The formula is as follows:[[]]

[0059] E n = x(n) 2 * ω(n) 2 where E n ∈[0, 100dB]

[0060] where E n is the energy of the nth frame, x(n) is the frame signal, and ω(n) is the Hamming window; then, divide the frames into four categories according to the magnitude of the frame signal energy value. The classification rules are as follows:[[]]

[0061]

[0062] where t represents the frame category, t = 1, 2, 3, 4; E t represents the range of the frame energy corresponding to the tth frame category.

[0063] Then calculate the time ratio of each category of frames in the total number of frames (here, the total number of frames refers to the total number of frames of a user in a certain emotional state):

[0064]

[0065] Take l t as the time dimension feature, where Time tThe time length of the t-th type of frame of a certain user in the q-th emotional state for statistics; select the short-time average frequency a1, short-time mean square deviation frequency a2, and formant frequency a3 of each user in the q-th emotional state as the prosodic features of each segment of sound data, and fuse the three features to obtain the fused prosodic feature m. The formula is as follows:

[0066] m = w1·a1 + w2·a2 + w3·a3

[0067] Among them, w1, w2, and w3 respectively represent the relative importance of the features corresponding to a1, a2, and a3.

[0068] For Q emotional states, the time proportion l of the t-th frame category of the user in the Q emotional states t , and the fused prosodic feature m of the user in the q-th emotional state form the three-dimensional sound feature dataset of the user.

[0069] Step 2: Construct a clustering multi-model training method based on three-dimensional sound features. The developer obtains sample users of different personnel categories through a clustering method based on three-dimensional sound features, and then uses the sample user data of different personnel categories to train a Gaussian mixture model to form a voice emotion recognition model for different personnel categories.

[0070] Step 2.1, first, use all three-dimensional sound feature datasets as samples, and randomly select K samples as clustering centers in the user three-dimensional sound feature dataset according to the personnel category; then, calculate the distance from the i-th sample to the k-th clustering center. The formula is as follows:

[0071]

[0072] Among them, q is the q-th emotion, l iqt is the time proportion of the t-th frame category of the user corresponding to the i-th sample in the q-th emotion, l kqt is the time proportion of the t-th frame category of the user corresponding to the k-th clustering center in the q-th emotion, m iq is the fused prosodic feature of the user corresponding to the i-th sample in the q-th emotion, m kq is the fused prosodic feature of the k-th clustering center in the q-th emotion, l iqs is the short-time average energy of the sound data of the i-th sample user in the q-th emotion; then evaluate the distance of each sample user to the clustering center and assign it to the cluster to which the nearest center belongs. After the calculation, update the clustering centers of each cluster again, and iterate until the number of iterations is greater than the preset number of times to obtain a set of clustering centers {o(1), o(2),... o(K)} corresponding to K personnel categories.

[0073] Step 2.2: For the K clustering centers obtained after clustering, K voice emotion recognition models (i.e., Gaussian mixture models) are respectively formed using the Gaussian mixture model, denoted as {G1, G2,..., G K}, and the calculation process is as follows:

[0074] First, use the K-means algorithm to initialize a set of parameters, and the formula is as follows:

[0075] θ d = {α d , μ d , σ d}

[0076] Among them, D is the number of Gaussian sub-models, θ d is the parameter of the d-th Gaussian sub-model, μ d is the mean of the d-th Gaussian sub-model, σ d is the variance of the d-th Gaussian sub-model, α d is the standard deviation of the d-th Gaussian sub-model, d = 1, 2,..., D; then calculate the probability that each data (the data is θ d ) comes from sub-model d, and the formula is as follows:

[0077]

[0078] Among them, γ d is the probability that the data comes from sub-model d, Φ(x|θ d ) is the Gaussian density function of the d-th Gaussian sub-model; Φ(x q |θ d ) is the Gaussian density function obtained when the input feature vector x q of the d-th Gaussian sub-model; then find the new parameter values according to the probability Replace the old parameters with the new parameters and start iteration. When is less than the threshold, terminate the iteration. Finally, obtain the voice emotion probability density function (i.e., Gaussian mixture model), and the formula is as follows:

[0079]

[0080] Among them is the probability of the q-th emotion output by the k-th Gaussian mixture model; repeat the above steps for the sample user data of each personnel category, and finally obtain the voice emotion recognition models {G1, G2,..., G K} corresponding to K personnel categories.

[0081] Step 3: Construct the initialization classification step. The user inputs the voice in their normal emotional state during initialization for initialization classification to obtain a general benchmark model and benchmark parameters.

[0082] The user inputs the voice in the normal emotional state during initialization, and a general reference model and reference parameters are obtained according to the distance function. The formula is as follows:

[0083]

[0084] k * =argmin(d 1 (∧,o(k)))

[0085] d * =d 1 (∧,o(k * ))

[0086] Where the first emotional state is set as the normal emotional state, ∧ represents the user, l ∧1t is the time proportion of the t-th frame category in the normal emotional state when the user initializes, l o(k)1t is the time proportion of the t-th frame category in the normal emotional state of the k-th cluster center, l ∧1s is the short-time average energy of the voice data in the normal emotional state when the user initializes, m ∧1 is the fused prosody feature in the normal emotional state when the user initializes, m o(k)1 is the fused prosody feature of the k-th cluster center in the normal emotional state, k * is the optimal person category corresponding to the voice input by the user in the normal emotional state during initialization (that is, the cluster center with the shortest distance from the voice input by the user in the normal emotional state during initialization), d * is the distance from the voice input by the user in the normal emotional state during initialization to the optimal person category, k * The corresponding Gaussian mixture model is recorded as the reference model The voice of the user in the normal emotional state during initialization enters to obtain the probability of the normal emotional state under this model Thus, the performance index of the reference model is

[0087] Step 4: Construct the post-fusion step of multiple models. During the running recognition stage, the voice of the user collected in real time is input. After the voice sample passes through the reference model, it is then input into other models for multi-modal information fusion and judgment, and finally the recognition result is output.

[0088] During the actual use of the user, the real-time voice of the user is continuously collected. During the recognition stage, the voice data collected in real time is first input into the reference model to obtain the probabilities of Q emotions Among them, the maximum probability is recorded as The corresponding emotional state is recorded as class q* .

[0089] Assume that the current sound corresponds to the q-th emotional state, then calculate the shortest distance from the current sound data to all cluster centers, and obtain the model to which it belongs under the conditional probability. The calculation formula is as follows:

[0090]

[0091]

[0092]

[0093] where ∧ c represents the current user, is the time proportion of the t-th frame category of the current user in the q-th emotional state, l o(k)qt is the time proportion of the t-th frame category of the k-th cluster center in the q-th emotional state, is the short-time average energy of the sound data of the current user in the q-th emotional state, is the fused prosody feature of the current user in the q-th emotional state, m o(k)q is the fused prosody feature of the k-th cluster center in the q-th emotional state; is the optimal personnel category to which the current sound numerical control belongs under the conditional probability, is the distance to the optimal personnel category, The corresponding Gaussian mixture model is Repeat the above steps to obtain the set of optimal personnel categories corresponding to Q emotional probabilities distance set and Gaussian mixture model set

[0094] Then input the current sound data into the model to obtain the probabilities of Q emotions Take the maximum probability among them and record it as The emotional state corresponding to the maximum probability is recorded as and calculate the performance index of the model The formula is as follows: The formula is as follows:

[0095]

[0096] Repeat the above steps to obtain the set of performance indices

[0097] Finally, calculate the maximum value of the performance index, and take the corresponding emotional state as the result output. The formula is as follows:

[0098]

[0099]

[0100] Among them, OP * is the maximum value of the performance index. If OP * is one of them, for example then the emotional state corresponding to the maximum probability output by the corresponding Gaussian mixture model is recorded as the emotional state of the current user; if then the emotional state corresponding to the maximum probability output by the initial reference model

[0101] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention does not further explain various possible combination methods.

Claims

1. A method for driver voice emotion recognition for intelligent driving, characterized in that, Specifically, it includes the following steps: Step 1: Collect voice data of different users in Q different emotional states in a driving scenario, and construct a three-dimensional voice feature dataset for each user; Step 2: Use the clustering method to find K clustering centers {o(1), o(2),..., o(k),..., o(K)} in the constructed three-dimensional sound feature dataset, where o(k) represents the k-th clustering center, and k = 1, 2,..., K; Use the K clustering centers to train the Gaussian mixture model to obtain K Gaussian mixture models {G1, G2,..., G k ,...., G K}, where G k represents the k-th Gaussian mixture model; Step 3: The user inputs the voice in the normal emotional state during initialization. According to the distance between the voice and each cluster center, the optimal cluster center k corresponding to the voice input by the user in the normal emotional state during initialization is obtained. * ; Take the Gaussian mixture model corresponding to k * as the reference model Calculate of the performance metrics Step 4: During driving, the sound of the user is collected in real time, and the currently collected sound is input into G k* to obtain the probabilities of Q emotional states where q represents the q-th emotional state, q = 1, 2, …, Q, denotes the probability of the q-th emotional state output; denote the emotional state corresponding to the maximum probability as class q* ; Step 5: Calculate the shortest distance between the currently collected sound in the q-th emotional state and all cluster centers; Denote the cluster center corresponding to the shortest distance as The corresponding Gaussian mixture model is denoted as Obtain the set of cluster centers corresponding to each emotional state And the set of Gaussian mixture models corresponding to each cluster center Input the currently collected sound into to obtain The probabilities of Q emotional states output The maximum probability is denoted as The emotional state corresponding to the maximum probability is denoted as And calculate The performance metrics of Step 6: According to the performance indicators and the performance indicators judge the emotional state of the current user; Specifically, Step 1 is: frame and window the voice data of each user in the q-th emotional state; and calculate the energy contained in each frame according to the following formula: E n = x(n) 2 * ω(n) 2 where E n is the energy of the n-th frame, x(n) is the frame signal, and ω(n) is the Hamming window; Classify all frame signals of each user in the q-th emotional state according to the following formula: where t represents the frame category, t = 1, 2, 3, 4; E t represents the range of frame energy corresponding to the t-th frame category; Calculate the time proportion \(l\) of each user in the \(q\)th emotional state for the \(t\)th frame category t ; Among them, Time t represents the total duration of the t-th frame category in the q-th emotional state of the user; Calculate the short-time average frequency a1, short-time mean square deviation frequency a2, and formant frequency a3 of each user in the q-th emotional state to obtain the fused prosody feature m of the user in the q-th emotional state: m = w1·a1 + w2·a2 + w3·a3 where w1, w2, and w3 all represent relative importance; The time ratio l of the Q emotional states and the t-th frame category when the user is in the Q emotional states t , and the fused prosody feature m of the user in the q-th emotional state are combined to form the three-dimensional voice feature dataset of the user; In step 3, the performance index is calculated according to the following formula Among them, represents the probability of the sound in the normal emotional state when the user initializes entering to obtain the probability of the normal emotional state, d * represents the distance between the sound data input by the user when initializing and the optimal clustering center k * ; In step 5, calculate the performance metrics of the Gaussian mixture model set according to the following formula ​ Among them, represents the distance between the voice of the current user in the q-th emotional state and the cluster center .

2. The method for recognizing the voice emotion of a driver for intelligent driving according to claim 1, characterized in that The method for finding K clustering centers in Step 2 is specifically: first, use all the three-dimensional voice feature datasets as samples; randomly select K samples from all the samples as clustering centers, and then calculate the distance d(i,k) from the i-th sample to the k-th clustering center according to the following formula: where \(l\) iqt is the time ratio of the \(t\)-th frame category of the user corresponding to the \(i\)-th sample in the \(q\)-th emotional state, and \(l\) kqt is the time ratio of the \(t\)-th frame category of the user corresponding to the \(k\)-th cluster center in the \(q\)-th emotional state; \(m\) iq is the fused prosodic feature of the user corresponding to the \(i\)-th sample in the \(q\)-th emotional state, and \(m\) kq is the fused prosodic feature of the user corresponding to the \(k\)-th cluster center in the \(q\)-th emotional state, and \(l\) iqs is the short-time average energy of the voice data of the user corresponding to the \(i\)-th sample in the \(q\)-th emotional state; Evaluate the distance from each sample to the clustering center, and assign the sample to the cluster to which the clustering center with the closest distance to the sample belongs, then update the clustering centers of each cluster to obtain new clustering centers, and calculate the distance between each sample and the new clustering centers until the number of iterations is greater than the preset number of times, and finally obtain the set {o(1), o(2),... o(K)} of the final K clustering centers.

3. The method for driver voice emotion recognition for intelligent driving according to claim 1, characterized in that In Step 2, use the K clustering centers to train Gaussian mixture models to obtain K Gaussian mixture models corresponding to the K clustering centers, specifically: Initialize a set of parameters using the K-means algorithm according to the following formula: θ d = {α d , μ d , σ d} where D is the number of Gaussian sub-models, and θ d is the parameter of the d-th Gaussian sub-model, d = 1, 2, …, D, μ d is the mean of the d-th Gaussian sub-model, and σ d is the variance of the d-th Gaussian sub-model, and α d is the standard deviation of the d-th Gaussian sub-model; Calculate θ d Probability γ from the d-th Gaussian sub-model d : where, x q is the three-dimensional sound feature vector of the q-th emotional state, and Φ(x q |θ d ) is the input feature vector of the d-th Gaussian sub-model when the input feature vector is x q to obtain the Gaussian density function; Then, according to the probability γ d , calculate the new parameter values Calculate the value of. If is less than the threshold, terminate the iterative calculation; otherwise, replace the old parameters with the new parameters and start the iteration until the absolute value of the difference between the parameter value obtained in the last iteration and the parameter value obtained in the previous iteration is less than the preset threshold; finally, obtain the k-th Gaussian mixture model G k The expression of is as follows:

4. The method for driver voice emotion recognition for intelligent driving according to claim 1, characterized in that In step 3, the distance d between the voice input by the user in the normal emotional state during initialization and each cluster center is calculated according to the following formula 1 (∧, o(k)): where ∧ represents the user, l ∧1t is the time ratio of the t-th frame category in the normal emotional state when the user is initialized, l o(k)1t is the time ratio of the t-th frame category in the normal emotional state for the k-th cluster center, l ∧1s is the short-time average energy of the voice data in the normal emotional state when the user is initialized, m ∧1 is the fused prosody feature in the normal emotional state when the user is initialized, m o(k)1 is the fused prosody feature in the normal emotional state for the k-th cluster center.

5. The method for driver voice emotion recognition for intelligent driving according to claim 1, wherein The specific content of step 5 is as follows: Assume that the currently collected voice of the user corresponds to the q-th emotional state, and calculate the distances d from the current voice data to all cluster centers according to the following formula q (∧ c , o(k)): Among them, ∧ c represents the current user, is the time ratio of the t-th frame category of the current user in the q-th emotional state, l o(k)qt is the time ratio of the t-th frame category of the k-th cluster center in the q-th emotional state, is the short-time average energy of the voice data of the current user in the q-th emotional state, is the fusion prosody feature of the current user in the q-th emotional state, m o(k)q is the fusion prosody feature of the k-th cluster center in the q-th emotional state.

6. The method for identifying the voice emotion of a driver for intelligent driving according to claim 1, wherein Specifically, in Step 6: The Gaussian mixture model set The performance metrics of each high-speed mixture model in the set form a performance set Among and the performance set select the largest performance metric and denote it as OP * ; judge the emotional state q of the current user according to the following formula end :

Citation Information

Patent Citations

  • Voiceprint identification method based on Gauss mixing model and system thereof

    CN102324232A

  • Layered speaker recognition method based on model clustering

    CN109961794A