A method and system for recognizing speech emotion
By extracting and analyzing the time domain, frequency domain and channel domain features in the voice collection information, and combining the emotional tone definition library, we can identify the emotional voice information in the car-free taxi, which solves the problem that the existing system cannot accurately recognize voice emotions and realizes the recognition of the true intention of the voice content.
Patent Information
- Application Number
- CN202411568954.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-05
AI Technical Summary
The existing car-free taxi voice recognition system cannot accurately identify voice information with emotion, resulting in the inability to express the true intention of the voice content.
By obtaining the user's voice collection information, including time domain feature information, frequency domain feature information and channel domain feature information, the sound amplitude, tone characteristics and clearness rate are extracted, the comprehensive correlation value feature value is calculated, and it is compared with the preset emotional tone definition library to extract emotional tone keywords to identify voice emotions.
Accurate recognition of voice emotions is achieved, and the emotional characteristics of the voice in the process of telling can be obtained, thereby obtaining the true intention of the voice content, and improving the naturalness and intelligence of human-computer interaction.
Smart Images

Figure CN119296587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular to a method and system for recognizing speech emotions. Background Art
[0002] With the rapid progress of computer network communication and multimedia technology, human-computer interaction technology has gradually become an important part of the field of artificial intelligence. Especially with the rise of digital entertainment, the popularization of smart home appliances, driverless taxis, and the increasing widespread use of computer applications, the naturalness and intelligence of human-computer interaction have become particularly crucial. In this context, the processing of emotional information has become an important research direction for enhancing human-computer interaction capabilities. Therefore, how to enhance the adaptability of speech emotion recognition technology to users' emotional changes, enabling the computer to accurately recognize speech information with emotions, thereby improving the naturalness and intelligence of human-computer interaction, has become increasingly urgent;
[0003] As driverless taxis become more and more popular, existing driverless taxi speech recognition usually can only convert speech into text and then recognize the meaning of the language through the text. However, speech has emotional characteristics during the process of narration, which leads to the inability to express the true intention of the speech content by simply recognizing speech. Therefore, a method and system for recognizing speech emotions are needed to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for recognizing speech emotions to solve the technical problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A method for recognizing speech emotions, including:
[0007] Obtain the speech acquisition information of the user, wherein obtaining the speech acquisition information includes time-domain feature information, frequency-domain feature information, and channel-domain feature information;
[0008] Obtain the sound amplitude information and the sound zero-crossing rate according to the time-domain feature information, and obtain the sound intensity according to the sound amplitude information;
[0009] Obtain the sound resonance peak value and the frequency band according to the frequency-domain feature information, perform envelope extraction on the frequency band to obtain a plurality of envelope instantaneous curve values, and calculate the timbre feature value within a preset time according to the oscillation of the plurality of envelope instantaneous curve values;
[0010] Obtain a plurality of sound transmission rates according to the channel-domain feature information, obtain the sound transmission rate feature information according to the plurality of sound transmission rates, and obtain the voicelessness / voicedness ratio according to the sound transmission rate feature information and the sound zero-crossing rate;
[0011] Obtain a comprehensive correlation value eigenvalue based on the sound intensity, timbre eigenvalue, and voicelessness rate;
[0012] Extract an emotional timbre keyword according to the comprehensive correlation value eigenvalue, and compare the emotional timbre keyword with a preset emotional timbre definition library to obtain an emotional timbre definition.
[0013] Preferably, the step of obtaining the sound amplitude information and the sound zero-crossing rate according to the time-domain characteristic information includes:
[0014] Obtain a plurality of first sound amplitudes corresponding to the time domain according to the time-domain characteristic information;
[0015] Obtain the corresponding average first sound amplitude according to the plurality of first sound amplitudes, and use the information corresponding to the average first sound amplitude as the sound amplitude information;
[0016] Obtain a plurality of time-domain sampling points according to the time-domain characteristic information, compare the plurality of time-domain sampling points in adjacent order to obtain a plurality of comparison information, wherein the plurality of time-domain sampling points reflect the active deviation degree of the signal through positive and negative responses;
[0017] Successively determine whether there are adjacent positive and negative signals in the plurality of comparison information;
[0018] If so, count the time-domain sampling points corresponding to the plurality of comparison information with adjacent positive and negative signals as the zero-crossing times;
[0019] Obtain the corresponding sampling times according to the time-domain sampling points, compare the zero-crossing times with the sampling times to obtain the zero-crossing passing rate, and use the zero-crossing passing rate as the sound zero-crossing rate.
[0020] Preferably, the step of calculating the timbre eigenvalue within a preset time by the plurality of envelope instantaneous curve values includes:
[0021] Obtain the starting acquisition moment and the ending acquisition moment according to the preset time;
[0022] Calculate the average envelope instantaneous curve value according to the plurality of envelope instantaneous curve values, the starting acquisition moment, and the ending acquisition moment, wherein the calculation formula is:
[0023]
[0024] Wherein, represents the average envelope instantaneous curve value, t 1 represents the starting acquisition moment, t 2 represents the ending acquisition moment, w(d) represents the envelope instantaneous curve value, and n represents the acquisition times of the envelope instantaneous curve value, where n = 1, 2, 3... n;
[0025] Calculate the standard envelope instantaneous curve value based on multiple said envelope instantaneous curve values and the average envelope instantaneous curve value, where the calculation formula is:
[0026]
[0027] where j(z) represents the standard envelope instantaneous curve value, and p(j) l represents the l-th collected envelope instantaneous curve value, N represents the number of times of collecting the envelope instantaneous curve value, and l represents the serial number of the envelope instantaneous curve value, represents the average envelope instantaneous curve value;
[0028] Calculate the envelope instantaneous coefficient based on the said standard envelope instantaneous curve value and the average envelope instantaneous curve value, where the calculation formula is:
[0029]
[0030] where X(s) represents the envelope instantaneous coefficient, represents the average envelope instantaneous curve value, and j(z) represents the standard envelope instantaneous curve value;
[0031] Obtain the timbre starting transient value according to the envelope instantaneous coefficient, and use the said timbre starting transient value as the timbre feature value.
[0032] Preferably, the step of obtaining the voiceless / voiced ratio according to the said sound transmission rate characteristic information and the sound zero-crossing rate includes:
[0033] Obtain the transmission rate spectrum according to the said sound transmission rate characteristic information;
[0034] Obtain multiple transmission peaks according to the transmission rate spectrum, and extract the multiple said transmission peaks based on the preset first appearance to obtain the first transmission peak;
[0035] Obtain the falling time when falling back to the preset standard transmission value according to the said first transmission peak, and use the said falling time as the fundamental frequency period;
[0036] Obtain the corresponding fundamental frequency according to the said fundamental frequency period;
[0037] Compare the fundamental frequency with the sound zero-crossing rate to obtain the fundamental frequency-zero crossing ratio, and use the said fundamental frequency-zero crossing ratio as the voiceless / voiced ratio.
[0038] Preferably, the step of obtaining the comprehensive correlation value feature value according to the said sound intensity, timbre feature value and voiceless / voiced ratio includes:
[0039] Obtain the corresponding first weight value according to the said sound intensity;
[0040] Calculate the comprehensive correlation value eigenvalue according to the first weight value, sound intensity, timbre eigenvalue, and voicelessness rate, where the calculation formula is:
[0041] z(h) = [P(Z)*a + P(q)*(1 - a)]*P(H);
[0042] Among them, z(h) represents the comprehensive correlation value eigenvalue, P(Z) represents the sound intensity, a represents the first weight value, P(q) represents the timbre eigenvalue, and P(H) represents the voicelessness rate.
[0043] Preferably, the step of comparing the emotional timbre keyword with the preset emotional timbre definition library to obtain the emotional timbre definition includes:
[0044] Encode the emotional timbre keyword to obtain an emotional timbre encoding value;
[0045] Encode each word in the preset emotional timbre definition library in turn to obtain multiple emotional timbre encoding values;
[0046] Traverse and match the emotional timbre encoding value with multiple emotional timbre encoding values based on the similarity model to obtain a timbre similarity value, where the function of the similarity model is:
[0047]
[0048] Among them, represents the timbre similarity value, j(z) represents the emotional timbre encoding value, Y(S) k represents the emotional timbre encoding value, k represents the serial number of the emotional timbre encoding value, k = 1, 2, 3...k;
[0049] Obtain the corresponding emotional timbre encoding value according to the timbre similarity value, and obtain the corresponding emotional timbre definition according to the emotional timbre encoding value.
[0050] This application also provides a speech emotion recognition system, including:
[0051] The first acquisition module is used to acquire the voice acquisition information of the user. Among them, acquiring the voice acquisition information includes time domain feature information, frequency domain feature information, and channel domain feature information;
[0052] The second acquisition module is used to obtain the sound amplitude information and the sound zero-crossing rate according to the time domain feature information, and obtain the sound intensity according to the sound amplitude information;
[0053] The third acquisition module is used to obtain the sound resonance peak value and the frequency band according to the frequency domain feature information, perform envelope extraction on the frequency band to obtain multiple envelope instantaneous curve values, and calculate the timbre eigenvalue within a preset time according to the oscillation of multiple envelope instantaneous curve values;
[0054] A fourth acquisition module, configured to acquire a plurality of voice transmission rates according to the channel domain feature information, acquire voice transmission rate feature information according to the plurality of voice transmission rates, and acquire a voiceless / voiced ratio according to the voice transmission rate feature information and the voice zero-crossing rate;
[0055] A fifth acquisition module, configured to acquire a comprehensive correlation value feature according to the voice intensity, the timbre feature value, and the voiceless / voiced ratio;
[0056] A first comparison module, configured to extract an emotional timbre keyword according to the comprehensive correlation value feature, and compare the emotional timbre keyword with a preset emotional timbre definition library to obtain an emotional timbre definition.
[0057] Preferably, the second acquisition module includes:
[0058] A first acquisition unit, configured to acquire a plurality of first voice amplitudes in the corresponding time domain according to the time domain feature information;
[0059] A second acquisition unit, configured to acquire a corresponding average first voice amplitude according to the plurality of first voice amplitudes, and use the information corresponding to the average first voice amplitude as voice amplitude information;
[0060] A third acquisition unit, configured to acquire a plurality of time domain sampling points according to the time domain feature information, compare the plurality of time domain sampling points in adjacent order to obtain a plurality of comparison information, wherein the plurality of time domain sampling points reflect the active deviation degree of the signal through positive and negative;
[0061] A first judgment unit, configured to sequentially judge whether there are adjacent positive and negative signals in the plurality of comparison information;
[0062] If so, count the time domain sampling points corresponding to the plurality of comparison information with adjacent positive and negative signals as the zero-crossing times;
[0063] A fourth acquisition unit, configured to acquire a corresponding sampling number according to the time domain sampling points, compare the zero-crossing times with the sampling number to obtain a zero-crossing passing rate, and use the zero-crossing passing rate as the voice zero-crossing rate.
[0064] The present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0065] The present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0066] The beneficial effects of the present application are as follows: The present invention obtains the voice acquisition information of the user. Among them, obtaining the voice acquisition information includes time-domain feature information, frequency-domain feature information, and channel-domain feature information. Then, the sound amplitude information and the sound zero-crossing rate are obtained according to the time-domain feature information, and the sound intensity is obtained according to the sound amplitude information. Then, the sound resonance peak value and the frequency band are obtained according to the frequency-domain feature information, and the envelope of the frequency band is extracted to obtain a plurality of envelope instantaneous curve values. The timbre feature value within a preset time is calculated by oscillating according to the plurality of envelope instantaneous curve values. Then, a plurality of sound transmission rates are obtained according to the channel-domain feature information, and the sound transmission rate feature information is obtained according to the plurality of sound transmission rates. The voicelessness / voicedness rate is obtained according to the sound transmission rate feature information and the sound zero-crossing rate. Finally, the comprehensive correlation value feature value is obtained according to the sound intensity, the timbre feature value, and the voicelessness / voicedness rate. Finally, the emotional timbre keyword is extracted according to the comprehensive correlation value feature value, and the emotional timbre keyword is compared with the preset emotional timbre definition library to obtain the emotional timbre definition. In this way, by comparing the preset emotional timbre definition library, the keyword is mapped to a specific emotional category, such as happy, sad, angry, etc., so as to realize that the emotion recognition can know the emotional characteristics of the voice during the narration process. In this way, the true intention of the voice content can be obtained through the emotional timbre definition. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present application.
[0068] Figure 2 It is a schematic structural diagram of the system according to an embodiment of the present application.
[0069] Figure 3 It is a schematic internal structure diagram of a computer device according to an embodiment of the present application.
[0070] The realization, functional characteristics, and advantages of the purpose of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0072] As Figures 1-3 shown, the present application provides a method for recognizing voice emotion, including:
[0073] S1. Obtain the voice acquisition information of the user. Among them, obtaining the voice acquisition information includes time-domain feature information, frequency-domain feature information, and channel-domain feature information;
[0074] S2. Obtain the sound amplitude information and the sound zero-crossing rate according to the time-domain feature information, and obtain the sound intensity according to the sound amplitude information;
[0075] S3. Obtain the sound resonance peak value and frequency band according to the frequency domain characteristic information, perform envelope extraction on the frequency band to obtain a plurality of envelope instantaneous curve values, and calculate the timbre characteristic value within a preset time by oscillating according to the plurality of envelope instantaneous curve values;
[0076] S4. Obtain a plurality of sound transmission rates according to the channel domain characteristic information, obtain the sound transmission rate characteristic information according to the plurality of sound transmission rates, and obtain the voiceless / voiced ratio according to the sound transmission rate characteristic information and the sound zero-crossing rate;
[0077] S5. Obtain the comprehensive correlation value characteristic value according to the sound intensity, timbre characteristic value and voiceless / voiced ratio;
[0078] S6. Extract the emotional timbre keyword according to the comprehensive correlation value characteristic value, and compare the emotional timbre keyword with a preset emotional timbre definition library to obtain the emotional timbre definition.
[0079] As described in the above steps S1 - S6, as driverless taxis become more and more popular, the existing voice recognition for driverless taxis usually only converts voice into text and then recognizes the meaning of the language through the text. However, the voice has emotional characteristics during the speaking process, which leads to the inability to express the true intention of the voice content by simply recognizing the voice. At the same time, since voice emotion recognition infers the emotional state of the speaker by analyzing various features in the voice signal, time-domain features, frequency-domain features, and channel-domain features capture information of the voice signal from different perspectives, and these information together reflect the emotional state of the speaker. Therefore, the present invention first obtains the voice acquisition information of the user, where obtaining the voice acquisition information includes time-domain feature information, frequency-domain feature information, and channel-domain feature information, which can ensure the acquisition of the features of the voice signal from multiple perspectives (time domain, frequency domain, channel domain), providing a rich data basis for subsequent feature extraction and emotion recognition. At the same time, multi-dimensional data helps to improve the accuracy and robustness of feature extraction. Then, according to the time-domain feature information, the sound amplitude information and the sound zero-crossing rate are obtained, and the sound intensity is obtained according to the sound amplitude information. In this way, the instantaneous amplitude of the sound is reflected by the amplitude information, which is used to evaluate the loudness of the sound, and the zero-crossing rate reflects the activity of the sound, which is used to distinguish between voiceless and voiced sounds. Furthermore, the sound intensity can reflect the comprehensive amplitude information, providing a quantitative sound loudness index, which helps to judge the emotional state of the user (such as excited, calm, etc.). Then, according to the frequency-domain feature information, the sound resonance peak value and the frequency band are obtained. The sound resonance peak value reflects the resonance frequency of the sound, and the frequency band provides the energy distribution of the sound in different frequency ranges, which helps to analyze the complexity of the timbre, helps to distinguish different timbres, and envelope extraction is performed on the frequency band to obtain multiple envelope instantaneous curve values, which reflect the amplitude change of the sound over time, helping to capture the dynamic characteristics of the timbre, and the timbre feature value within a preset time is calculated by oscillating the multiple envelope instantaneous curve values. In this way, the timbre feature value can reflect the comprehensive envelope instantaneous curve values, providing a quantitative sound timbre index, which helps to judge the emotional state of the user (such as happy, sad, etc.). Since the channel domain refers to the propagation rate of the physical path or medium that the signal experiences during transmission, then according to the channel-domain feature information, multiple sound transmission rates are obtained, and the sound transmission rate feature information is obtained according to the multiple sound transmission rates, and the voiceless / voiced rate is obtained according to the sound transmission rate feature information and the sound zero-crossing rate. In this way, by combining the transmission rate feature information and the zero-crossing rate, a quantitative sound voiceless / voiced degree index is provided, which helps to judge the emotional state of the user (such as angry, calm, etc.). Then, according to the sound intensity, the timbre feature value, and the voiceless / voiced rate, the comprehensive correlation value feature value is obtained. In this way, by integrating multiple features, a comprehensive emotion recognition index is provided, which helps to more accurately judge the emotional state of the user. Finally, according to the comprehensive correlation value feature value, the emotional timbre keywords are extracted, and the emotional timbre keywords are compared with the preset emotional timbre definition library.Obtain the emotional timbre definition, so that by comparing with the preset emotional timbre definition library, the keywords can be mapped to specific emotional categories, such as happy, sad, angry, etc., so as to realize that emotional recognition can obtain the emotional characteristics of the speech during the speaking process, and thus the true intention of the speech content can be obtained through the emotional timbre definition.
[0080] In one embodiment, step S2 of obtaining the sound amplitude information and the sound zero-crossing rate according to the time-domain feature information includes:
[0081] S201. Obtain a plurality of first sound amplitudes corresponding to the time domain according to the time-domain feature information;
[0082] S202. Obtain the corresponding average first sound amplitude according to the plurality of first sound amplitudes, and use the information corresponding to the average first sound amplitude as the sound amplitude information;
[0083] S203. Obtain a plurality of time-domain sampling points according to the time-domain feature information, compare the plurality of time-domain sampling points in adjacent order to obtain a plurality of comparison information, wherein the plurality of time-domain sampling points reflect the active deviation degree of the signal through positive and negative;
[0084] S204. Sequentially judge whether there are adjacent positive and negative signals in the plurality of comparison information;
[0085] If there are, count the time-domain sampling points corresponding to the plurality of comparison information with adjacent positive and negative signals as the zero-crossing times;
[0086] S205. Obtain the corresponding sampling times according to the time-domain sampling points, compare the zero-crossing times with the sampling times to obtain the zero-crossing passing rate, and use the zero-crossing passing rate as the sound zero-crossing rate.
[0087] As described in the above steps S201 - S205, the present invention first obtains a plurality of first sound amplitudes in the corresponding time domain according to the time domain feature information, then obtains the corresponding average first sound amplitude according to the plurality of first sound amplitudes, and takes the information corresponding to the average first sound amplitude as the sound amplitude information. In this way, the short - term fluctuations can be smoothed out through the average first sound amplitude, and a more stable amplitude index can be obtained. At the same time, the sound amplitude information can reflect the overall loudness of the sound, which helps to evaluate the intensity of the sound. Then, according to the time domain feature information, a plurality of time domain sampling points are obtained, and the plurality of time domain sampling points are compared in adjacent order to obtain a plurality of comparison information. Among them, the plurality of time domain sampling points reflect the active deviation degree of the signal through positive and negative responses. Among them, the zero - crossing rate is for each sampling point, checking whether the signs of its two adjacent sampling points before and after are opposite, and then judging that if the signs are opposite, it is considered that there is a zero - crossing between these two sampling points. On this basis, it can be successively judged whether there are adjacent positive and negative signals in the plurality of comparison information. If there are, the time domain sampling points corresponding to the plurality of comparison information with adjacent positive and negative signals are counted as the zero - point times. Then, the corresponding sampling times are obtained according to the time domain sampling points, and the zero - point times are compared with the sampling times to obtain the zero - point passing rate, and the zero - point passing rate is used as the sound zero - crossing rate. And the zero - point times reflect the active degree of the signal. The zero - point times of high - frequency signals are usually more, and it can also provide a basis for the subsequent calculation of the voiceless - voiced ratio.
[0088] In one embodiment, step S3 of calculating the timbre feature value within a preset time by oscillating the plurality of envelope instantaneous curve values includes:
[0089] S301. Obtain the starting acquisition moment and the ending acquisition moment according to the preset time;
[0090] S302. Calculate the average envelope instantaneous curve value according to the plurality of envelope instantaneous curve values, the starting acquisition moment, and the ending acquisition moment. The calculation formula is:
[0091]
[0092] Wherein, represents the average envelope instantaneous curve value, t 1 represents the starting acquisition moment, t 2 represents the ending acquisition moment, w(d) represents the envelope instantaneous curve value, and n represents the acquisition times of the envelope instantaneous curve value, where n = 1, 2, 3... n;
[0093] S303. Calculate the standard envelope instantaneous curve value according to the plurality of envelope instantaneous curve values and the average envelope instantaneous curve value. The calculation formula is:
[0094]
[0095] Among them, j(z) represents the standard envelope instantaneous curve value, and p(j) l represents the l-th collected envelope instantaneous curve value, N represents the number of times of collecting the envelope instantaneous curve value, and l represents the serial number of the envelope instantaneous curve value. represents the average envelope instantaneous curve value;
[0096] S304. Calculate the envelope instantaneous coefficient according to the standard envelope instantaneous curve value and the average envelope instantaneous curve value, and the calculation formula is:
[0097]
[0098] Among them, X(s) represents the envelope instantaneous coefficient, represents the average envelope instantaneous curve value, and j(z) represents the standard envelope instantaneous curve value;
[0099] S305. Obtain the timbre starting transient value according to the envelope instantaneous coefficient, and use the timbre starting transient value as the timbre feature value.
[0100] As described in the above steps S301 - S305, the present invention first obtains the starting acquisition moment and the ending acquisition moment according to the preset time, and then calculates the average envelope instantaneous curve value according to the multiple envelope instantaneous curve values, the starting acquisition moment and the ending acquisition moment. Among them, the envelope instantaneous curve value refers to the envelope value of the signal at a certain moment. The envelope can be regarded as the instantaneous change of the signal amplitude, which reflects the amplitude change trend of the signal. In signal processing, envelope extraction is a common technique for analyzing the amplitude characteristics of the signal, and the envelope instantaneous curve value acquisition method is to perform Hilbert transform on the signal through the Hilbert transform method to obtain the analytic signal, and the modulus of the analytic signal is the envelope of the signal. Then, calculate the standard envelope instantaneous curve value according to the multiple envelope instantaneous curve values and the average envelope instantaneous curve value. In this way, the standard envelope instantaneous curve value reflects the degree of dispersion or volatility of the data. Then, calculate the envelope instantaneous coefficient according to the standard envelope instantaneous curve value and the average envelope instantaneous curve value. In this way, the envelope instantaneous coefficient reflects the relative volatility of the data. Finally, obtain the timbre starting transient value according to the envelope instantaneous coefficient. Among them, the timbre starting transient value refers to the envelope instantaneous value of the signal in the starting stage, which is usually used to describe the initial characteristics of the timbre. In this way, through the initial characteristics, crosstalk generated in subsequent sound production can be prevented from affecting the timbre acquisition. Furthermore, the timbre starting transient value can be used as the timbre feature value. In this way, the timbre starting transient value is deduced through the envelope instantaneous coefficient and used as the timbre feature value for further timbre analysis or recognition, and provides data support for the calculation of the comprehensive correlation value feature value.
[0101] In one embodiment, step S4 of obtaining the voiceless / voiced rate according to the sound transmission rate characteristic information and the sound zero-crossing rate includes:
[0102] S401. Obtain a transmission rate spectrum based on the voice transmission rate characteristic information;
[0103] S402. Obtain multiple transmission peaks based on the transmission rate spectrum, and extract the multiple transmission peaks based on a preset first occurrence to obtain a first transmission peak;
[0104] S403. Obtain the fall time when falling back to a preset standard transmission value based on the first transmission peak, and use the fall time as the fundamental frequency period;
[0105] S404. Obtain the corresponding fundamental frequency based on the fundamental frequency period;
[0106] S405. Compare the fundamental frequency with the voice zero-crossing rate to obtain a fundamental frequency-zero crossing ratio, and use the fundamental frequency-zero crossing ratio as the voiceless-voiced ratio.
[0107] As described in the above steps S401 - S405, the present invention first obtains a transmission rate spectrum based on the voice transmission rate characteristic information, so as to convert the transmission rate in the time domain into a frequency domain representation, and different frequency components can be seen more intuitively. Then, multiple transmission peaks are obtained based on the transmission rate spectrum, and the multiple transmission peaks are extracted based on a preset first occurrence to obtain a first transmission peak. Selecting the peak that appears first in this way can avoid the interference of subsequent peaks and ensure that the most significant periodic component is extracted. Then, the fall time when falling back to a preset standard transmission value is obtained based on the first transmission peak. In this way, the fall time reflects the time when the transmission rate falls back from the peak to the standard value, and this time interval usually corresponds to the period of the fundamental frequency. The fall time is used as the fundamental frequency period. Immediately afterwards, the corresponding fundamental frequency is obtained based on the fundamental frequency period. Finally, the fundamental frequency is compared with the voice zero-crossing rate to obtain a fundamental frequency-zero crossing ratio, and the fundamental frequency-zero crossing ratio is used as the voiceless-voiced ratio. In this way, the ratio of the fundamental frequency to the zero-crossing rate can reflect the periodic characteristics and activity level of the voice, which helps to distinguish between voiceless and voiced sounds. At the same time, the voiceless-voiced ratio is a comprehensive index for judging the voiceless-voiced degree of the voice. Usually, a signal with a higher fundamental frequency and a lower zero-crossing rate is considered a voiced sound, and vice versa is a voiceless sound. Therefore, the voiceless-voiced ratio is of great significance in applications such as speech recognition, sentiment analysis, and speech synthesis.
[0108] In one embodiment, step S5 of obtaining the comprehensive correlation value characteristic value according to the voice intensity, timbre characteristic value, and voiceless-voiced ratio includes:
[0109] S501. Obtain the corresponding first weight value according to the voice intensity;
[0110] S502. Calculate the comprehensive correlation value characteristic value according to the first weight value, voice intensity, timbre characteristic value, and voiceless-voiced ratio, where the calculation formula is:
[0111] z(h) = [P(Z) * a + P(q) * (1 - a)] * P(H);
[0112] Among them, z(h) represents the comprehensive correlation value eigenvalue, P(Z) represents the sound intensity, a represents the first weight value, P(q) represents the timbre eigenvalue, and P(H) represents the voicing rate.
[0113] As described in the above steps S501 - S502, the present invention first obtains the corresponding first weight value according to the sound intensity, and then calculates the comprehensive correlation value eigenvalue according to the first weight value, sound intensity, timbre eigenvalue and voicing rate. In this way, the comprehensive correlation value eigenvalue can improve the richness and expression ability of the features, contribute to improving the accuracy and robustness of emotion recognition, and at the same time provide a basis for judging the definition of emotional timbre.
[0114] In one embodiment, the step S6 of comparing the emotional timbre keyword with the preset emotional timbre definition library to obtain the emotional timbre definition includes:
[0115] S601. Encode the emotional timbre keyword to obtain an emotional timbre encoding value;
[0116] S602. Encode each word in the preset emotional timbre definition library in turn to obtain a plurality of emotional timbre encoding values;
[0117] S603. Traverse and match the emotional timbre encoding value with a plurality of emotional timbre encoding values based on a similarity model to obtain a timbre similarity value, where the function of the similarity model is:
[0118]
[0119] Among them, represents the timbre similarity value, j(z) represents the emotional timbre encoding value, Y(S) k represents the emotional timbre encoding value, and k represents the serial number of the emotional timbre encoding value, k = 1, 2, 3... k;
[0120] S604. Obtain the corresponding emotional timbre encoding value according to the timbre similarity value, and obtain the corresponding emotional timbre definition according to the emotional timbre encoding value.
[0121] As described in the above steps S601-S604, the present invention first encodes the emotional timbre keywords to obtain emotional timbre coding values, and then encodes each word in the preset emotional timbre definition library in turn to obtain multiple emotional timbre coding values, so that the emotional timbre keywords and the preset emotional timbre definition library can be standardized, and then the emotional timbre coding values are traversed and matched with multiple emotional timbre coding values based on a similarity model to obtain timbre similarity values, and finally, the corresponding emotional timbre coding values are obtained according to the timbre similarity values, and the corresponding emotional timbre definitions are obtained according to the emotional timbre coding values, so that by comparing the preset emotional timbre definition library, the keywords are mapped to specific emotional categories, such as happiness, sadness, anger, etc., so that emotion recognition can obtain the emotional characteristics of the voice in the narration process, so that the true intention of the voice content can be obtained through the emotional timbre definition.
[0122] The present application also provides a speech emotion recognition system, comprising:
[0123] The first acquisition module 1 is used to acquire the user's voice collection information, wherein the acquired voice collection information includes time domain feature information, frequency domain feature information and channel domain feature information;
[0124] A second acquisition module 2, used to acquire sound amplitude information and sound zero-crossing rate according to the time domain feature information, and acquire sound intensity according to the sound amplitude information;
[0125] The third acquisition module 3 is used to obtain the sound resonance peak and frequency band according to the frequency domain feature information, and perform envelope extraction on the frequency band to obtain multiple envelope instantaneous curve values, and calculate the timbre feature value within a preset time according to the oscillation of the multiple envelope instantaneous curve values;
[0126] A fourth acquisition module 4 is used to acquire multiple sound transmission rates according to the channel domain characteristic information, acquire sound transmission rate characteristic information according to the multiple sound transmission rates, and acquire clear and voiced rate according to the sound transmission rate characteristic information and the sound zero crossing rate;
[0127] A fifth acquisition module 5, used for acquiring a comprehensive correlation value characteristic value according to the sound intensity, timbre characteristic value and voicing rate;
[0128] The first comparison module 6 is used to extract emotional timbre keywords according to the comprehensive correlation value feature value, and compare the emotional timbre keywords with a preset emotional timbre definition library to obtain an emotional timbre definition.
[0129] In one embodiment, the second acquisition module includes:
[0130] A first acquisition unit, configured to acquire a plurality of first sound amplitudes in a corresponding time domain according to the time domain feature information;
[0131] A second acquisition unit, configured to obtain a corresponding average first sound amplitude according to the plurality of first sound amplitudes, and use the information corresponding to the average first sound amplitude as sound amplitude information;
[0132] A third acquisition unit, configured to obtain a plurality of time-domain sampling points according to the time-domain feature information, compare the plurality of time-domain sampling points in adjacent order to obtain a plurality of comparison information, wherein the plurality of time-domain sampling points reflect the active deviation degree of the signal through positive and negative;
[0133] A first determination unit, configured to sequentially determine whether there are adjacent positive and negative signals in the plurality of comparison information;
[0134] If so, count the time-domain sampling points corresponding to the plurality of comparison information with adjacent positive and negative signals as the zero-crossing times;
[0135] A fourth acquisition unit, configured to obtain a corresponding sampling number according to the time-domain sampling points, compare the zero-crossing times with the sampling number to obtain a zero-crossing passing rate, and use the zero-crossing passing rate as the sound zero-crossing rate.
[0136] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method for identifying speech emotion are implemented.
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and the embodiments can include non-volatile and / or volatile memories. The non-volatile memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0138] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, apparatus, article or method comprising such element.
[0139] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for recognizing speech emotion, characterized in that: include: Acquire the user's voice collection information, wherein the acquired voice collection information includes time domain feature information, frequency domain feature information and channel domain feature information; Acquire sound amplitude information and sound zero-crossing rate according to the time domain feature information, and acquire sound intensity according to the sound amplitude information; Acquire the sound resonance peak and frequency band according to the frequency domain feature information, perform envelope extraction on the frequency band to obtain a plurality of envelope instantaneous curve values, and calculate the timbre feature value within a preset time according to the oscillation of the plurality of envelope instantaneous curve values; The step of calculating the timbre characteristic value within a preset time according to the oscillation of the multiple envelope instantaneous curve values comprises: Get the start and end time of collection according to the preset time; The average envelope instantaneous curve value is calculated according to the multiple envelope instantaneous curve values, the starting acquisition time and the ending acquisition time, wherein the calculation formula is: in, represents the average envelope instantaneous curve value, t1 represents the start acquisition time, t2 represents the end acquisition time, w(d) represents the envelope instantaneous curve value, n represents the number of acquisitions of the envelope instantaneous curve value, where n=1, 2, 3...n; The standard envelope instantaneous curve value is calculated according to the plurality of envelope instantaneous curve values and the average envelope instantaneous curve value, wherein the calculation formula is: Among them, j(z) represents the instantaneous curve value of the standard envelope, p(j) l Indicates the lth acquisition of the envelope instantaneous curve value, N indicates the number of acquisitions of the envelope instantaneous curve value, l indicates the sequence number of the envelope instantaneous curve value, Indicates the average envelope instantaneous curve value; The envelope instantaneous coefficient is calculated according to the standard envelope instantaneous curve value and the average envelope instantaneous curve value, wherein the calculation formula is: Among them, X(s) represents the instantaneous coefficient of the envelope, represents the average envelope instantaneous curve value, j(z) represents the standard envelope instantaneous curve value; Acquire a timbre starting transient value according to the envelope transient coefficient, and use the timbre starting transient value as a timbre characteristic value; Acquire multiple sound transmission rates according to the channel domain characteristic information, acquire sound transmission rate characteristic information according to the multiple sound transmission rates, and acquire clear and voiced rate according to the sound transmission rate characteristic information and sound zero crossing rate; Obtaining a comprehensive correlation value characteristic value according to the sound intensity, timbre characteristic value and voicing rate; Emotional timbre keywords are extracted according to the comprehensive correlation value feature value, and the emotional timbre keywords are compared with a preset emotional timbre definition library to obtain an emotional timbre definition.
2. The method for recognizing speech emotion according to claim 1, characterized in that: The step of acquiring the sound amplitude information and the sound zero-crossing rate according to the time domain feature information comprises: Acquire a plurality of first sound amplitudes corresponding to the time domain according to the time domain feature information; Acquire a corresponding average first sound amplitude according to the multiple first sound amplitudes, and use information corresponding to the average first sound amplitude as sound amplitude information; Acquire multiple time domain sampling points according to the time domain feature information, compare the multiple time domain sampling points in adjacent order, and obtain multiple comparison information, wherein the multiple time domain sampling points reflect the active deviation degree of the positive and negative reaction signals; Determining in sequence whether there are adjacent positive and negative signals in the plurality of comparison information; If so, the time domain sampling points corresponding to the multiple comparison information of adjacent positive and negative signals are counted as zero points; The corresponding sampling times are obtained according to the time domain sampling points, and the zero point times are compared with the sampling times to obtain the zero point passing rate, and the zero point passing rate is used as the sound zero crossing rate.
3. The method for recognizing speech emotion according to claim 1, characterized in that: The step of obtaining the clarity and voicing rate according to the sound transmission rate characteristic information and the sound zero-crossing rate comprises: Acquire a transmission rate spectrum according to the sound transmission rate characteristic information; Acquire multiple transmission peaks according to the transmission rate spectrum, and extract the multiple transmission peaks based on a preset first appearance to obtain a first transmission peak; Acquire a fallback time to a preset standard transmission value according to the first transmission peak value, and use the fallback time as a baseband period; Acquire a corresponding fundamental frequency according to the fundamental frequency period; The fundamental frequency is compared with the sound zero-crossing rate to obtain a fundamental frequency-zero-crossing ratio, and the fundamental frequency-zero-crossing ratio is used as the voicing rate.
4. The method for recognizing speech emotion according to claim 1, characterized in that: The step of obtaining a comprehensive correlation value characteristic value according to the sound intensity, timbre characteristic value and voicing rate comprises: Acquire a corresponding first weight value according to the sound intensity; The comprehensive correlation value characteristic value is calculated according to the first weight value, sound intensity, timbre characteristic value and voicing rate, wherein the calculation formula is: z(h)=[P(Z)*a+P(q)*(1-a)]*P(H); Among them, z(h) represents the comprehensive correlation value characteristic value, P(Z) represents the sound intensity, a represents the first weight value, P(q) represents the timbre characteristic value, and P(H) represents the voicing rate.
5. The method for recognizing speech emotion according to claim 1, characterized in that: The step of comparing the emotional timbre keyword with a preset emotional timbre definition library to obtain an emotional timbre definition includes: Encoding the emotional timbre keywords to obtain emotional timbre coding values; Encode each word in the preset emotional timbre definition library in sequence to obtain a plurality of emotional timbre coding values; The emotional timbre coding value is traversed and matched with a plurality of emotional timbre coding values based on a similarity model to obtain a timbre similarity value, wherein the function of the similarity model is: in, represents the timbre similarity value, j(z) represents the emotional timbre encoding value, Y(S) k represents the emotional timbre coding value, k represents the sequence number of the emotional timbre coding value, k=1, 2, 3...k; The corresponding emotional timbre coding value is obtained according to the timbre similarity value, and the corresponding emotional timbre definition is obtained according to the emotional timbre coding value.
6. A speech emotion recognition system is used to execute the speech emotion recognition method according to any one of claims 1 to 5, characterized in that: include: A first acquisition module is used to acquire the user's voice collection information, wherein the acquired voice collection information includes time domain feature information, frequency domain feature information and channel domain feature information; A second acquisition module, configured to acquire sound amplitude information and sound zero-crossing rate according to the time domain feature information, and acquire sound intensity according to the sound amplitude information; A third acquisition module is used to obtain a sound resonance peak value and a frequency band according to the frequency domain feature information, and to perform envelope extraction on the frequency band to obtain a plurality of envelope instantaneous curve values, and to calculate a timbre feature value within a preset time according to the oscillation of the plurality of envelope instantaneous curve values; A fourth acquisition module, configured to acquire a plurality of sound transmission rates according to the channel domain characteristic information, acquire sound transmission rate characteristic information according to the plurality of sound transmission rates, and acquire a clear and voiced rate according to the sound transmission rate characteristic information and a sound zero-crossing rate; A fifth acquisition module, used for acquiring a comprehensive correlation value characteristic value according to the sound intensity, timbre characteristic value and voicing rate; The first comparison module is used to extract emotional timbre keywords according to the comprehensive correlation value feature value, and compare the emotional timbre keywords with a preset emotional timbre definition library to obtain an emotional timbre definition.
7. The speech emotion recognition system according to claim 6, characterized in that: The second acquisition module includes: A first acquisition unit, configured to acquire a plurality of first sound amplitudes in a corresponding time domain according to the time domain feature information; A second acquisition unit, configured to acquire a corresponding average first sound amplitude according to the plurality of first sound amplitudes, and use information corresponding to the average first sound amplitude as sound amplitude information; A third acquisition unit is used to acquire a plurality of time domain sampling points according to the time domain feature information, and compare the plurality of time domain sampling points in an adjacent order to obtain a plurality of comparison information, wherein the plurality of time domain sampling points are active deviation degrees of positive and negative reaction signals; A first judging unit, used for sequentially judging whether there are adjacent positive and negative signals in the plurality of comparison information; If so, the time domain sampling points corresponding to the multiple comparison information of adjacent positive and negative signals are counted as zero points; The fourth acquisition unit is used to acquire the corresponding sampling times according to the time domain sampling points, and compare the zero point times with the sampling times to obtain the zero point passing rate, and use the zero point passing rate as the sound zero crossing rate.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Formant envelope estimation method and device, voice processing method and device, storage medium and terminal
CN112397087A
Emotion recognition method and device based on multi-modal signal, equipment and storage medium
CN113420556A