Speech Emotion Recognition Method, Device and Electronic Device

By extracting the context-dependent features of local emotional features in the speech emotion recognition model, combining convolution and LSTM models, the problem of insufficient accuracy of low-level feature recognition in the prior art is solved, and a higher accuracy of speech emotion recognition is achieved.

CN114495987BActive Publication Date: 2025-07-08BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011159427.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-26
Publication Date
2025-07-08
Estimated Expiration
2040-10-26

AI Technical Summary

Technical Problem

In the prior art, the speech emotion recognition method based on neural networks can only be classified by extracting low-level features, resulting in poor recognition accuracy.

Method used

The emotion recognition model is used to extract local emotional features from the to-process speech, and through multiple parallel first feature extraction modules and a second feature extraction module, combining the convolution layer, normalization layer, activation function layer and pooling layer, the context-dependent features of local emotional features are extracted, and the LSTM model is used to learn long-term contextual connections, and finally classified through the Softmax function.

Benefits of technology

The accuracy of speech emotion recognition is improved, and by simultaneously referring to local and global features, a more comprehensive feature reference is provided, enhancing the accuracy and generalization ability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495987B_ABST
    Figure CN114495987B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, and electronic device for voice emotion recognition, including: obtaining a voice to be processed; inputting the voice to be processed into an emotion recognition model, so that the emotion recognition model extracts local emotion features of the voice to be processed from the voice to be processed, extracts context correlation features of the local emotion features from the local emotion features, and recognizes the voice emotion of the voice to be processed based on the local emotion features and the context correlation features of the local emotion features. In this way, the emotion recognition model can recognize the voice emotion based on the local emotion features of the voice to be processed and the context correlation features of the local emotion features. Compared with the local emotion features, the context correlation features can reflect the voice emotion of the voice from a global perspective. Therefore, when recognizing the voice emotion in this way, both local features and global features are referred to, and the referred features are more comprehensive, thereby improving the accuracy of voice emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular to a method, device and electronic device for speech emotion recognition. Background Art

[0002] With the rapid development of technology, automatic speech recognition services have gradually penetrated into all aspects of daily life. Automatic speech recognition usually converts speech content into text content with corresponding meanings. However, in addition to the text content, the speech content also includes volume, tone, speech emotion, etc. These contents will have a great impact on the understanding of the text content. For example, different speech emotions may lead to completely opposite understandings of the same text content.

[0003] In related technologies, speech emotion is usually recognized based on a neural network. First, the neural network extracts features from the input speech, and then classifies the extracted features to obtain the speech emotion corresponding to the input speech. However, the neural network in this method can only classify emotions through the extracted low-level features, resulting in poor accuracy of speech emotion recognition. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, device and electronic device for speech emotion recognition to improve the accuracy of speech emotion recognition.

[0005] In a first aspect, an embodiment of the present invention provides a method for speech emotion recognition. The method includes: obtaining a speech to be processed; inputting the speech to be processed into an emotion recognition model, and outputting a recognition result for indicating the speech emotion of the speech to be processed; wherein, the emotion recognition model is used for: extracting local emotion features of the speech to be processed from the speech to be processed; extracting context correlation features of the local emotion features from the local emotion features; and recognizing the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain a recognition result.

[0006] In an optional implementation manner, the emotion recognition model includes a plurality of first feature extraction modules connected in parallel and a second feature extraction module; the step of extracting local emotion features of the speech to be processed from the speech to be processed and extracting context correlation features of the local emotion features from the local emotion features includes: for each first feature extraction module, extracting local emotion features of the basic elements corresponding to the first feature extraction module in the speech to be processed; wherein, the speech to be processed includes a plurality of basic elements, and the plurality of basic elements include: words, phrases, tones and volumes in the speech to be processed; and through the second feature extraction module, performing feature extraction on the local emotion features corresponding to each basic element to obtain context correlation features of the local emotion features.

[0007] In an alternative embodiment, the above first feature extraction module includes a convolutional layer, a normalization layer, an activation function layer, and a pooling layer connected in sequence; wherein, the convolutional layer is used to extract features from the speech to be processed to obtain feature data; the normalization layer is used to perform normalization processing on the feature data to obtain normalized feature data; the activation function layer is used to perform a function transformation on the normalized feature data; and the pooling layer is used to perform a pooling operation on the feature data after the function transformation to obtain local emotion features.

[0008] In an alternative embodiment, the above normalization layer uses the BN method to perform normalization processing on the feature data; the above activation function layer uses the ELU function to perform a function transformation on the normalized feature data.

[0009] In an alternative embodiment, the step of identifying the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain an identification result includes: classifying the local emotion features and the context correlation features of the local emotion features under a preset plurality of emotion categories to obtain the emotion category to which the speech to be processed belongs; and determining the emotion category to which the speech to be processed belongs as the identification result.

[0010] In an alternative embodiment, the step of classifying the local emotion features and the context correlation features of the local emotion features under a preset plurality of emotion categories to obtain the emotion category to which the speech to be processed belongs includes: integrating the local emotion features and the context correlation features of the local emotion features to obtain an integration result; and classifying the integration result under a plurality of emotion categories through a preset classifier to obtain the emotion category to which the speech to be processed belongs.

[0011] In an alternative embodiment, the step of classifying the integration result under a plurality of emotion categories through a preset classifier to obtain the emotion category to which the speech to be processed belongs includes: calculating the probability value of the integration result under each emotion category through the Softmax function; and determining the emotion category corresponding to the maximum value of the probability value as the emotion category to which the speech to be processed belongs.

[0012] In a second aspect, an embodiment of the present invention provides a speech emotion recognition device, which includes: a speech acquisition module, configured to acquire the speech to be processed; an emotion recognition module, configured to input the speech to be processed into an emotion recognition model and output an identification result for indicating the speech emotion of the speech to be processed; wherein, the emotion recognition model is used to: extract local emotion features of the speech to be processed from the speech to be processed; extract context correlation features of the local emotion features from the local emotion features; and identify the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain an identification result.

[0013] In a third aspect, an embodiment of the present invention provides an electronic device, which includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the voice emotion recognition method described in any one of the foregoing embodiments.

[0014] In a fourth aspect, an embodiment of the present invention provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the voice emotion recognition method described in any one of the foregoing embodiments.

[0015] The embodiments of the present invention bring the following beneficial effects:

[0016] A voice emotion recognition method, device and electronic device provided by the present invention first obtain a voice to be processed, and then input the voice to be processed into an emotion recognition model, so that the emotion recognition model extracts local emotion features of the voice to be processed from the voice to be processed, extracts context correlation features of the local emotion features from the local emotion features, and based on the local emotion features and the context correlation features of the local emotion features, recognize the voice emotion of the voice to be processed, and output a recognition result for indicating the voice emotion of the voice to be processed. The emotion recognition model in this method can recognize the voice emotion based on the local emotion features of the voice to be processed and the context correlation features of the local emotion features. Compared with the local emotion features, the context correlation features can reflect the voice emotion of the voice from a more global perspective. Therefore, when recognizing the voice emotion, this method refers to both local features and global features, and the reference features are more comprehensive, thereby improving the accuracy of voice emotion recognition.

[0017] Other features and advantages of the present invention will be described in the following specification, or some features and advantages can be inferred from the specification or determined without doubt, or can be known by implementing the above technologies of the present invention.

[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and detailed descriptions are made in conjunction with the accompanying drawings as follows. Description of the Drawings

[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1Flow chart of a voice emotion recognition method provided by an embodiment of the present invention;

[0021] Figure 2 Flow chart of another voice emotion recognition method provided by an embodiment of the present invention;

[0022] Figure 3 Flow chart of another voice emotion recognition method provided by an embodiment of the present invention;

[0023] Figure 4 Structural schematic diagram of a voice emotion recognition device provided by an embodiment of the present invention;

[0024] Figure 5 Structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0025] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0027] With the rapid development of technology, automatic speech recognition services have gradually penetrated into all aspects of daily life. Automatic speech recognition usually converts speech content into text content with corresponding meanings. However, in addition to the text content, the speech content also includes other contents, such as volume, tone, and speech emotion. These other contents will have a great impact on the understanding of the text content. For example, different speech emotions may lead to completely opposite understandings of the same text content.

[0028] Voice emotion is a specific and intense psychological activity, and recognizing voice emotion is an important research direction in voice interaction. Generally, to recognize the emotional state of a speaker, it is necessary to extract discriminative language features from the voice that are independent of the speaker or lexical content. Generally speaking, there are two types of information in voice: linguistic information and paralinguistic information. Among them, linguistic information refers to the context or meaning of the voice, and paralinguistic information refers to implicit information, such as the emotion contained in the voice. There are many unique acoustic features that can be used to identify the continuous features, qualitative features, and spectral features of voice emotion. In related technologies, many features have been studied to identify voice emotion, but so far, it has not been determined which type of feature is the best.

[0029] In related technologies, the methods for voice emotion recognition usually include two types. One is to manually extract emotion features for voice emotion recognition, and the other is to extract emotion features based on neural networks for voice emotion recognition. Usually, the classification accuracy of manually extracted emotion features is relatively high, but the extracted emotion features rely on professional knowledge and are time-consuming and laborious. Moreover, the emotion features extracted manually usually ignore high-level features, which are derived from low-level features, thus affecting the accuracy of voice emotion recognition.

[0030] When extracting emotion features based on neural networks, since machine automatic extraction of emotion features is used, the network has high requirements for the number of frames of the input voice data. When performing voice emotion recognition, the number of frames of the input voice data is not the same. Usually, in order to meet the input length requirements, after extracting the voice features, the unequal-length features are padded with zeros to the same length, and then voice emotion recognition is performed. This method loses part of the content of the original voice and the accuracy is not high. In addition, most neural networks can only extract low-level features to classify emotions, and most neural network models can only learn one emotion-related feature to recognize voice emotion, and the recognition accuracy is limited. The above neural networks include algorithm models based on LSTM (Long-Short Term Memory) or CNN (convnet, Convolutional Neural Network).

[0031] Based on the above problems, the embodiments of the present invention provide a method, device, and electronic device for voice emotion recognition. This technology can be applied to voice recognition scenarios, especially in voice emotion recognition scenarios. To facilitate the understanding of this embodiment, first, a method for voice emotion recognition disclosed in the embodiments of the present invention will be introduced in detail, as Figure 1 shown, the method includes the following specific steps:

[0032] Step S102, obtain the voice to be processed.

[0033] The above voice to be processed can be the words spoken by the user, and the voice to be processed contains content such as voice text, volume, tone, and voice emotion. In specific implementation, voice data can be collected through a voice acquisition device connected by communication (such as a smart speaker, a voice interaction device, etc.) or a voice recording device, or the voice to be processed can be obtained from a storage device of a voice file.

[0034] Step S104: Input the above voice to be processed into the emotion recognition model, and output a recognition result indicating the voice emotion of the voice to be processed.

[0035] The above emotion recognition model is used to: extract local emotion features of the voice to be processed from the voice to be processed; extract context correlation features of the local emotion features from the local emotion features; and recognize the voice emotion of the voice to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain a recognition result.

[0036] The above emotion recognition model can be a deep learning model or a neural network model, etc. The above emotion recognition model first extracts local emotion features from the voice to be processed. The local emotion features include basic voice recognition elements such as words, phrases, tones, and volumes, and belong to short-term semantic partial features. Then, according to the extracted local emotion features, emotional analysis is performed in connection with the context to obtain context correlation features of the local emotion features. Then, according to the local emotion features and the context correlation features of the local emotion features, classification prediction is performed on the voice to be processed to obtain the predicted voice emotion of the voice to be processed, that is, a recognition result indicating the voice emotion of the voice to be processed is obtained.

[0037] The voice emotion of the above voice to be processed includes one of a plurality of preset emotion categories. The plurality of preset emotion categories are set according to user needs and include anger, disgust, fear, happiness, sadness, surprise, etc., and are not specifically limited herein.

[0038] A voice emotion recognition method provided by an embodiment of the present invention first obtains a voice to be processed, and then inputs the voice to be processed into an emotion recognition model, so that the emotion recognition model extracts local emotion features of the voice to be processed from the voice to be processed, extracts context correlation features of the local emotion features from the local emotion features, and based on the local emotion features and the context correlation features of the local emotion features, recognizes the voice emotion of the voice to be processed, and outputs a recognition result indicating the voice emotion of the voice to be processed. The emotion recognition model in this method can recognize the voice emotion based on the local emotion features of the voice to be processed and the context correlation features of the local emotion features. Compared with the local emotion features, the context correlation features can reflect the voice emotion of the voice from a more global perspective. Therefore, when recognizing the voice emotion, this method refers to both local features and global features, and the reference features are more comprehensive, thereby improving the accuracy of voice emotion recognition.

[0039] Another voice emotion recognition method is also provided by an embodiment of the present invention, which is implemented on the basis of the method in the above embodiment; when the emotion recognition model includes a plurality of parallel first feature extraction modules and a second feature extraction module, this method focuses on describing the specific process of inputting the voice to be processed into the emotion recognition model and outputting a recognition result indicating the voice emotion of the voice to be processed (implemented through the following steps S204 - S210); as Figure 2 shown, this method includes the following specific steps:

[0040] Step S202, obtain the voice to be processed.

[0041] Step S204, input the above voice to be processed into an emotion recognition model, where the emotion recognition model includes a plurality of parallel first feature extraction modules and a second feature extraction module.

[0042] The above emotion recognition model includes a plurality of parallel first feature extraction modules and a second feature extraction module connected to each first feature extraction module. Each first feature extraction module is used to extract different local emotion features; the second feature extraction module is used to extract context correlation features of the local emotion features from the local emotion features extracted by the plurality of first feature extraction modules. In specific implementation, the number of first feature extraction modules in the emotion recognition model can be set according to user requirements, or determined according to the model calculation amount and model recognition accuracy. For example, three first feature extraction modules can be set, and the network parameters in each first feature extraction model can be the same or different.

[0043] Step S206: For each first feature extraction module, extract the local emotion features of the basic elements corresponding to the first feature extraction module in the speech to be processed; wherein, the speech to be processed includes multiple basic elements, and the multiple basic elements include: words, phrases, tones, and volumes in the speech to be processed.

[0044] In specific implementation, the basic elements corresponding to each first feature extraction module can be one or more, and the basic elements corresponding to different first feature extraction modules are different. For example, the emotion recognition model includes three first feature extraction modules. The first first feature extraction module is used to extract the features corresponding to the words and phrases in the speech to be processed, the second first feature extraction module is used to extract the features corresponding to the tones in the speech to be processed, and the third first feature extraction module is used to extract the features corresponding to the volumes in the speech to be processed.

[0045] In some embodiments, the above first feature extraction module includes a convolutional layer, a normalization layer, an activation function layer, and a pooling layer connected in sequence; wherein, the convolutional layer is used to extract features from the speech to be processed to obtain feature data; the normalization layer is used to perform normalization processing on the feature data to obtain the normalized feature data; the activation function layer is used to perform a function transformation on the normalized feature data; the pooling layer is used to perform a pooling operation on the feature data after the function transformation to obtain local emotion features.

[0046] The above convolutional layer and pooling layer are the core layers of the first feature extraction module. The convolutional layer has the characteristics of spatial local connectivity and shared weights, and these attributes allow the convolutional layer to perform the function of learning kernels. The above normalization layer uses the BN (Batch Normalization) method to perform normalization processing on the feature data. The BN method standardizes the activation of each batch of convolutional layers and improves the performance and stability of the deep network. The BN method keeps the average activation value at a level close to zero and keeps the activation standard deviation at a level close to 1.

[0047] In specific implementation, the convolution kernel of the above convolutional layer and the pooling kernel of the pooling layer can be one-dimensional, and the size of the convolution kernel, the size of the pooling kernel, and the step size can be set according to user needs. For example, the size of the convolution kernel is set to 3*3, the number of steps is 1, the size of the pooling kernel is set to 4*4, and the number of steps is 4. At the same time, the number of convolution kernels of multiple first feature extraction modules in the emotion recognition model can be determined according to user needs and recognition accuracy. Taking three first feature extraction modules as an example, the number of convolution kernels in the first and second first feature extraction modules can be 64, and the number of convolution kernels in the third first feature extraction module is 128.

[0048] The above activation function layer can use the ELU (Exponential Linear Units) function to perform a function transformation on the normalized feature data. Different from other activation functions, the ELU has negative values, which can make the activation average close to zero. Therefore, it can accelerate the learning speed in the designed network and improve the recognition accuracy. The above pooling layer can use max pooling, which can resist noise and distortion. It divides the input into a group of non-overlapping regions and outputs the maximum value of each such sub-region. Different configurations can also be made for the local feature learning block according to different tasks.

[0049] In specific implementation, first, the speech to be processed x(n) is input into the convolutional layer for convolutional processing to obtain convolutional features (equivalent to the above feature data). The formula for convolutional processing is:

[0050]

[0051] In the formula, n represents the length of the time series of the speech to be processed, w(n) represents the convolutional kernel, l represents the length of the convolutional kernel, and z(n) represents the convolutional features. The convolutional features are input into the BN layer (equivalent to the above normalization layer), which can keep the mean of the convolutional features close to zero and the variance of the convolutional features close to one. Then, the normalized features (equivalent to the above normalized feature data) are input into the ELU layer (equivalent to the above activation function layer), and the feature data after function transformation is output:

[0052] z1 = A·z(n) + β;

[0053]

[0054] Among them, the above β represents the learning parameter, which is a constant set by the user and is actually a value debugged according to experience; the above z1 represents the normalized feature data, and the above var represents the variance. After obtaining the normalized feature data z1, it is subjected to a function transformation through the ELU function to obtain the feature data after function transformation ELU(z1):

[0055]

[0056] Among them, α represents a constant, which can be an empirical value. After obtaining the feature data after function transformation, it is input into the pooling layer to perform a non-linear downsampling function and reduce the resolution of the elements. The local emotion features z2 generated by the max pooling of the feature data after function transformation by the pooling layer can be expressed as:

[0057]

[0058] Among them, Ω represents the pooling range, z pRepresents the p-th output feature in z1.

[0059] Step S208: Through the second feature extraction module, perform feature extraction on the local emotion features corresponding to each basic element to obtain the context correlation features of the local emotion features.

[0060] The above-mentioned second feature extraction module combines the time series of the speech to be processed and the multiple local feature emotion features extracted by the first feature extraction model to obtain the context correlation features of the local emotion features. In specific implementation, the above-mentioned second feature extraction module can adopt an LSTM model. This LSTM model is a recursive neural network architecture used to learn long-term context connection relationships. The LSTM is specifically designed to learn long-term dependencies from sequences. Therefore, placing it after multiple first feature extraction modules can learn context correlation from the extracted local emotion features to obtain the context correlation features of the local emotion features.

[0061] Step S210: Based on the local emotion features and the context correlation features of the local emotion features, identify the speech emotion of the speech to be processed to obtain the recognition result.

[0062] The above speech emotion recognition method first obtains the speech to be processed; then inputs the speech to be processed into an emotion recognition model. The emotion recognition model includes multiple parallel first feature extraction modules and a second feature extraction module. For each first feature extraction module, extract the local emotion features of the basic element corresponding to the first feature extraction module in the speech to be processed; then through the second feature extraction module, perform feature extraction on the local emotion features corresponding to each basic element to obtain the context correlation features of the local emotion features, and then based on the local emotion features and the context correlation features of the local emotion features, identify the speech emotion of the speech to be processed to obtain the recognition result. In this method, the designed network structure of the emotion recognition model not only has high emotion recognition accuracy but also has better generalization ability, thus providing guarantee for applications in the field of speech interaction.

[0063] An embodiment of the present invention also provides another speech emotion recognition method, which is implemented on the basis of the method in the above embodiment; this method focuses on describing the specific process of identifying the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain the recognition result (implemented through the following step S306); as Figure 3 shown, this method includes the following specific steps:

[0064] Step S302: Obtain the speech to be processed.

[0065] Step S304: Input the above-mentioned speech to be processed into the emotion recognition model, so that the emotion recognition model extracts the local emotion features of the speech to be processed from the speech to be processed, and extracts the context correlation features of the local emotion features from the local emotion features.

[0066] Step S306: Under a preset multiple emotion categories, classify the local emotion features and the context correlation features of the local emotion features to obtain the emotion category to which the speech to be processed belongs; determine the emotion category to which the speech to be processed belongs as the recognition result.

[0067] The above-mentioned preset multiple emotion categories are set according to user needs, including emotions such as anger, disgust, fear, happiness, sadness, and surprise. In specific implementation, the above step S306 can be implemented through the following steps 10-11:

[0068] Step 10: Integrate the local emotion features and the context correlation features of the local emotion features to obtain an integration result.

[0069] In specific implementation, the above local emotion features and the context correlation features of the local emotion features can be input into the fully connected layer of the emotion recognition model, and the result output by the fully connected layer is the integration result.

[0070] Step 11: Through a preset classifier, classify the above integration result under multiple emotion categories to obtain the emotion category to which the speech to be processed belongs.

[0071] The above preset classifier can be a Softmax function, can be a Bayesian classification, etc. The specific classifier adopted can be set according to user needs. As a preferred embodiment, a Softmax function can be used as the preset classifier. Through this Softmax function, calculate the probability value of the integration result under each emotion category; then determine the emotion category corresponding to the maximum value of the probability value as the emotion category to which the speech to be processed belongs.

[0072] In specific implementation, after obtaining the speech emotion of the speech to be processed, a preset intention can be matched according to the speech emotion. Usually, the classification labels of multiple speech emotions correspond to user intentions. For example, the intention corresponding to the happy emotion is the next action of the user after long-term training. When happy, the user likes to eat, read books and other actions; after recognizing the emotion recognition result, match it with the next intention, determine the interaction instruction, and then determine the corresponding interaction instruction according to the emotion recognition result and the intention information. After confirmation, send and execute the instruction.

[0073] In the above speech emotion recognition method, the emotion recognition model can extract the local emotion features and the context correlation features of the local emotion features from the speech to be processed, thereby improving the accuracy of emotion recognition.

[0074] Corresponding to the above method embodiments, an embodiment of the present invention provides a speech emotion recognition device, as Figure 4 shown, the device includes:

[0075] A speech acquisition module 40, configured to acquire the speech to be processed.

[0076] An emotion recognition module 41, configured to input the speech to be processed into an emotion recognition model, and output a recognition result for indicating the speech emotion of the speech to be processed; wherein, the emotion recognition model is used for: extracting local emotion features of the speech to be processed from the speech to be processed; extracting context correlation features of the local emotion features from the local emotion features; and recognizing the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain a recognition result.

[0077] For the above speech emotion recognition device, first, the speech to be processed is acquired, and then the speech to be processed is input into the emotion recognition model, so that the emotion recognition model extracts local emotion features of the speech to be processed from the speech to be processed, extracts context correlation features of the local emotion features from the local emotion features, and recognizes the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features, and outputs a recognition result for indicating the speech emotion of the speech to be processed. In this method, the emotion recognition model can recognize speech emotions based on local emotion features of the speech to be processed and context correlation features of the local emotion features. Compared with local emotion features, context correlation features can reflect the speech emotion of speech from a more global perspective. Therefore, when recognizing speech emotions, this method refers to both local features and global features, and the reference features are more comprehensive, thereby improving the accuracy of speech emotion recognition.

[0078] Specifically, the above emotion recognition model includes a plurality of first feature extraction modules connected in parallel and a second feature extraction module; the emotion recognition module 41 is further configured to: for each first feature extraction module, extract local emotion features of basic elements corresponding to the first feature extraction module in the speech to be processed; wherein, the speech to be processed includes a plurality of basic elements, and the plurality of basic elements include: words, phrases, tones, and volumes in the speech to be processed; and through the second feature extraction module, perform feature extraction on the local emotion features corresponding to each basic element to obtain context correlation features of the local emotion features.

[0079] In specific implementation, the above-mentioned first feature extraction module includes a convolutional layer, a normalization layer, an activation function layer, and a pooling layer connected in sequence; among them, the convolutional layer is used to extract features from the speech to be processed to obtain feature data; the normalization layer is used to perform normalization processing on the feature data to obtain normalized feature data; the activation function layer is used to perform a function transformation on the normalized feature data; the pooling layer is used to perform a pooling operation on the feature data after the function transformation to obtain local emotion features.

[0080] Furthermore, the above-mentioned normalization layer uses the BN method to perform normalization processing on the feature data; the activation function layer uses the ELU function to perform a function transformation on the normalized feature data.

[0081] Furthermore, the above-mentioned emotion recognition module 41 is used to: classify the local emotion features and the context correlation features of the local emotion features under a preset plurality of emotion categories to obtain the emotion category to which the speech to be processed belongs; determine the emotion category to which the speech to be processed belongs as the recognition result.

[0082] Specifically, the above-mentioned emotion recognition module 41 is further used to: integrate the local emotion features and the context correlation features of the local emotion features to obtain an integration result; classify the integration result under a plurality of emotion categories through a preset classifier to obtain the emotion category to which the speech to be processed belongs.

[0083] In specific implementation, the above-mentioned emotion recognition module 41 is further used to: calculate the probability value of the integration result under each emotion category through the Softmax function; determine the emotion category corresponding to the maximum value of the probability value as the emotion category to which the speech to be processed belongs.

[0084] For the speech emotion recognition device provided by the embodiments of the present invention, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments.

[0085] Embodiments of the present invention also provide an electronic device. Refer to Figure 5 As shown, the electronic device includes a processor 101 and a memory 100. The memory 100 stores machine-executable instructions that can be executed by the processor 101, and the processor 101 executes the machine-executable instructions to implement the above-mentioned speech emotion recognition method.

[0086] Furthermore, Figure 5 As shown, the electronic device further includes a bus 102 and a communication interface 103. The processor 101, the communication interface 103, and the memory 100 are connected through the bus 102.

[0087] Among them, the memory 100 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0088] The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 101 or the instructions in the form of software. The above-mentioned processor 101 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 100, and the processor 101 reads the information in the memory 100 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0089] The embodiments of the present invention also provide a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above-mentioned speech emotion recognition method. For the specific implementation, reference can be made to the method embodiments, which will not be elaborated here.

[0090] A computer program product of a voice emotion recognition method, apparatus, and electronic device provided by an embodiment of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated herein. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0091] Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program code.

[0092] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for speech emotion recognition, characterized in that, The method includes: Obtaining the speech to be processed; Inputting the speech to be processed into an emotion recognition model, and outputting a recognition result for indicating the speech emotion of the speech to be processed; the emotion recognition model includes a plurality of first feature extraction modules connected in parallel and a second feature extraction module; Wherein, the emotion recognition model is used for: for each of the first feature extraction modules, extracting the local emotion features of the basic elements corresponding to the first feature extraction module in the speech to be processed; wherein, the speech to be processed includes a plurality of basic elements, and the plurality of basic elements include: words, phrases, tones, and volumes in the speech to be processed; through the second feature extraction module, performing feature extraction on the local emotion features corresponding to each of the basic elements to obtain the context correlation features of the local emotion features; based on the local emotion features and the context correlation features of the local emotion features, identifying the speech emotion of the speech to be processed to obtain the recognition result.

2. The method according to claim 1, wherein The first feature extraction module includes a convolutional layer, a normalization layer, an activation function layer, and a pooling layer connected in sequence; wherein, the convolutional layer is used for performing feature extraction on the speech to be processed to obtain feature data; the normalization layer is used for performing normalization processing on the feature data to obtain the normalized feature data; the activation function layer is used for performing function transformation on the normalized feature data; the pooling layer is used for performing a pooling operation on the feature data after function transformation to obtain local emotion features.

3. The method according to claim 2, wherein The normalization layer uses the BN method to perform normalization processing on the feature data; the activation function layer uses the ELU function to perform function transformation on the normalized feature data.

4. The method according to claim 1, characterized in that The step of identifying the speech emotion of the speech to be processed based on the local emotion features and the context correlation features of the local emotion features to obtain the recognition result includes: Under a preset plurality of emotion categories, classifying the local emotion features and the context correlation features of the local emotion features to obtain the emotion category to which the speech to be processed belongs; determining the emotion category to which the speech to be processed belongs as the recognition result.

5. The method according to claim 4, wherein The step of classifying the local emotion features and the context correlation features of the local emotion features under a preset plurality of emotion categories to obtain the emotion category to which the speech to be processed belongs includes: Integrating the local emotion features and the context correlation features of the local emotion features to obtain an integration result; Classifying the integration result under the plurality of emotion categories through a preset classifier to obtain the emotion category to which the speech to be processed belongs.

6. The method according to claim 5, characterized in that, The step of classifying the integration result under the plurality of emotion categories through the preset classifier to obtain the emotion category to which the speech to be processed belongs includes: Calculating the probability value of the integration result under each of the emotion categories through the Softmax function; Determining the emotion category corresponding to the maximum value of the probability value as the emotion category to which the speech to be processed belongs.

7. A voice emotion recognition device, characterized in that, The device includes: A speech acquisition module for acquiring the speech to be processed; An emotion recognition module, configured to input the to-be-processed speech into an emotion recognition model, and output a recognition result indicating the speech emotion of the to-be-processed speech; the emotion recognition model includes a plurality of first feature extraction modules connected in parallel and a second feature extraction module; Wherein, the emotion recognition model is configured to: for each of the first feature extraction modules, extract local emotion features of basic elements corresponding to the first feature extraction module in the to-be-processed speech; wherein, the to-be-processed speech includes a plurality of basic elements, and the plurality of basic elements include: words, phrases, tones and volumes in the to-be-processed speech; through the second feature extraction module, perform feature extraction on the local emotion features corresponding to each basic element to obtain context correlation features of the local emotion features; based on the local emotion features and the context correlation features of the local emotion features, recognize the speech emotion of the to-be-processed speech to obtain the recognition result.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the speech emotion recognition method according to any one of claims 1 to 6.

9. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the speech emotion recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Model generation method and system, emotion recognition method and system, equipment and storage medium

    CN110909131A

  • Voice emotional state evaluation method and device based on attention, medium and equipment

    CN111402928A