Speech Emotion Recognition Method, Terminal Device and Computer Readable Storage Medium
By mapping text features into image space and fusing them with audio features in image space, the problem of missing semantic information of speech emotion recognition in the prior art is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111615879.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-27
AI Technical Summary
When extracting language and paralinguistic features, existing speech emotion recognition methods lose important semantic information, affecting the accuracy of the recognition results.
Map text features into image space, acquire image features, and fuse information with audio features in image space to form fusion features to identify speech emotion categories.
By retaining language and semantic information, the accuracy of speech emotion recognition is improved and more feature information is obtained to support more accurate emotion recognition.
Smart Images

Figure CN114203159B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of speech processing, and particularly relates to a speech emotion recognition method, device, terminal device, and computer-readable storage medium. Background Art
[0002] Speech emotion recognition refers to the technology of recognizing the emotional state of a speaker based on the input speech information. Speech is the most direct way of communication between people, containing information such as the speaker's semantics, intonation, and tone, and can well reflect emotional characteristics. Therefore, speech emotion recognition technology is a very popular research direction in the field of emotion recognition.
[0003] In the prior art, usually, the linguistic features (such as text features) and paralinguistic features (such as pitch, rhythm, etc.) in the speech to be processed are extracted separately, and then the emotional type is recognized based on these two types of features. Such a feature extraction method often loses some important semantic information, thereby affecting the subsequent recognition results. Summary of the Invention
[0004] The embodiments of this application provide a speech emotion recognition method, device, terminal device, and computer-readable storage medium, which can effectively improve the accuracy of speech emotion recognition.
[0005] In a first aspect, the embodiments of this application provide a speech emotion recognition method, including:
[0006] Obtaining the text features obtained by performing speech recognition on the speech to be processed, and the audio features obtained by performing audio feature extraction on the speech to be processed;
[0007] Mapping the text features to the image space to obtain image features;
[0008] According to the correspondence between the audio features and the text features, fusing the audio features and the image features to obtain fused features;
[0009] Recognizing the emotional category of the speech to be processed according to the fused features.
[0010] Since the text features include language information and the image features include richer semantic information, in the embodiments of the present application, the text features of the speech to be processed are mapped to the image space, and the obtained image features not only retain the original language information of the speech to be processed but also add the semantic information of the information to be processed. Since the audio features include paralinguistic features such as pitch and rhythm, the audio features of the text to be processed are fused with the image features, and the obtained fused features include the language information, paralinguistic information, and semantic information of the speech to be processed. Through the above method, more feature information in the speech to be processed can be obtained, and using this more feature information to identify the emotion category of the speech to be processed can effectively improve the accuracy of speech emotion recognition.
[0011] In a possible implementation manner of the first aspect, before mapping the text features to the image space to obtain image features, the method further includes:
[0012] Obtain a training text and a training image that semantically matches the training text;
[0013] Input the features of the training text into a preset generator to obtain the features of the generated image;
[0014] Input the features of the generated image and the features of the training image into a preset discriminator to obtain a discrimination result;
[0015] Update the parameters of the generator according to the discrimination result to obtain the trained generator;
[0016] Correspondingly, mapping the text features to the image space to obtain image features includes:
[0017] Input the text features into the trained generator to obtain the image features.
[0018] In a possible implementation manner of the first aspect, fusing the audio features and the image features according to the corresponding relationship between the audio features and the text features to obtain fused features includes:
[0019] Calculate a first mapping relationship between a first local feature in the image features and a second local feature in the text features, where the first local region is used to represent a region in the image, and the second local feature is used to represent a word in the text;
[0020] Calculate a second mapping relationship between a third local feature in the audio features and a fourth local feature in the text features, where the third local feature is used to represent a phoneme in the audio, and the fourth local feature is used to represent a word in the text;
[0021] According to the first mapping relationship and the second mapping relationship, perform information fusion on the audio feature and the image feature to obtain a fusion feature.
[0022] In a possible implementation manner of the first aspect, the performing information fusion on the audio feature and the image feature according to the first mapping relationship and the second mapping relationship to obtain a fusion feature includes:
[0023] For each group of third local features, obtain a target feature according to the first mapping relationship and the second mapping relationship, where the target feature is a first local feature in the image feature corresponding to the third local feature;
[0024] Add the third local feature to the target feature to obtain the fused target feature;
[0025] After processing all the third local features, generate the fusion feature from the fused target feature and the unfused first local feature.
[0026] In a possible implementation manner of the first aspect, the identifying the emotion category of the speech to be processed according to the fusion feature includes:
[0027] Perform feature extraction processing on the fusion feature to obtain a target feature;
[0028] Input the target feature into a preset classifier to output the emotion category.
[0029] In a possible implementation manner of the first aspect, the steps of obtaining the audio feature of the speech to be processed include:
[0030] Perform frame splitting on the speech to be processed to obtain a plurality of audio segments;
[0031] Obtain the spectrum corresponding to each of the plurality of audio segments;
[0032] Perform filtering processing on the spectrum according to a preset filter to obtain the audio feature.
[0033] In a possible implementation manner of the first aspect, the steps of obtaining the text feature of the speech to be processed include:
[0034] Input the speech to be processed into a preset speech recognition model to output the text feature;
[0035] Correspondingly, the second mapping relationship between the third local feature in the audio feature and the fourth local feature in the text feature is determined according to the intermediate calculation result of the speech recognition model.
[0036] Second aspect, an embodiment of the present application provides a voice emotion recognition device, including:
[0037] An acquisition unit, configured to acquire text features obtained by performing speech recognition on the speech to be processed, and audio features obtained by performing audio feature extraction on the speech to be processed;
[0038] A mapping unit, configured to map the text features to an image space to obtain image features;
[0039] A fusion unit, configured to perform information fusion on the audio features and the image features according to the corresponding relationship between the audio features and the text features to obtain fusion features;
[0040] An identification unit, configured to identify the emotion category of the speech to be processed according to the fusion features.
[0041] Third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the voice emotion recognition method according to any one of the above first aspects is implemented.
[0042] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the voice emotion recognition method according to any one of the above first aspects is implemented.
[0043] Fifth aspect, an embodiment of the present application provides a computer program product, when the computer program product runs on a terminal device, the terminal device is enabled to execute the voice emotion recognition method according to any one of the above first aspects.
[0044] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 It is a flowchart of the existing method provided by the embodiment of the present application;
[0047] Figure 2 It is a schematic flowchart of a voice emotion recognition method provided by an embodiment of the present application;
[0048] Figure 3 It is a flowchart of the recognition method provided by an embodiment of the present application;
[0049] Figure 4 It is a structural block diagram of a voice emotion recognition device provided by an embodiment of the present application;
[0050] Figure 5 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0051] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0052] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0053] It should also be understood that the term "and / or" as used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0054] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.
[0055] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0056] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0057] Voice emotion recognition refers to the technology of recognizing the emotional state of a speaker based on the input voice information. When existing voice emotion recognition methods perform language and paralinguistic feature modeling, they adopt methods based on sequence models (such as neural network models). For example, a recurrent neural network, as follows:
[0058] h0 = 0;
[0059] h t = σ(Wx t + Uh t-1 );
[0060]
[0061] where x [1:N] is the input sequence, W and U are matrices respectively, and σ is a non-linear activation function (such as sigmoid, etc.). is the finally obtained feature, w k is the weight corresponding to the sequence element h k . After obtaining the paralinguistic and linguistic features and respectively, they are concatenated and classified. For example:
[0062]
[0063] where concat represents concatenation, f represents the classifier, and y represents the result of emotion recognition.
[0064] The flowchart of the existing voice emotion recognition method is shown in Figure 1 . As Figure 1 shown, in the prior art, the voice to be processed is obtained through voice recognition technology (such as ASR shown in Figure 1 ), such as Figure 1The text information of the voice shown is obtained, and then word features are obtained from the text information, and then a text sequence representation is obtained. Then, the frame features of the voice to be processed (referring to the segments obtained by segmenting the voice at intervals) are obtained, and then a voice sequence representation is obtained. Then, the voice sequence representation and the text sequence representation are concatenated together, which is equivalent to concatenating the audio features and the text features. Finally, emotion recognition is performed using the concatenated features.
[0065] As Figure 1 shown, in the prior art, the linguistic features (text features) and paralinguistic features (frame features) in the voice to be processed are extracted separately. Such a feature extraction method often loses some important semantic information in the voice to be processed. In addition, directly mapping the linguistic features and paralinguistic features to the emotion recognition result cannot ensure that the emotion recognition model can learn the required feature information.
[0066] To solve the above problems, an embodiment of the present application provides a voice emotion recognition method. Refer to Figure 2 , which is a schematic flowchart of the voice emotion recognition method provided by the embodiment of the present application. By way of example and not limitation, the method may include the following steps:
[0067] S201, obtain the text features obtained by performing speech recognition on the voice to be processed, and the audio features obtained by performing audio feature extraction on the voice to be processed.
[0068] The voice to be processed in the embodiment of the present application may be digital audio, such as a sampling point sequence with a sampling rate greater than 16KHz.
[0069] In one embodiment, the method for obtaining text features is:
[0070] Input the voice to be processed into a preset speech recognition model, and output the text features.
[0071] The purpose of speech recognition is to recognize speech as text. In the embodiment of the present application, existing technologies, such as speech recognition tools like Kaldi, can be used to perform speech recognition processing on the voice to be processed.
[0072] Optionally, the result of the speech recognition processing can be further converted into a representation vector. For example, existing neural network models (such as RNN, Transformer, etc.) can be used to convert the variable-length input into a fixed-length representation vector.
[0073] Correspondingly, the speech recognition model in the embodiment of the present application includes both a speech recognition tool for recognizing speech as text and a model for converting the speech recognition result into a representation vector.
[0074] One way to obtain audio features is as follows: Input the speech to be processed into a trained audio feature extraction model to obtain the audio features. Among them, the audio feature extraction model can be a neural network.
[0075] Another way to obtain audio features is:
[0076] Perform frame splitting on the speech to be processed to obtain multiple audio segments; obtain the spectra corresponding to the multiple audio segments respectively; perform filtering processing on the spectra according to a preset filter to obtain the audio features.
[0077] In the embodiments of the present application, the Fourier transform can be used to obtain the spectra of the audio segments. The preset filter can be a Mel filter. Specifically, the process of performing filtering processing according to the Mel filter may include: inputting the spectra into the Mel filter to obtain the Mel spectra; taking the logarithm of the Mel spectra and performing inverse transformation processing to obtain the cepstral coefficients on the Mel spectra. These cepstral coefficients are called Mel-frequency cepstral coefficients, and the Mel-frequency cepstral coefficients are used as the audio features.
[0078] The Mel-frequency cepstral coefficients can accurately describe the characteristics of the vocal tract shape in the envelope of the short-time power spectrum of speech, which is beneficial to subsequent speech emotion recognition.
[0079] It should be noted that in the embodiments of the present application, the mapping from text features to image features and the mapping from audio features to image features are taken as examples for description, that is, text features and audio features are used as input data, and image features are used as output data. In practical applications, text and audio can also be used as input data, and images can be used as output data.
[0080] S202, map the text features to the image space to obtain image features.
[0081] In the embodiments of the present application, a generator model can be used to obtain the mapping relationship between text features and image features.
[0082] In one embodiment, the way to obtain image features is: input the text features into the trained generator to obtain the image features.
[0083] Correspondingly, the generator needs to be pre-trained. The specific steps include:
[0084] Obtain training texts and training images that match the semantics expressed by the training texts; input the features of the training texts into a preset generator to obtain the features of the generated images; input the features of the generated images and the features of the training images into a preset discriminator to obtain a discrimination result; update the parameters of the generator according to the discrimination result to obtain the trained generator.
[0085] During the process of training the generator, training texts and corresponding training images can be collected through the technique of generating descriptions from images. Then, the generator is trained through a generative adversarial approach as follows:
[0086]
[0087]
[0088] Among them is the text representation, V T represents the image that meets the conditions, and G represents the generator (such as CNN, RNN, etc.). represents the image of the real corresponding text, D represents the discriminator, which is used to distinguish between real images and generated images The learning objective is for D to accurately distinguish between real and fake images as much as possible, while G makes it as difficult as possible for D to distinguish. Exemplarily, the training loss function, that is, the objective function, can be as follows:
[0089]
[0090] Among them, obj refers to the objective to be solved by the model. N is the number of data pairs of the pre-collected text (T i ) and the image (physical space representation (I i ). θ represents the parameters of the probability model, and p θ (I i |T i ) represents the probability function of the known text T i and its corresponding physical space representation I i . According to the conditional probability p θ (I i |T i ), the physical space representation corresponding to the text can be obtained. The θ obtained from the above formula is the parameter of the generator.
[0091] S203. According to the corresponding relationship between the audio feature and the text feature, the audio feature and the image feature are fused to obtain a fused feature.
[0092] In the embodiments of the present application, the information fusion method adopted is to map the audio feature and the text feature to the image space (i.e., the physical space) respectively and fuse the two in the image space.
[0093] In one embodiment, the information fusion method includes:
[0094] I. Calculate a first mapping relationship between a first local feature in the image feature and a second local feature in the text feature, where the first local region is used to represent a region in the image, and the second local feature is used to represent a word in the text.
[0095] In this step, a generative model can also be used. However, here the input to the generative model is the image feature, and the output is the text feature. The training of the generative model can also adopt the generative adversarial method, as described in S202, which will not be elaborated here.
[0096] The generator includes an encoder and a decoder, and only a feature vector representing the global information of the original text is used to connect them, which will seriously affect the decoding effect. If the length of the original text is too long, it is difficult for the Decoder to perform accurate decoding word by word based on the global information of the original text.
[0097] To solve the above problems, optionally, the attention mechanism can be introduced and the attention model is combined with the generative model. For example, the attention mechanism is introduced in the decoder part. The encoder part still uses a series of LSTM (Long Short-Term Memory) units to encode the original text. The input of each LSTM unit in the decoder part is the concatenation of the output of the attention unit and the output vector of the previous LSTM unit. In this way, the decoder can focus on different parts of the original text for decoding at each step of decoding. The output of each attention unit is the superposition of the weighted average of the output vectors of each LSTM unit in the encoder and the hidden layer vector of the previous LSTM unit in the decoder. The output after combining the attention model and the generative model is the attention distribution (weight distribution) of each word in the text corresponding to each region in the image, and this attention distribution can be used as the first mapping relationship.
[0098] II. Calculate a second mapping relationship between a third local feature in the audio feature and a fourth local feature in the text feature, where the third local feature is used to represent a phoneme in the audio, and the fourth local feature is used to represent a word in the text.
[0099] As described in S201, the text feature is obtained through a speech recognition model. Since in the recognition process of the speech recognition model, it is necessary to know the phoneme corresponding to each frame of audio and then form words from the phonemes. Therefore, correspondingly, the second mapping relationship between the third local feature in the audio feature and the fourth local feature in the text feature can be determined according to the intermediate calculation results of the speech recognition model.
[0100] III. According to the first mapping relationship and the second mapping relationship, fuse the audio feature and the image feature to obtain a fused feature.
[0101] In step I, it is equivalent to obtaining the regions in the images corresponding to each word. In step II, it is equivalent to obtaining the audio features corresponding to each word. The audio features can be mapped to the image space through the words. Specifically, one implementation of step III is as follows:
[0102] For each group of third local features, obtain the target features according to the first mapping relationship and the second mapping relationship. The target features are the first local features in the image features corresponding to the third local features; add the third local features to the target features to obtain the fused target features;
[0103] After processing all the third local features, generate the fused features from the fused target features and the unfused first local features.
[0104] In one implementation, the above information fusion method can be implemented through a mapping model. As described in S202, a mapping model from audio features to the image space can be trained. Exemplarily, the objective function for training can be:
[0105] obj align = argmax θ ∏p θ (warp|I,T);
[0106] The above formula describes the alignment model between the paralinguistic features (aligned with the text during the speech recognition process, denoted by T for the paralinguistic features) and the physical space (image space). Here, I refers to the physical space representation corresponding to T, and warp refers to the mapping relationship between the manually labeled text words and the local physical space representations. The p θ (warp|I,T) trained with data is used for the alignment between the paralinguistic features and the physical space representations.
[0107] The image space usually includes multiple channels, and different channels contain different feature information. Optionally, the audio features can be spliced as additional channels onto the image regions.
[0108] The fused features obtained by the above method include both the language information of the text, the semantic information of the image space, and the paralinguistic information of the audio. Moreover, these feature information are correlated with each other in the image space, which is beneficial to analyzing the inducement of the emotional state during the subsequent emotion recognition process, thereby improving the accuracy of emotion recognition.
[0109] S204, recognize the emotion category of the to-be-processed speech according to the fused features.
[0110] In one embodiment, the fused features can be input into a preset classifier to obtain the emotion category of the speech to be processed.
[0111] To further extract more accurate feature information and further improve the recognition accuracy. In another embodiment, S204 includes:
[0112] Perform feature extraction processing on the fused features to obtain target features; input the target features into a preset classifier and output the emotion category.
[0113] Models such as convolutional neural networks or local statistical features can be used to perform feature extraction on the fused features.
[0114] The preset classifier uses a classifier that has been trained to reach the preset accuracy.
[0115] Since the text features include language information, and the image features include richer semantic information, in the embodiments of the present application, the text features of the speech to be processed are mapped to the image space, and the obtained image features not only retain the original language information of the speech to be processed, but also add the semantic information of the information to be processed. Since the audio features include paralinguistic features such as pitch and prosody, the audio features of the text to be processed are fused with the image features, and the obtained fused features contain the language information, paralinguistic information and semantic information of the speech to be processed. Through the above method, more feature information in the speech to be processed can be obtained, and the emotion category of the speech to be processed can be recognized by using the more feature information, which can effectively improve the accuracy of speech emotion recognition.
[0116] See Figure 3 , which is a flowchart of the recognition method provided by the embodiments of the present application. As Figure 3 shown, in the embodiments of the present application, first, the speech to be processed (such as the speech shown in Figure 3 ) is subjected to speech recognition processing through a speech recognition model (ASR, such as Kaldi, etc.) to obtain text; then, the recognized text is converted into word features by using a BERT (Bidirectional Encoder Representation from Transformers, language representation) model; the word features are mapped to the image space to align the text features with the image features. Then, audio feature extraction is performed on the speech to be processed to obtain frame features; the frame features are mapped to the image space to align the audio features (acoustics) with the image features. Finally, according to the alignment results of the text features and the image features, and the alignment results of the audio features and the image features, the emotion type of the speech to be processed is recognized.
[0117] Compared with Figure 1Compared with the prior art shown, the feature extraction method in the speech emotion recognition method provided in the embodiments of the present application can obtain more feature information and can obtain the correlation relationship between various feature information, thereby providing a more reliable data basis for subsequent emotion recognition.
[0118] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0119] Corresponding to the speech emotion recognition method described in the above embodiments, Figure 4 is a structural block diagram of a speech emotion recognition device provided in an embodiment of the present application. For the convenience of description, only the parts related to the embodiments of the present application are shown.
[0120] Referring to Figure 4 , the device includes:
[0121] An acquisition unit 41, configured to acquire text features obtained by performing speech recognition on the speech to be processed, and audio features obtained by performing audio feature extraction on the speech to be processed.
[0122] A mapping unit 42, configured to map the text features to an image space to obtain image features.
[0123] A fusion unit 43, configured to perform information fusion on the audio features and the image features according to the corresponding relationship between the audio features and the text features to obtain fusion features.
[0124] An identification unit 44, configured to identify the emotion category of the speech to be processed according to the fusion features.
[0125] Optionally, the device 4 further includes:
[0126] A training unit 45, configured to acquire training texts and training images that match the semantics expressed by the training texts before mapping the text features to an image space to obtain image features; input the features of the training texts into a preset generator to obtain the features of the generated images; input the features of the generated images and the features of the training images into a preset discriminator to obtain a discrimination result; update the parameters of the generator according to the discrimination result to obtain the trained generator.
[0127] Correspondingly, the mapping unit 42 is further configured to:
[0128] Input the text features into the trained generator to obtain the image features.
[0129] Optionally, the fusion unit 43 is further configured to:
[0130] Calculate a first mapping relationship between a first local feature in the image feature and a second local feature in the text feature, where the first local region is used to represent a region in the image, and the second local feature is used to represent a word in the text;
[0131] Calculate a second mapping relationship between a third local feature in the audio feature and a fourth local feature in the text feature, where the third local feature is used to represent a phoneme in the audio, and the fourth local feature is used to represent a word in the text;
[0132] According to the first mapping relationship and the second mapping relationship, perform information fusion on the audio feature and the image feature to obtain a fusion feature.
[0133] Optionally, the fusion unit 43 is further configured to:
[0134] For each group of third local features, obtain a target feature according to the first mapping relationship and the second mapping relationship, where the target feature is the first local feature in the image feature corresponding to the third local feature;
[0135] Add the third local feature to the target feature to obtain the fused target feature;
[0136] After processing all the third local features, generate the fusion feature from the fused target feature and the unfused first local feature.
[0137] Optionally, the recognition unit 44 is further configured to:
[0138] Perform feature extraction processing on the fusion feature to obtain a target feature;
[0139] Input the target feature into a preset classifier to output the sentiment category.
[0140] Optionally, the acquisition unit 41 is further configured to:
[0141] Perform frame splitting on the to-be-processed speech to obtain a plurality of audio segments;
[0142] Obtain the respective spectra corresponding to the plurality of audio segments;
[0143] Perform filtering processing on the spectrum according to a preset filter to obtain the audio feature.
[0144] Optionally, the acquisition unit 41 is further configured to:
[0145] Input the to-be-processed speech into a preset speech recognition model to output the text feature.
[0146] Correspondingly, the second mapping relationship between the third local feature in the audio feature and the fourth local feature in the text feature is determined according to the intermediate calculation result of the speech recognition model.
[0147] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0148] In addition, Figure 4 The shown speech emotion recognition device can be a software unit, a hardware unit, or a unit combining software and hardware built into an existing terminal device, can also be integrated into the terminal device as an independent pendant, or can exist as an independent terminal device.
[0149] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, can also exist physically as individual units, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details are not described herein again.
[0150] Figure 5 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. As Figure 5 shown, the terminal device 5 in this embodiment includes: at least one processor 50 ( Figure 5 only one is shown in the figure), a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in any of the above-mentioned speech emotion recognition method embodiments are implemented.
[0151] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand, Figure 5The terminal device 5 is merely an example and does not limit the terminal device 5. It may include more or fewer components than those shown, or combine certain components, or have different components. For example, it may also include input / output devices, network access devices, etc.
[0152] The so-called processor 50 may be a central processing unit (CPU). The processor 50 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0153] In some embodiments, the memory 51 may be an internal storage unit of the terminal device 5, such as the hard disk or memory of the terminal device 5. In other embodiments, the memory 51 may also be an external storage device of the terminal device 5, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the terminal device 5. Further, the memory 51 may also include both the internal storage unit and the external storage device of the terminal device 5. The memory 51 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or will be output.
[0154] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0155] An embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device can implement the steps in the above method embodiments when executed.
[0156] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0157] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0158] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0159] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0160] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for speech emotion recognition, characterized in that, Including: Obtaining text features obtained by performing speech recognition on the speech to be processed, and audio features obtained by performing audio feature extraction on the speech to be processed; Mapping the text features to an image space to obtain image features; Calculating a first mapping relationship between a first local feature in the image features and a second local feature in the text features, where the first local feature is used to characterize a region in an image, and the second local feature is used to characterize a word in the text; Calculating a second mapping relationship between a third local feature in the audio features and a fourth local feature in the text features, where the third local feature is used to characterize a phoneme in the audio, and the fourth local feature is used to characterize a word in the text; According to the first mapping relationship and the second mapping relationship, performing information fusion on the audio features and the image features to obtain fusion features; Identifying an emotional category of the speech to be processed according to the fusion features.
2. The voice emotion recognition method according to claim 1, wherein Before mapping the text features to an image space to obtain image features, the method further includes: Obtaining a training text and a training image that semantically matches the expression of the training text; Inputting the features of the training text into a preset generator to obtain the features of a generated image; Inputting the features of the generated image and the features of the training image into a preset discriminator to obtain a discrimination result; Updating the parameters of the generator according to the discrimination result to obtain the trained generator; Correspondingly, the mapping the text features to an image space to obtain image features includes: Inputting the text features into the trained generator to obtain the image features.
3. The voice emotion recognition method according to claim 1, characterized in that, The performing information fusion on the audio features and the image features according to the first mapping relationship and the second mapping relationship to obtain fusion features includes: For each group of third local features, obtaining target features according to the first mapping relationship and the second mapping relationship, where the target features are the first local features in the image features corresponding to the third local features; Adding the third local features to the target features to obtain the fused target features; After processing all the third local features, generating the fusion features from the fused target features and the unfused first local features.
4. The speech emotion recognition method according to claim 1, wherein The identifying an emotional category of the speech to be processed according to the fusion features includes: Performing feature extraction processing on the fusion features to obtain target features; Inputting the target features into a preset classifier to output the emotional category.
5. The speech emotion recognition method according to claim 1, wherein The step of obtaining the audio features of the speech to be processed includes: Performing frame splitting on the speech to be processed to obtain a plurality of audio segments; Obtaining the respective spectra corresponding to the plurality of audio segments; Performing filtering processing on the spectra according to a preset filter to obtain the audio features.
6. The speech emotion recognition method according to claim 1, characterized in that The step of obtaining the text features of the speech to be processed includes: Inputting the speech to be processed into a preset speech recognition model to output the text features; Correspondingly, the second mapping relationship between the third local feature in the audio feature and the fourth local feature in the text feature is the intermediate calculation result of the speech recognition model.
7. A voice emotion recognition device, characterized in that, Including: An acquisition unit configured to acquire the text feature obtained by performing speech recognition on the speech to be processed, and the audio feature obtained by performing audio feature extraction on the speech to be processed; A mapping unit configured to map the text feature into an image space to obtain an image feature; A fusion unit configured to calculate a first mapping relationship between a first local feature in the image feature and a second local feature in the text feature, where the first local feature is used to characterize a region in the image, and the second local feature is used to characterize a word in the text; calculate a second mapping relationship between a third local feature in the audio feature and a fourth local feature in the text feature, where the third local feature is used to characterize a phoneme in the audio, and the fourth local feature is used to characterize a word in the text; and perform information fusion on the audio feature and the image feature according to the first mapping relationship and the second mapping relationship to obtain a fusion feature; An identification unit configured to identify the emotion category of the speech to be processed according to the fusion feature.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Multi-modal emotion recognition method and device
CN111564164A
Voice-based image generation method and device, equipment and storage medium
CN113256751A