Emotion recognition method, apparatus, device, and storage medium

By processing users' voice data to obtain spectrograms and text, and extracting voice, image, and text features, the problem of low accuracy in emotion recognition in existing technologies is solved, achieving more efficient and accurate emotion recognition.

CN116013369BActive Publication Date: 2026-05-29GREAT WALL MOTOR CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GREAT WALL MOTOR CO LTD
Filing Date
2022-11-22
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in emotion recognition in complex or special scenarios, poor resistance to interference and robustness, resulting in inaccurate emotion recognition for users.

Method used

By acquiring users' voice data, processing it to obtain spectrograms and text, extracting features from the voice, images, and text, and combining these features to identify users' emotions.

Benefits of technology

It improves the accuracy of emotion recognition, reduces the complexity and cost of data acquisition, and enhances the robustness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116013369B_ABST
    Figure CN116013369B_ABST
Patent Text Reader

Abstract

The application discloses an emotion recognition method and device, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: obtaining voice data of a target object; processing the voice data to obtain a spectrogram and text corresponding to the voice data; extracting features of the voice data to obtain voice features of the voice data; extracting features of the spectrogram to obtain image features of the spectrogram; extracting features of the text to obtain text features of the text; and determining the emotion of the target object based on the voice features, the image features and the text features. The application obtains features of three modalities, i.e., an image, voice and text, by processing voice data of a target object, and then comprehensively recognizes the emotion of the target object based on the features of the three modalities, so that the accuracy of emotion recognition of the target object can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an emotion recognition method, apparatus, device, and storage medium. Background Technology

[0002] Today, artificial intelligence (AI) technology is developing rapidly, and AI products are emerging in large numbers. In some cases, AI products can replace human operators in performing certain tasks, such as interacting with humans (e.g., conversing). To enhance the user experience of AI products, features such as emotion recognition (joy, anger, sorrow, happiness, etc.) can be added, enabling AI products to respond accordingly based on human emotions.

[0003] In related technologies, user voice data is acquired, the user voice data is converted into text to obtain corresponding text data, and then the text data can be analyzed to realize the recognition of user emotions.

[0004] However, due to the difficulty of emotion recognition in some complex or special scenarios, the above-mentioned emotion recognition methods are difficult to accurately identify users' emotions in such cases, resulting in poor anti-interference and robustness, which will reduce the accuracy of user emotion recognition. Summary of the Invention

[0005] This application provides an emotion recognition method, apparatus, device, and storage medium, which can improve the accuracy of emotion recognition by acquiring only the user's voice data, thereby enhancing the user experience. The technical solution is as follows:

[0006] Firstly, an emotion recognition method is provided, the method comprising:

[0007] Acquire the voice data of the target object;

[0008] The speech data is processed to obtain the spectrogram and text corresponding to the speech data;

[0009] The speech data is subjected to feature extraction to obtain speech features; the spectrogram is subjected to feature extraction to obtain image features; and the text is subjected to feature extraction to obtain text features.

[0010] The emotion of the target object is determined based on the speech features, the image features, and the text features.

[0011] In this application, speech data of the target object is acquired, processed to obtain the corresponding spectrogram and text, and then feature extraction is performed on the speech data, spectrogram, and text to obtain speech features, image features, and text features. Thus, features in three modalities—image, speech, and text—can be obtained simply by processing the speech data of the target object. These three modal features can more comprehensively represent the emotional characteristics of the target object. Subsequently, based on the speech features, image features, and text features, the emotion of the target object is determined. By comprehensively determining the emotion of the target object using these three modal features, the accuracy of emotion recognition for the target object can be improved.

[0012] Optionally, processing the speech data to obtain the spectrogram and text corresponding to the speech data includes:

[0013] The voice data is divided into multiple voice segments;

[0014] Perform time-frequency transformation on the multiple speech segments to obtain the spectrogram corresponding to the speech data;

[0015] Text recognition is performed on the multiple speech segments to obtain the text corresponding to the speech data.

[0016] Optionally, the step of performing time-frequency transformation on the plurality of speech segments to obtain the spectrogram corresponding to the speech data includes:

[0017] For any one of the plurality of speech segments, perform a Fourier transform or wavelet transform on the speech segment to obtain a target spectrum; based on the target spectrum, generate a spectrogram corresponding to the speech segment;

[0018] The spectrograms of the multiple speech segments are spliced ​​together to obtain the spectrogram corresponding to the speech data.

[0019] Optionally, the step of extracting features from the speech data to obtain the speech features of the speech data includes:

[0020] The voice data is divided into multiple voice segments;

[0021] For any one of the plurality of speech segments, the spectrum of the speech segment is filtered to obtain filtering information; based on the filtering information, the speech segment features of the speech segment are determined.

[0022] The speech features of the multiple speech segments are concatenated to obtain the speech features of the speech data.

[0023] Optionally, determining the emotion of the target object based on the speech features, the image features, and the text features includes:

[0024] The speech features, image features, and text features are concatenated to obtain the multimodal features of the speech data;

[0025] The emotion of the target object is determined based on the multimodal features of the speech data.

[0026] Optionally, determining the emotion of the target object based on the multimodal features of the speech data includes:

[0027] The multimodal features of the speech data are encoded based on an attention mechanism to obtain the attention features of the speech data;

[0028] Based on the attentional features of the voice data, the emotion of the target object is determined.

[0029] Optionally, determining the emotion of the target object based on the attention features of the voice data includes:

[0030] The attention features of the speech data are sequence encoded to obtain the sequence encoded features of the speech data; the sequence encoded features are normalized to obtain the probability that the target object corresponds to multiple candidate emotions; the candidate emotion with the highest probability among the multiple candidate emotions is determined as the emotion of the target object.

[0031] Secondly, an emotion recognition device is provided, the device comprising:

[0032] The acquisition module is used to acquire the voice data of the target object;

[0033] The processing module is used to process the speech data to obtain the spectrogram and text corresponding to the speech data;

[0034] The feature extraction module is used to extract features from the speech data to obtain the speech features of the speech data; to extract features from the spectrogram to obtain the image features of the spectrogram; and to extract features from the text to obtain the text features of the text.

[0035] The determination module is used to determine the emotion of the target object based on the speech features, the image features, and the text features.

[0036] Optionally, the processing module is used for:

[0037] The voice data is divided into multiple voice segments;

[0038] Perform time-frequency transformation on the multiple speech segments to obtain the spectrogram corresponding to the speech data;

[0039] Text recognition is performed on the multiple speech segments to obtain the text corresponding to the speech data.

[0040] Optionally, the processing module is used to:

[0041] For any one of the plurality of speech segments, perform a Fourier transform or wavelet transform on the speech segment to obtain a target spectrum; based on the target spectrum, generate a spectrogram corresponding to the speech segment;

[0042] The spectrograms of the multiple speech segments are spliced ​​together to obtain the spectrogram corresponding to the speech data.

[0043] Optionally, the feature extraction module is used for:

[0044] The voice data is divided into multiple voice segments;

[0045] For any one of the plurality of speech segments, the spectrum of the speech segment is filtered to obtain filtering information; based on the filtering information, the speech segment features of the speech segment are determined.

[0046] The speech features of the multiple speech segments are concatenated to obtain the speech features of the speech data.

[0047] Optionally, the determining module includes:

[0048] The feature concatenation unit is used to concatenate the speech features, the image features, and the text features to obtain the multimodal features of the speech data;

[0049] The determining unit is used to determine the emotion of the target object based on the multimodal features of the speech data.

[0050] Optionally, the determining unit is used for:

[0051] The multimodal features of the speech data are encoded based on an attention mechanism to obtain the attention features of the speech data;

[0052] Based on the attentional features of the voice data, the emotion of the target object is determined.

[0053] Optionally, the determining unit is used for:

[0054] The attention features of the speech data are sequence encoded to obtain the sequence encoded features of the speech data; the sequence encoded features are normalized to obtain the probability that the target object corresponds to multiple candidate emotions; the candidate emotion with the highest probability among the multiple candidate emotions is determined as the emotion of the target object.

[0055] Thirdly, a computer device is provided, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the above-described emotion recognition method.

[0056] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described emotion recognition method.

[0057] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the steps of the aforementioned emotion recognition method.

[0058] It is understood that the beneficial effects of the second, third, fourth, and fifth aspects mentioned above can be found in the relevant descriptions in the first aspect above, and will not be repeated here. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a flowchart of an emotion recognition method provided in an embodiment of this application;

[0061] Figure 2 This is a flowchart of a speech feature extraction method provided in an embodiment of this application;

[0062] Figure 3 This is a flowchart of another emotion recognition method provided in the embodiments of this application;

[0063] Figure 4 This is a schematic diagram of the structure of an emotion recognition device provided in an embodiment of this application;

[0064] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0066] It should be understood that "multiple" as mentioned in this application refers to two or more. In the description of this application, unless otherwise stated, " / " indicates "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, to facilitate a clear description of the technical solutions of this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and that "first," "second," etc., do not necessarily imply differences.

[0067] Before providing a detailed explanation of the embodiments of this application, the application scenarios of these embodiments will be described first.

[0068] In related technologies, a common method for emotion recognition is to acquire the user's voice data, extract features from the voice data to obtain corresponding voice features, and then use these voice features to recognize the user's emotions. However, this method relies on relatively single features, and the accuracy of emotion recognition is low in some complex or special scenarios.

[0069] Another common approach to emotion recognition involves acquiring the user's voice data, corresponding text data, and video data including the user's image. Feature extraction is then performed on each of these data types to obtain voice features, text features, and corresponding image features. These features are then used to identify the user's emotions. However, this method requires acquiring multiple types of data, increasing data acquisition costs. Furthermore, in some special scenarios (such as telephone customer service), certain types of data may be unavailable, making emotion recognition impossible.

[0070] Therefore, this application provides an emotion recognition method that can be applied in human-computer interaction scenarios, especially in scenarios where artificial intelligence products recognize users' emotions.

[0071] Specifically, the process involves acquiring the user's voice data, converting it into text data, and then processing the voice data to obtain the corresponding image data. Next, voice features, image features, and text features are extracted from the voice data, image features, and text features. Finally, these features are combined to achieve user emotion recognition. In this way, features in three modalities—image, voice, and text—can be obtained simply by processing the user's voice data. Since these three modalities can more comprehensively represent the user's emotional characteristics, using these features to identify the user's emotional characteristics can improve the accuracy of the emotion recognition results.

[0072] The emotion recognition method provided in the embodiments of this application will be explained in detail below.

[0073] Figure 1 This is a flowchart illustrating an emotion recognition method provided in an embodiment of this application. This method can be applied to computer devices, which can be artificial intelligence products with human-computer interaction capabilities. See also... Figure 1 The method includes the following steps.

[0074] Step 101: The computer device acquires the voice data of the target object.

[0075] The target object is the object that needs to be sentiment-recognized; the target object is the user.

[0076] As an example, when a target object interacts with a computer device, the target object can output a voice message. If the computer device receives this voice message, then the computer device has obtained the voice data of the target object.

[0077] Step 102: The computer device processes the speech data to obtain the spectrogram and text corresponding to the speech data.

[0078] A spectrogram, also called a time-domain spectrogram, is a two-dimensional or three-dimensional spectrum. It is a graph that represents the frequency spectrum of speech data as it changes over time, with time on the horizontal axis and frequency on the vertical axis. A spectrogram can reflect information related to the characteristics of speech data, combining the features of the speech data's spectrogram and time-domain waveform to show how the speech data's spectrum changes over time.

[0079] In this case, the computer device obtains the spectrogram of the speech data, that is, it obtains information related to the sentence features of the speech data, so that the computer device can analyze the emotions of the target audience based on this.

[0080] Furthermore, in this embodiment, by processing the voice data, three types of data can be obtained: voice data, image data, and text data. These three types of data can reflect the emotions of the target audience from different perspectives. This reduces the complexity of data acquisition while increasing data diversity, thereby lowering the cost of data acquisition.

[0081] Specifically, the operation of step 102 includes the following steps (1)-(3).

[0082] (1) Computer equipment divides speech data into multiple speech segments.

[0083] These multiple speech segments have a temporal relationship, meaning that the computer device has divided the speech data into multiple speech segments with a temporal relationship.

[0084] In this case, by processing multiple speech segments to obtain the spectrogram and text corresponding to the speech data, processing resources can be saved and processing efficiency can be improved.

[0085] (2) The computer device performs time-frequency transformation on the multiple speech segments to obtain the spectrogram corresponding to the speech data.

[0086] Time-frequency transformation refers to converting speech segments between the time and frequency domains; that is, converting a speech segment from the time domain to the frequency domain, or vice versa. Because the domains differ, the perspective from which to analyze the speech segment also differs. Thus, time-frequency transformation allows for the analysis of the spectral characteristics of a speech segment from different angles, resulting in a more accurate spectrogram of the speech data.

[0087] Optionally, the time-frequency transformation can be a Fourier Transform (FT) or a wavelet transform (WT), etc. In this case, step (2) can be performed as follows: for any one of the multiple speech segments, perform a Fourier Transform or a wavelet transform on the speech segment to obtain the target spectrum; based on the target spectrum, generate the spectrogram corresponding to the speech segment; and concatenate the spectrograms of the multiple speech segments to obtain the spectrogram corresponding to the speech data.

[0088] Optionally, before performing Fourier transform or wavelet transform on the speech segment to obtain the target spectrum, the computer device can also perform frame-by-frame processing on the speech segment to obtain multiple speech signals corresponding to the speech segment. Subsequently, Fourier transform or wavelet transform can be performed on the obtained speech signals.

[0089] Alternatively, the Fourier transform can be a Fast Fourier Transform (FFT), which greatly reduces the number of multiplications required for the computer to compute the Fourier transform, thus saving processing resources.

[0090] For example, assuming the speech segment is X, the process of obtaining the spectrogram corresponding to the speech segment can include the following steps a-e.

[0091] a. When a computer device divides a speech segment X into frames, it will obtain X(m, n) (where m is the number of frames and n is the frame length).

[0092] b. The computer device performs a Fast Fourier Transform on each speech signal X(m,n) to obtain the spectrum of each speech signal X(m,n).

[0093] c. The computer device generates a periodic graph Y(m,n) of the spectrogram of each speech signal X(m,n).

[0094] Specifically, computer equipment can generate a periodic graph Y(m,n) using the formula Y(m,n) = X(m,n) × T·X(m,n), where T is the period.

[0095] d. The computer equipment performs logarithmic operations on Y(m, n). That is, it calculates using the formula 10lgY(m, n). Then, it transforms m into scale P based on time and n into scale Q based on frequency.

[0096] e. Finally, the computer device generates a spectrogram of the speech segment based on the obtained scale P, scale Q, and 10lgY(m,n) calculated by logarithm.

[0097] In this way, the spectrogram of each speech segment in multiple speech segments can be obtained. Then, the spectrograms of multiple speech segments can be spliced ​​together to obtain the spectrogram corresponding to the speech data.

[0098] The operation of splicing spectrograms of multiple speech segments by a computer device is similar to the operation of splicing multiple spectrograms by a device in related technologies, and will not be described in detail in the embodiments of this application.

[0099] (3) The computer device performs text recognition on the multiple speech segments to obtain the text corresponding to the speech data.

[0100] Specifically, the computer device performs text recognition on the multiple speech segments to obtain the text corresponding to the multiple speech segments. Then, the text corresponding to the multiple speech segments is concatenated to obtain the text corresponding to the speech data.

[0101] The operation of the computer device to concatenate the text corresponding to the multiple speech segments is similar to the operation of a device to concatenate multiple text segments in the prior art, and will not be described in detail in the embodiments of this application.

[0102] Step 103: The computer device performs feature extraction on the speech data to obtain the speech features of the speech data; performs feature extraction on the spectrogram to obtain the image features of the spectrogram; and performs feature extraction on the text to obtain the text features of the text.

[0103] Specifically, the operation of step 103 may include the following steps (1)-(3).

[0104] (1) The computer equipment processes the spectrum of the speech data to obtain the speech features of the speech data.

[0105] Since the spectrum of speech data can fully reflect the characteristics of speech data in the frequency domain, as well as the relationship between the frequency and energy of speech data, analyzing the speech data based on its spectrum can yield relatively accurate speech features.

[0106] The computer device can obtain the spectrum of the speech data by performing a Fourier transform or wavelet transform on the speech data. Optionally, the Fourier transform can be a Fast Fourier Transform, which greatly reduces the number of multiplications required for the computer device to calculate the Fourier transform, thereby saving processing resources.

[0107] Specifically, step (1) can be performed as follows: the computer device divides the speech data into multiple speech segments; for any one of the multiple speech segments, the spectrum of the speech segment is filtered to obtain filtering information; based on the filtering information, the speech segment features of the speech segment are determined; the speech segment features of the multiple speech segments are spliced ​​together to obtain the speech features of the speech data.

[0108] Similarly, when extracting speech features from this speech data, the speech data is divided into multiple speech segments, and then the spectrum of the speech segments is processed to obtain the speech features of the speech data. In this way, processing resources can be saved and processing efficiency can be improved.

[0109] Optionally, the speech segment can be preprocessed before the computer device determines the spectrum of the speech segment, thereby facilitating subsequent processing operations and reducing the difficulty of subsequent processing.

[0110] The computer equipment can perform the following preprocessing operations on the speech segment: pre-emphasize the speech segment; perform frame segmentation on the pre-emphasized speech segment; and then perform windowing on each speech signal obtained from the frame segmentation to increase the continuity between the two ends of the speech signal.

[0111] One method for computer equipment to filter the spectrum of a speech segment is to input the spectrum of the speech segment into a triangular bandpass filter and output the filtered information.

[0112] This triangular bandpass filter is a set of triangular filter banks, which includes K triangular filters. Optionally, the value of K can be between 22 and 26, for example, K can be set to 25.

[0113] In this scenario, computer equipment can smooth the spectrum of a speech segment and eliminate harmonics to highlight the formants of the speech data. Generally, the spectrum of speech data has an envelope and fine structure, corresponding to timbre and pitch, respectively. For speech feature extraction, timbre is the primary useful information. Therefore, by inputting the spectrum of a speech segment into a triangular bandpass filter, the fine structure can be eliminated, retaining only the envelope structure, that is, preserving the timbre information.

[0114] The operation of determining the speech segment features based on the filtering information can be as follows: the computer device determines the logarithmic energy of the filtering information; the discrete cosine transform (DCT) is performed on the logarithmic energy of the filtering information to obtain the cepstral coefficients corresponding to the speech segment; and the speech segment features are determined based on the cepstral coefficients corresponding to the speech segment.

[0115] The computer equipment determines the logarithmic energy of the filtered information by performing a logarithmic operation on it. This logarithmic operation includes taking the absolute value and taking the logarithm. Since the phase information of the spectrum is not valuable in the feature extraction process of speech data, the absolute value of the filtered information is calculated first, retaining only the amplitude value and ignoring the phase effect. Furthermore, a crucial feature in speech is volume (i.e., energy), which can be obtained through logarithmic operations. Therefore, taking the logarithm of the absolute value of the filtered information yields this speech feature.

[0116] Since the standard cepstral coefficients obtained after discrete cosine transform only reflect the static features of speech data, it is necessary to extract the dynamic features of the speech data to ensure the accuracy of the extracted speech features. The dynamic features of speech data can be described by the difference spectrum of these static features. In this case, the operation of a computer device to determine the speech segment features based on the cepstral coefficients corresponding to a speech segment can be as follows: determine the first and second differences of the cepstral coefficients corresponding to the speech segment, and determine the cepstral coefficients (static features), the first and second differences of the cepstral coefficients (dynamic features), and the logarithmic energy of the speech segment as the speech segment features.

[0117] Thus, the speech segment features of this speech segment integrate the static, dynamic, and energy features of the speech data, making the speech segment features of this speech segment more accurate, and consequently, the speech features of the speech data obtained from this will also be more accurate.

[0118] The operation of the computer device to determine the first and second differences of the cepstral coefficients corresponding to the speech segment is similar to the operation of determining the first and second differences of a certain cepstral coefficient in related technologies, and will not be described in detail in the embodiments of this application.

[0119] In this scenario, the computer extracts the MFCCs (Mel-scale Frequency Cepstral Coefficients) of the speech data as its speech features. Since Mel-scale cepstral coefficients take into account human auditory characteristics, and the human auditory system can perceive linear spectra, Mel-scale cepstral coefficients first map the linear spectrum to a Mel-scale nonlinear spectrum based on auditory perception, and then transform it to the cepstral spectrum. Thus, Mel-scale cepstral coefficients can more accurately represent the speech features of the speech data.

[0120] For example: Figure 2 A flowchart for extracting speech features from speech data. See also... Figure 2 , Figure 2 This includes steps 201-207.

[0121] Step 201: The computer device divides the voice data into multiple voice segments and preprocesses each voice segment.

[0122] Step 202: The computer device performs a Fast Fourier Transform on the preprocessed speech segment to obtain the spectrum of the speech segment.

[0123] Step 203: The computer device inputs the spectrum of the speech segment into a triangular bandpass filter to obtain filtering information.

[0124] Step 204: The computer device determines the logarithmic energy of the filtered information, that is, performs a logarithmic operation on the filtered information.

[0125] Step 205: The computer device performs a discrete cosine transform on the logarithmic energy of the filtered information to obtain the cepstral coefficients.

[0126] Step 206: The computer device extracts the dynamic features of the speech segment, which involves calculating the first and second differences of the cepstral coefficients. Then, the logarithmic energy, cepstral coefficients, and the first and second differences of the cepstral coefficients are determined as the speech segment features.

[0127] Step 207: The computer device splices the speech segment features of the multiple speech segments to obtain the speech features of the speech data.

[0128] The above method can obtain the speech features of the speech data by determining the MFCC features of the speech data. Of course, computer devices can also use other methods to characterize the speech features of the speech data, such as the Bark spectrum, LPC (Linear Predictive Coding) features, etc., as the speech features of the speech data.

[0129] Optionally, after dividing the speech data into multiple speech segments, the computer device can also extract the speech segment features of the multiple speech segments using a neural network.

[0130] Specifically, for any one of the multiple speech segments, the speech segment is input into the speech feature extraction model, and the speech feature extraction model is used to extract features from the speech segment and output the speech segment features.

[0131] It is worth noting that before a computer device inputs a speech segment into a speech feature extraction model, extracts features from the speech segment through the speech feature extraction model, and outputs the speech segment features, it also needs to train the speech feature extraction model.

[0132] Specifically, the computer device can acquire multiple first training samples and use these multiple first training samples to train the neural network model to obtain the speech feature extraction model.

[0133] The plurality of first training samples can be pre-set. Each of the plurality of first training samples includes sample data, which can be a segment of speech data of a sample object.

[0134] This neural network model can include multiple network layers, including an input layer, multiple hidden layers, and an output layer. The input layer is responsible for receiving input data; the output layer is responsible for outputting the processed data; the multiple hidden layers are located between the input and output layers and are responsible for processing the data. These hidden layers are not visible to the outside world. For example, this neural network model can be an unsupervised pre-trained model, such as the wav2vec model.

[0135] In this process, when a computer device trains a neural network model using multiple first training samples, for each of these first training samples, the input data from that first training sample is input into the neural network model to obtain output data. A loss function is then used to determine the loss value between the output data and the sample labels in that first training sample. The parameters of the neural network model are then adjusted based on this loss value. After adjusting the parameters of the neural network model based on each of these first training samples, the adjusted neural network model is the speech feature extraction model.

[0136] The operation of adjusting the parameters in the neural network model based on the loss value by the computer device can refer to relevant technologies, and will not be described in detail in the embodiments of this application.

[0137] For example, computer equipment can use formulas This allows for the adjustment of any parameter in the neural network model. is the adjusted parameter. w is the parameter before adjustment. α is the learning rate, which can be preset, such as 0.001, 0.000001, etc., and this application does not limit this to a single value. dw is the partial derivative of the loss function with respect to w, which can be obtained from the loss value.

[0138] (2) The computer device inputs the spectrogram into the image feature extraction model, extracts features from the spectrogram through the image feature extraction model, and outputs the image features of the spectrogram.

[0139] Optionally, the image feature extraction model may include multiple network layers, including an input layer, multiple hidden layers, and an output layer. The input layer is responsible for receiving input data; the output layer is responsible for outputting the processed data; the multiple hidden layers are not visible to the outside world and are responsible for processing the data.

[0140] Optionally, the hidden layers may include multiple convolutional layers, pooling layers, and activation layers. After the input layer receives the image (spectral graph) corresponding to the speech data, the convolutional layers, pooling layers, and activation layers can perform convolution, pooling, and activation operations on the spectrogram to extract its image features. After processing the spectrogram by multiple hidden layers, the features extracted by the last hidden layer can be output.

[0141] From the first hidden layer to the last, the features extracted from the spectrogram become increasingly refined. In other words, the receptive field of the image feature extraction model increases from the first to the last hidden layer, enabling it to extract detailed features from the spectrogram. Therefore, by outputting the features extracted from the last hidden layer out of multiple hidden layers, a more accurate image feature of the spectrogram can be obtained.

[0142] Optionally, the image feature extraction model can be a deep neural network model, and can be a convolutional neural network (CNN) model, etc., which is not limited in this application embodiment.

[0143] (3) The computer device inputs the text into the text feature extraction model, extracts features from the text through the text feature extraction model, and outputs the text features.

[0144] Optionally, the text feature extraction model can be a word vector-based text feature extraction model, for example, the text feature extraction model can be a word2vector model.

[0145] Thus, based on the above steps (1)-(3), the speech features, spectrogram image features, and text features of the speech data can be obtained. The computer can then identify the emotion of the target object, that is, perform the following step 104.

[0146] Step 104: The computer device determines the emotion of the target object based on voice features, image features, and text features.

[0147] Since the features of these three modalities can more comprehensively represent the emotional characteristics of the target object when only the voice data of the target object is acquired, computer devices can determine the emotion of the target object more accurately based on voice features, image features and text features.

[0148] One possible implementation involves a computer device inputting speech features, image features, and text features into an emotion recognition model. This model then makes predictions based on the speech features, image features, and text features, and outputs the emotion recognition result for the target object.

[0149] The emotion recognition result for the target object is the emotion reflected in the target object's speech data. For example, emotions can include positive emotions (happiness, joy, surprise, anticipation, etc.), negative emotions (sadness, grief, anger, fear, etc.), and other emotions. The emotion recognition result for the target object can be one of these emotions.

[0150] It is worth noting that before a computer device inputs speech features, image features, and text features into an emotion recognition model, and the model makes predictions based on these features and outputs the emotion recognition result for the target object, the emotion recognition model needs to be trained.

[0151] Specifically, the computer device can acquire multiple second training samples and use these multiple second training samples to train the neural network model to obtain the emotion recognition model.

[0152] The plurality of second training samples can be pre-set. Each of the plurality of second training samples includes sample data and sample labels. The sample data contains speech features, image features, and text features of the speech data of the sample object, and the sample labels are the emotion recognition results of the sample object corresponding to the sample data. That is, the input data in each of the plurality of second training samples is speech features, image features, and text features of the speech data of the sample object, and the sample labels are the emotion recognition results of the sample object corresponding to the sample data.

[0153] Specifically, the operation of the computer device using multiple second training samples to train the neural network model is similar to the operation of the computer device using multiple first training samples to train the neural network model described above, and will not be described in detail in the embodiments of this application.

[0154] Optionally, after the computer device inputs speech features, image features, and text features into the emotion recognition model, the emotion recognition model can first concatenate the speech features, image features, and text features of the speech data to obtain multimodal features of the speech data. In this case, the emotion recognition model can make predictions based on the multimodal features of the speech data to output the emotion recognition result of the target object.

[0155] Since speech features, image features, and text features are different types of data, these three types of features are concatenated to obtain the multimodal features of the speech data. This facilitates the subsequent processing of these multimodal features by the emotion recognition model, thereby improving the processing efficiency of the emotion recognition model.

[0156] The emotion recognition model predicts based on the multimodal features of the speech data, and the operation of outputting the emotion recognition result of the target object may include the following steps (1) and (2).

[0157] (1) The multimodal features of the speech data are encoded based on the attention mechanism to obtain the attention features of the speech data.

[0158] Attention mechanisms can capture global and local relationships in data from a large number of features and highlight important features of the data.

[0159] In this case, encoding the multimodal features of the speech data based on the attention mechanism can yield the global and local correlations of these multimodal features, identify the more important features among them, and thus obtain the attention features of the speech data. This can then help the emotion recognition model to complete the recognition task effectively and efficiently.

[0160] Optionally, the attention mechanism can be a self-attention mechanism, a multi-head attention mechanism, etc., and this application embodiment does not limit it.

[0161] For example, if the attention mechanism is a multi-head attention mechanism, then the multimodal features of the speech data can be encoded based on the multi-head attention mechanism to obtain the attention features of the speech data.

[0162] Multi-head attention is an improved attention mechanism that maps the multimodal features of speech data to different subspaces through various linear transformations. Attention features are then extracted from these different subspaces. Finally, these extracted attention features are concatenated to obtain a fused attention feature set.

[0163] Thus, by extracting the attention features of this multimodal feature from different perspectives, the accuracy of the attention features obtained from the speech data can be improved.

[0164] (2) The emotion recognition model predicts based on the attention features of the speech data and outputs the emotion recognition result of the target object.

[0165] Since the attention feature of this speech data is a relatively important feature among the multimodal features, the emotion recognition model can make more accurate emotion recognition results for the target object by making predictions based on the attention feature of this speech data.

[0166] Specifically, step (2) can be performed as follows: using the emotion recognition model, the attention features of the speech data are sequence encoded to obtain the sequence encoded features of the speech data; the sequence encoded features are normalized to output the probability of the target object corresponding to multiple candidate emotions; and the candidate emotion with the highest probability among the multiple candidate emotions is determined as the emotion of the target object.

[0167] The multiple candidate emotions refer to all possible emotions.

[0168] Specifically, the attention features of the speech data are sequence encoded. That is, the emotion recognition model inputs the attention features of the speech data into a recursive module, so that the attention features of the speech data are processed by the recursive module to obtain the sequence encoded features of the speech data.

[0169] The recursive module can be a module composed of a recursive neural network (RNN). For example, the recursive neural network can be an LSTM (Long Short-Term Memory) network, a BLSTM (Bi-Long Short-Term Memory) network, or a GRU (Gate Recurrent Unit) network. The embodiments of this application do not limit the type of recursive neural network.

[0170] In this embodiment of the application, by normalizing the sequence encoding features, all output values ​​can be converted into values ​​between 0 and 1, that is, the probability values ​​of the multiple candidate emotions can be obtained, and the emotion of the target object can be determined based on the probability values ​​of the multiple candidate emotions.

[0171] Optionally, the emotion recognition model can input the sequence-encoded features into the softmax (normalization exponent) function to normalize the sequence-encoded features.

[0172] To facilitate understanding, the following will be combined with... Figure 3 The following is an exemplary description of the emotion recognition method provided in the embodiments of this application.

[0173] Assuming that the emotion recognition method is applied in a scenario where a user interacts with a voice assistant, the emotion recognition method provided in this application embodiment may include the following steps 301-311.

[0174] Step 301: The voice assistant obtains the user's voice data.

[0175] Step 302: The voice assistant converts the user's voice data into text.

[0176] Step 303: The voice assistant processes the user's voice data to obtain the spectrogram corresponding to the voice data.

[0177] Step 304: The voice assistant extracts features from the obtained text to obtain the text features.

[0178] Step 305: The voice assistant extracts features from the obtained spectrogram to obtain the image features of the spectrogram.

[0179] Step 306: The voice assistant extracts features from the user's voice data to obtain the voice features of the voice data.

[0180] Specifically, the voice assistant can extract the MFCC features from the user's voice data and use the MFCC features as the voice features of the voice data.

[0181] Step 307: The voice assistant performs feature concatenation on the text features, image features, and speech features to obtain the multimodal features of the speech data.

[0182] Step 308: The voice assistant inputs the multimodal features of the voice data into the emotion recognition model, and the emotion recognition model performs attention encoding on the multimodal features to obtain the attention features of the voice data.

[0183] Step 309: Sequence encoding of the attention features of the speech data is performed using an emotion recognition model to obtain sequence encoded features.

[0184] Step 310: Input the obtained sequence encoding features into the softmax function through the emotion recognition model to output the probabilities of multiple candidate emotions.

[0185] Step 311: Output the emotion recognition result of the target object through the emotion recognition model. The emotion recognition result of the target object is the candidate emotion with the highest probability value among multiple candidate emotions.

[0186] In this embodiment, a computer device acquires the speech data of a target object, processes the speech data to obtain a spectrogram and text corresponding to the speech data, and then extracts features from the speech data, spectrogram, and text to obtain speech features, image features, and text features. Thus, by processing only the speech data of the target object, features in three modalities—image, speech, and text—can be obtained. These three modal features can more comprehensively represent the emotional characteristics of the target object. Subsequently, based on the speech features, image features, and text features, the emotion of the target object is determined. By comprehensively determining the emotion of the target object using features from these three modalities, the accuracy of emotion recognition of the target object can be improved.

[0187] Figure 4This is a schematic diagram of the structure of an emotion recognition device provided in an embodiment of this application. The emotion recognition device can be implemented as part or all of a computer device by software, hardware, or a combination of both. This computer device can be as described below. Figure 5 The computer equipment shown. See also Figure 4 The device includes: an acquisition module 401, a processing module 402, a feature extraction module 403, and a determination module 404.

[0188] Acquisition module 401 is used to acquire the voice data of the target object;

[0189] Processing module 402 is used to process the speech data to obtain the spectrogram and text corresponding to the speech data;

[0190] The feature extraction module 403 is used to extract features from the speech data to obtain the speech features of the speech data; to extract features from the spectrogram to obtain the image features of the spectrogram; and to extract features from the text to obtain the text features of the text.

[0191] The determination module 404 is used to determine the emotion of the target object based on speech features, image features and text features.

[0192] Optionally, the processing module 402 is used for:

[0193] The speech data is divided into multiple speech segments;

[0194] Perform time-frequency transformation on the multiple speech segments to obtain the spectrograms corresponding to the speech data;

[0195] Text recognition is performed on the multiple speech segments to obtain the text corresponding to the speech data.

[0196] Optionally, the processing module 402 is used for:

[0197] For any one of the multiple speech segments, perform a Fourier transform or wavelet transform on the speech segment to obtain the target spectrum; based on the target spectrum, generate the spectrogram corresponding to the speech segment.

[0198] The spectrograms of the multiple speech segments are stitched together to obtain the image corresponding to the speech data.

[0199] Optionally, the feature extraction module 403 is used for:

[0200] The speech data is divided into multiple speech segments;

[0201] For any one of the multiple speech segments, the spectrum of the speech segment is filtered to obtain filtering information; based on the filtering information, the speech segment features of the speech segment are determined.

[0202] The speech features of the multiple speech segments are concatenated to obtain the speech features of the speech data.

[0203] Optionally, the determining module 404 includes:

[0204] The feature concatenation unit is used to concatenate speech features, image features, and text features to obtain the multimodal features of the speech data.

[0205] The determining unit is used to determine the emotion of the target object based on the multimodal features of the speech data through the emotion recognition model.

[0206] Optionally, the determining unit is used for:

[0207] The multimodal features of the speech data are encoded based on the attention mechanism to obtain the attention features of the speech data;

[0208] Based on the attentional features of voice data, determine the emotions of the target audience.

[0209] Optionally, the determining unit is used for:

[0210] The attention features of the speech data are sequence encoded to obtain the sequence encoded features of the speech data; the sequence encoded features are normalized to obtain the probability of the target object corresponding to multiple candidate emotions; the candidate emotion with the highest probability among the multiple candidate emotions is determined as the emotion of the target object.

[0211] In this embodiment, speech data of the target object is acquired, processed to obtain a spectrogram and text corresponding to the speech data, and then feature extraction is performed on the speech data, spectrogram, and text to obtain speech features, image features, and text features. Thus, features in three modalities—image, speech, and text—can be obtained simply by processing the speech data of the target object. These three modal features can more comprehensively represent the emotional characteristics of the target object. Subsequently, based on the speech features, image features, and text features, the emotion of the target object is determined. By comprehensively determining the emotion of the target object using features from these three modalities, the accuracy of emotion recognition for the target object can be improved.

[0212] It should be noted that the emotion recognition device provided in the above embodiments is only illustrated by the division of the above functional modules when recognizing the emotions of the target object. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0213] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.

[0214] The emotion recognition device and emotion recognition method embodiments provided in the above embodiments belong to the same concept. The specific working process and technical effects of the units and modules in the above embodiments can be found in the method embodiment section, and will not be repeated here.

[0215] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 5 As shown, the computer device 5 includes a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50. When the processor 50 executes the computer program 52, it implements the steps in the emotion recognition method in the above embodiments.

[0216] Computer device 5 can be a general-purpose computer device or a special-purpose computer device. In specific implementations, computer device 5 can be a desktop computer, portable computer, handheld computer, mobile phone, tablet computer, or other device with human-computer interaction functions. This application embodiment does not limit the type of computer device 5. Those skilled in the art will understand that... Figure 5 The computer device 5 is merely an example and does not constitute a limitation on the computer device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0217] Processor 50 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0218] In some embodiments, memory 51 may be an internal storage unit of the computer device 5, such as a hard disk or RAM of the computer device 5. In other embodiments, memory 51 may be an external storage device of the computer device 5, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, etc., provided on the computer device 5. Furthermore, memory 51 may include both internal and external storage units of the computer device 5. Memory 51 is used to store the operating system, applications, boot loader, data, and other programs. Memory 51 may also be used to temporarily store data that has been output or will be output.

[0219] This application also provides a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.

[0220] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the various method embodiments described above.

[0221] This application provides a computer program product that, when run on a computer, causes the computer to perform the steps described in the various method embodiments above.

[0222] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above method embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices. The computer-readable storage medium mentioned in this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.

[0223] It should be understood that all or part of the steps of the above embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-described computer-readable storage medium.

[0224] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0225] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0226] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0227] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0228] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An emotion recognition method, characterized in that, The method includes: Acquire the voice data of the target object; The speech data is processed to obtain the spectrogram and text corresponding to the speech data, wherein the spectrogram represents the frequency of the speech data changing over time. The speech data is subjected to feature extraction to obtain speech features; the spectrogram is subjected to feature extraction to obtain image features; and the text is subjected to feature extraction to obtain text features. Based on the speech features, image features, and text features, the emotion of the target object is determined; The step of determining the emotion of the target object based on the speech features, the image features, and the text features includes: The attention features of the speech data are sequence encoded to obtain the sequence encoded features of the speech data; the sequence encoded features are normalized to obtain the probability that the target object corresponds to multiple candidate emotions; the candidate emotion with the highest probability among the multiple candidate emotions is determined as the emotion of the target object. The attention features of the speech data are obtained by encoding the multimodal features of the speech data based on the attention mechanism. The multimodal features of the speech data are obtained by concatenating the speech features, the image features, and the text features.

2. The method as described in claim 1, characterized in that, The process of processing the speech data to obtain the spectrogram and text corresponding to the speech data includes: The voice data is divided into multiple voice segments; Perform time-frequency transformation on the multiple speech segments to obtain the spectrogram corresponding to the speech data; Text recognition is performed on the multiple speech segments to obtain the text corresponding to the speech data.

3. The method as described in claim 2, characterized in that, The step of performing time-frequency transformation on the plurality of speech segments to obtain the spectrogram corresponding to the speech data includes: For any one of the plurality of speech segments, perform a Fourier transform or wavelet transform on the speech segment to obtain a target spectrum; based on the target spectrum, generate a spectrogram corresponding to the speech segment; The spectrograms of the multiple speech segments are spliced ​​together to obtain the spectrogram corresponding to the speech data.

4. The method as described in claim 1, characterized in that, The step of extracting features from the speech data to obtain the speech features of the speech data includes: The voice data is divided into multiple voice segments; For any one of the plurality of speech segments, the spectrum of the speech segment is filtered to obtain filtering information; based on the filtering information, the speech segment features of the speech segment are determined. The speech features of the multiple speech segments are concatenated to obtain the speech features of the speech data.

5. The method as described in any one of claims 1 to 4, characterized in that, Determining the emotion of the target object based on the speech features, image features, and text features includes: The speech features, image features, and text features are concatenated to obtain the multimodal features of the speech data; The emotion of the target object is determined based on the multimodal features of the speech data.

6. The method as described in claim 5, characterized in that, Determining the emotion of the target object based on the multimodal features of the speech data includes: The multimodal features of the speech data are encoded based on an attention mechanism to obtain the attention features of the speech data; Based on the attentional features of the voice data, the emotion of the target object is determined.

7. An emotion recognition device, characterized in that, The device includes: The acquisition module is used to acquire the voice data of the target object; The processing module is used to process the speech data to obtain the spectrogram and text corresponding to the speech data, wherein the spectrogram represents the frequency of the speech data changing over time. The feature extraction module is used to extract features from the speech data to obtain the speech features of the speech data; to extract features from the spectrogram to obtain the image features of the spectrogram; and to extract features from the text to obtain the text features of the text. The determination module is used to determine the emotion of the target object based on the speech features, the image features, and the text features; The determining module is specifically used to perform sequence encoding on the attention features of the speech data to obtain the sequence encoded features of the speech data; normalize the sequence encoded features to obtain the probability that the target object corresponds to multiple candidate emotions; determine the candidate emotion with the highest probability among the multiple candidate emotions as the emotion of the target object. The attention features of the speech data are obtained by encoding the multimodal features of the speech data based on the attention mechanism, and the multimodal features of the speech data are obtained by feature concatenation based on the speech features, the image features, and the text features.

8. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 6.