Artificial intelligence-based emotion recognition method, device, computer equipment and medium

By introducing confidence to fuse speech and text features in speech emotion recognition, the problem of inaccurate recognition of speech emotion recognition models in high-noise environments is solved, and higher recognition accuracy and robustness are achieved.

CN114974310BActive Publication Date: 2025-10-03PING AN TECH (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210602736.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-10-03
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

Existing speech emotion recognition models lack recognition accuracy and robustness when faced with sparse and high-noise speech emotion data, especially because translation errors in automatic speech recognition technology affect the accuracy of emotion recognition.

Method used

By using the trained speech recognition network to translate the speech data into text data, the confidence of the text data is calculated and the acoustic feature vector of the speech data is obtained, the linguistic feature vector is extracted, and the two are fused and input into the trained emotion classification network to determine the emotion category.

Benefits of technology

The error tolerance and robustness of the emotion recognition model to translation accuracy are improved, especially in high-noise environments, which improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974310B_ABST
    Figure CN114974310B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of artificial intelligence technology, and in particular relates to an emotion recognition method, device, computer equipment and medium based on artificial intelligence. The method translates the acquired speech data into text data, calculates the confidence of the text data, obtains the acoustic feature vector of the speech data, extracts the linguistic feature vector of the text data, fuses the acoustic feature vector with the linguistic feature vector using the confidence, inputs the feature fusion result into a trained emotion classification network, outputs the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, determines the preset emotion classification whose probability meets the preset conditions as the emotion recognition result of the speech data, introduces the recognition result confidence to fuse the linguistic features and acoustic features of the text, so that the method has a certain fault tolerance for speech recognition translation, has a certain robustness for high-noise speech, and improves the recognition accuracy of the emotion recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application is applicable to the field of artificial intelligence technology, and in particular relates to an emotion recognition method, apparatus, computer equipment, and medium based on artificial intelligence. Background Art

[0002] Currently, speech emotion recognition involves the task of identifying the emotions in a speech stream. This can effectively improve the intelligence of robots during human-computer interaction, thereby enhancing the customer experience. With the development of deep learning technology, speech emotion recognition accuracy has made significant progress. However, due to the fuzzy boundaries of speech emotion categories and the resulting difficulty in data annotation, the amount of speech emotion annotation data is relatively small, and most of it consists of the emotional expressions of specific actors, which differ from human emotions in daily life. Therefore, the task of simply classifying speech emotion is quite challenging. Due to the sparsity of speech emotion data, speech emotion recognition currently relies primarily on automatic speech recognition (ASR) technology to translate speech into text, then combine the text and speech for comprehensive emotion recognition. However, this process is limited by the accuracy of speech recognition translation. Text with ASR translation errors can affect the accuracy of the joint modeling. The recognition process has poor tolerance for translation errors, resulting in low robustness of the recognition model and ultimately inaccurate recognition. Therefore, how to improve the emotion recognition process to enhance the error tolerance and robustness of the emotion recognition model to translation accuracy has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, the embodiments of the present application provide an artificial intelligence-based emotion recognition method, apparatus, computer equipment and medium to solve the problem of how to improve the emotion recognition process to enhance the fault tolerance and robustness of the emotion recognition model for translation accuracy.

[0004] In a first aspect, an embodiment of the present application provides an emotion recognition method based on artificial intelligence, the emotion recognition method comprising:

[0005] Use the trained speech recognition network to translate the acquired speech data into text data;

[0006] Calculating the confidence of the text data, and obtaining the acoustic feature vector of the speech data in the process of translating the speech data into the text data;

[0007] Extracting a linguistic feature vector of the text data, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a feature fusion result;

[0008] The feature fusion result is input into a trained emotion classification network, and the probability of the speech data being divided into each preset emotion category in the trained emotion classification network is output, and the preset emotion classification whose probability meets the preset conditions is determined as the emotion recognition result of the speech data.

[0009] In one embodiment, if the confidence level includes a confidence level for each word in the text data, the acoustic feature vector and the linguistic feature vector are subjected to feature fusion using the confidence level, and the feature fusion result obtained includes:

[0010] The confidence of each word is dot-multiplied by the acoustic feature vector, and the result is combined with the linguistic feature vector into an array, and the array is determined to be a feature fusion result.

[0011] In one embodiment, using a trained speech recognition network to translate the acquired speech data into text data includes:

[0012] Extracting Fbank features from the acquired speech data, and calculating the acoustic feature vector of each frame of speech based on the Fbank features;

[0013] The acoustic feature vector of each frame of speech is matched with the words in the vocabulary, and the matched words are serialized to obtain text data.

[0014] In one embodiment, calculating the confidence level of the text data includes:

[0015] For any target word in the text data, determining a start frame number and an end frame number corresponding to the target word in the acoustic feature vector;

[0016] The probability of outputting the target word under the condition of the acoustic feature vector corresponding to each frame number between the starting frame number and the ending frame number is calculated, and the average value of all probabilities is determined as the confidence of the target word.

[0017] In one embodiment, extracting the linguistic feature vector of the text data includes:

[0018] Extracting a character feature vector and a position feature vector of each character in the text data to obtain a semantic feature vector of each character;

[0019] The semantic feature vector of each word is input into the transformer model, and the output feature is the linguistic feature vector corresponding to the text data.

[0020] In one embodiment, the emotion classification network includes a two-layer feedforward neural network layer and a one-layer softmax layer, and uses a cross entropy function as a loss function; the speech recognition network includes a transformer model and a one-layer feedforward neural network layer, and uses CTC as a loss function, and the emotion classification network and the speech recognition network are jointly trained;

[0021] The joint training process is:

[0022] Use the speech recognition network to be trained to translate the training speech into training text and calculate the CTC loss;

[0023] Calculating the confidence of the training text and obtaining the acoustic feature vector of the training speech during the process of translating the training speech into the training text;

[0024] Extracting the linguistic feature vector of the training text, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a training feature fusion result;

[0025] The feature fusion result of the training is input into the emotion classification network to be trained, the emotion recognition result of the training is output and the cross entropy loss is calculated with the annotation result of the training speech, and the parameters of the speech recognition network to be trained and the parameters of the emotion classification network to be trained are reversely updated using the gradient descent method. It is iterated until the sum of the cross entropy loss and the CTC loss converges, thereby obtaining the parameters of the trained speech recognition network and the trained emotion classification network.

[0026] In a second aspect, an embodiment of the present application provides an emotion recognition device based on artificial intelligence, the emotion recognition device comprising:

[0027] A speech recognition module is used to translate the acquired speech data into text data using a trained speech recognition network;

[0028] A confidence calculation module, used to calculate the confidence of the text data;

[0029] A vector acquisition module, configured to acquire an acoustic feature vector of the speech data during the process of translating the speech data into the text data;

[0030] a feature fusion module, configured to extract the linguistic feature vector of the text data, and perform feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a feature fusion result;

[0031] The emotion recognition module is used to input the feature fusion result into a trained emotion classification network, output the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, and determine the preset emotion classification whose probability meets the preset conditions as the emotion recognition result of the speech data.

[0032] In one embodiment, if the confidence level includes a confidence level for each word in the text data, the feature fusion module includes:

[0033] The feature fusion unit is used to perform dot multiplication of the acoustic feature vector with the confidence of each word, and merge the acoustic feature vector with the linguistic feature vector into an array, and determine the array as the feature fusion result.

[0034] In one embodiment, the speech recognition module includes:

[0035] An acoustic vector extraction unit is used to extract Fbank features from the acquired speech data and calculate the acoustic feature vector of each frame of speech based on the Fbank features;

[0036] The text matching unit is used to match the acoustic feature vector of each frame of speech with the words in the word list, and serialize the matched words to obtain text data.

[0037] In one embodiment, the confidence calculation module includes:

[0038] a frame number determining unit, configured to determine, for any target word in the text data, a start frame number and an end frame number corresponding to the target word in the acoustic feature vector;

[0039] The confidence determination unit is used to calculate the probability of outputting the target word under the condition of the acoustic feature vector corresponding to each frame number between the starting frame number and the ending frame number, and determine the average value of all probabilities as the confidence of the target word.

[0040] In one embodiment, the feature fusion module includes:

[0041] a semantic vector determination unit, configured to extract a character feature vector and a position feature vector of each character in the text data to obtain a semantic feature vector of each character;

[0042] The linguistic vector output unit is used to input the semantic feature vector of each word into the transformer model, and the output feature is the linguistic feature vector corresponding to the text data.

[0043] In one embodiment, the emotion classification network includes a two-layer feedforward neural network layer and a one-layer softmax layer, and uses a cross entropy function as a loss function; the speech recognition network includes a transformer model and a one-layer feedforward neural network layer, and uses CTC as a loss function, and the emotion classification network and the speech recognition network are jointly trained;

[0044] The joint training process is:

[0045] Use the speech recognition network to be trained to translate the training speech into training text and calculate the CTC loss;

[0046] Calculating the confidence of the training text and obtaining the acoustic feature vector of the training speech during the process of translating the training speech into the training text;

[0047] Extracting the linguistic feature vector of the training text, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a training feature fusion result;

[0048] The feature fusion result of the training is input into the emotion classification network to be trained, the emotion recognition result of the training is output and the cross entropy loss is calculated with the annotation result of the training speech, and the parameters of the speech recognition network to be trained and the parameters of the emotion classification network to be trained are reversely updated using the gradient descent method. It is iterated until the sum of the cross entropy loss and the CTC loss converges, thereby obtaining the parameters of the trained speech recognition network and the trained emotion classification network.

[0049] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the emotion recognition method as described in the first aspect when executing the computer program.

[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the emotion recognition method as described in the first aspect is implemented.

[0051] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: the present application uses a trained speech recognition network to translate the acquired speech data into text data, calculates the confidence of the text data, and obtains the acoustic feature vector of the speech data in the process of translating the speech data into text data, extracts the linguistic feature vector of the text data, and uses the confidence to perform feature fusion on the acoustic feature vector and the linguistic feature vector to obtain a feature fusion result, inputs the feature fusion result into a trained emotion classification network, outputs the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, determines the preset emotion classification whose probability meets the preset conditions as the emotion recognition result of the speech data, introduces the recognition result confidence in the process to fuse the linguistic features and acoustic features of the text, so that the method has a certain fault tolerance for speech recognition translation, has a certain robustness for high-noise speech, and improves the recognition accuracy of the emotion recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0053] Figure 1 This is a schematic diagram of an application environment of an artificial intelligence-based emotion recognition method provided in Example 1 of the present application;

[0054] Figure 2 This is a flow chart of an emotion recognition method based on artificial intelligence provided in Example 2 of the present application;

[0055] Figure 3 This is a flowchart of an emotion recognition method based on artificial intelligence provided in Example 3 of the present application;

[0056] Figure 4 This is a schematic diagram of the structure of an emotion recognition device based on artificial intelligence provided in Example 4 of the present application;

[0057] Figure 5 This is a structural diagram of a computer device provided in Example 5 of the present application. DETAILED DESCRIPTION

[0058] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0059] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0060] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0061] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0062] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0063] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0064] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0065] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0066] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0067] In order to illustrate the technical solution of the present application, specific embodiments are provided below.

[0068] The first embodiment of the present application provides an emotion recognition method based on artificial intelligence, which can be applied in Figure 1 In an application environment, a client communicates with a server. Clients include, but are not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0069] See also Figure 2 , is a flow chart of an emotion recognition method based on artificial intelligence provided in Example 2 of this application, the above emotion recognition method is applied to Figure 1 The server in the server, the computer device corresponding to the server connects to the corresponding database to obtain the corresponding voice data in the database. The above computer device can also connect to the corresponding client, and the client sends the voice data to the server, realizing the function of the server obtaining the voice data. Figure 2 As shown, the emotion recognition method based on artificial intelligence may include the following steps:

[0070] Step S201: Using a trained speech recognition network, the acquired speech data is translated into text data.

[0071] In this application, a server connects to a corresponding client, which collects voice and sends it to the server, thereby implementing the server's voice collection step. The client can be a device equipped with voice collection equipment, such as a voice robot or an in-vehicle terminal. In one embodiment, the server retrieves the voice from a corresponding database.

[0072] Before using the trained speech recognition network for speech recognition, the server can also preprocess the speech, including noise reduction, enhancement, and other processing, to ensure the accuracy of subsequent speech recognition.

[0073] The trained speech recognition network can be a learning network built based on ASR technology or other speech recognition technologies. The trained speech recognition network can convert speech into text. The speech-to-text process can include identifying timbre, pitch, and punctuation, obtaining the phonemes or phonetic symbols corresponding to each frame of speech (i.e., the feature vector corresponding to each frame of speech), and then using matching to obtain the corresponding characters, thereby converting the speech into text.

[0074] Optionally, using a trained speech recognition network, translating the acquired speech data into text data includes:

[0075] Extract Fbank features from the acquired speech data and calculate the acoustic feature vector of each frame of speech based on the Fbank features;

[0076] The acoustic feature vector of each frame of speech is matched with the words in the vocabulary, and the matched words are serialized to obtain text data.

[0077] Among them, extracting Fbank features can refer to performing pre-emphasis, framing, windowing, short-time Fourier transform (STFT), mel filtering, de-averaging and other processing on the speech data to finally obtain the feature expression of the speech data.

[0078] The calculation of acoustic feature vectors can use an N-layer transformer structure to encode and decode the Fbank features, and then use a forward neural network to calculate the features to obtain the acoustic feature vector of each frame of speech.

[0079] Serializing the corresponding acoustic feature vectors according to the time sequence of each frame of speech and matching them with the characters in the word list. It should be understood that each character in the word list corresponds to a group of feature vectors, and the matching is a similarity matching. When matching, determine the similarity results of multiple matches in the ways of matching from one frame of speech and continuous multiple frames of speech, and determine the character with the highest similarity. After obtaining all the matching results, serialize all the characters according to the time sequence, and the serialization result is the text data.

[0080] For example, the speech data contains 3 frames of speech, which are divided into the first frame of speech, the second frame of speech and the third frame of speech according to the time sequence. When matching, calculate the similarity between the acoustic feature vectors of the first frame of speech and each group of feature vectors in the word list respectively, calculate the similarity between the continuous speech composed of the first frame of speech and the second frame of speech and each group of feature vectors in the word list respectively, and then calculate the similarity between the continuous speech composed of the first frame of speech, the second frame of speech and the third frame of speech and each group of feature vectors in the word list respectively to obtain all the similarities. Among them, the similarity between the continuous speech composed of the first frame of speech and the second frame of speech and the character "and" in the word list is the highest. Subsequently, calculate the similarity between the acoustic feature vectors of the third frame of speech and each group of feature vectors in the word list respectively, and determine that the similarity with the character "you" is the highest. Finally, the above speech data is translated into the text data "and you".

[0081] Step S202, calculate the confidence of the text data, and obtain the acoustic feature vectors of the speech data during the process of translating the speech data into the text data.

[0082] In this application, after translating the speech data into the text data, it is necessary to calculate the confidence of the text data, and corresponding parameters in the process of translating the speech data into the text data are required for the calculation. For the confidence of each character in the text data, the speech constituting the character includes at least one frame of speech. Therefore, the confidence of the character is the authenticity of each frame of speech constituting the character. In one embodiment, the confidence of the text data can also be the confidence of a word composed of two or more characters. Correspondingly, the process of fusing features can be adjusted according to requirements in subsequent use.

[0083] During the process of translating the speech data into the text data, it is necessary to extract the acoustic feature vectors of the speech data. The acoustic feature vectors are essentially used to characterize the acoustic features of the speech data. The acoustic feature vectors can be the feature vectors of each frame of speech after dividing the speech data into T frames of speech according to a pre-set frame division method. For example, assume the speech data is the sequence X = [x1, x2,..., x i , …, x T , where, x iFurthermore, the above-mentioned pre-set frame division method needs to divide each frame of the speech data into the smallest unit as much as possible, and ideally, one frame represents one phoneme or phonetic symbol.

[0084] The data generated in the above translation process can be directly obtained, or after being generated, this part of the data can be transferred to the corresponding task for caching processing. When the corresponding task is executed, this part of the data is called to participate in the calculation of the task.

[0085] Optionally, calculating the confidence of text data includes:

[0086] For any target word in the text data, determine the start frame number and end frame number corresponding to the target word in the acoustic feature vector;

[0087] The probability of outputting the target word under the condition of the acoustic feature vector corresponding to each frame number between the start frame number and the end frame number is calculated, and the average value of all probabilities is determined as the confidence level of the target word.

[0088] For any target word, the probability that the acoustic feature vector can output the target word within the time period corresponding to the target word is averaged. This means that the probability that the acoustic feature vector corresponding to each frame of speech can output the target word is determined. All probabilities are averaged, and the average value is used as the confidence level of the target word. Therefore, if the probability of the acoustic feature vector corresponding to each frame of speech outputting the target word is 1, then the confidence level is also 1, indicating that the target word is accurately recognized. Conversely, a lower confidence level indicates that the target word is inaccurately recognized.

[0089] If the confidence of each word in the text data is calculated, the calculation formula is as follows:

[0090]

[0091] In the formula, s is the word y j The starting frame number corresponding to the acoustic feature vector, s+d represents y j The end frame number corresponding to the acoustic feature vector, d is y j The number of frames corresponding to the acoustic feature vector, S(y j ) represents the character y j confidence level.

[0092] Step S203 : extracting the linguistic feature vector of the text data, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a feature fusion result.

[0093] In this application, the linguistic feature vector can be a linguistic feature that characterizes text data, that is, an analysis of text data, wherein the analysis of text data can refer to the analysis of the semantics, grammatical structure and pragmatics of the text in the text data based on language processing models such as natural language processing (NLP).

[0094] Linguistic feature vectors can integrate semantic analysis results, grammatical analysis results, pragmatic analysis results, etc. to form a vector representation.

[0095] In this application, emotion recognition for speech integrates acoustic features and linguistic features, which can better represent the true emotional expression of speech. The fusion of features is based on the confidence level as the connection condition, combining the two vectors into a group feature. The fusion here is to convert the two different dimensional information of acoustic features and linguistic features into information with the same dimensionality. For example, the acoustic feature vector is corrected with the confidence level, that is, the acoustic feature vector is converted into one-dimensional information. Since the linguistic feature vector is also one-dimensional, the corrected feature vector is combined with the linguistic feature vector with the corresponding confidence level into a group of feature vectors.

[0096] Optionally, if the confidence level includes a confidence level for each word in the text data, the confidence level is used to perform feature fusion on the acoustic feature vector and the linguistic feature vector, and the feature fusion result includes:

[0097] The confidence of each word is dot-multiplied with the acoustic feature vector, and the result is merged with the linguistic feature vector into an array, which is determined as the feature fusion result.

[0098] Among them, the confidence is the confidence of each word in the text data, and the confidence can be represented by a one-dimensional confidence matrix. Each frame of speech corresponds to an acoustic feature vector. For a word in the text data, the corresponding acoustic feature vector can be one or more. The acoustic feature vector is converted into an acoustic matrix representation. If the confidence matrix is ​​a one-dimensional row matrix, the row of the acoustic matrix corresponds to a word in the text data (that is, a confidence), and the column of the acoustic matrix is ​​the acoustic feature vector corresponding to the word in the row. The number of elements in the column can change randomly. If the acoustic feature vector corresponding to the word in the row is less than the number of elements in the column, the remaining elements are filled with 0. If the confidence matrix is ​​a one-dimensional column matrix, the structure of the acoustic matrix is ​​adjusted accordingly.

[0099] The confidence matrix is ​​dot-multiplied by the acoustic matrix to obtain a matrix after dot product. The matrix after dot product is combined with the linguistic feature vector into one data, which is the feature fusion result.

[0100] For example, the confidence matrix is ​​S = [S(y1)S(y2)...S(y j )...S(yN )], where S(y j ) represents the character y j The confidence of the acoustic matrix is Among them, f j1 Represents the word y j The corresponding acoustic feature vector, the dot product is S⊙feat a The final feature fusion result is feat = Concat (feat t ,S⊙feat a ), among which, feat t represents a linguistic feature vector.

[0101] Optionally, extracting a linguistic feature vector from text data includes:

[0102] Extracting the character feature vector and position feature vector of each character in the text data to obtain the semantic feature vector of each character;

[0103] The semantic feature vector of each word is input into the transformer model, and the output feature is the linguistic feature vector of the corresponding text data.

[0104] Among them, for the linguistic feature vector of text data, the character feature vector and position feature vector of each character are analyzed, and the two are used as the semantic features of the corresponding character. The semantic features are input into the transformer model for encoding and decoding, and the decoding results are passed through a layer of forward neural network for feature calculation to obtain the linguistic feature vector of the text data.

[0105] Step S204: input the feature fusion result into the trained emotion classification network, output the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, and determine the preset emotion category whose probability meets the preset conditions as the emotion recognition result of the speech data.

[0106] In this application, a plurality of preset emotion categories are set in the trained emotion classification network. After the feature fusion result is input into the trained emotion classification network, the output is the probability of each preset emotion category. The probability is used to characterize the degree of association or similarity between the feature fusion result and a preset emotion category. If the probability is high, it means that the feature fusion result should be classified into the corresponding preset emotion category. If the probability is low, it means that the feature fusion result should not be classified into the corresponding preset emotion category.

[0107] Set corresponding preset conditions to determine the final emotion recognition result of the language data. The emotion recognition result is the emotion category determined from all preset emotion categories. The preset condition can be a threshold. Of course, the preset condition can also be the maximum value in the comparison probability. During use, the specific conditions can be set according to actual needs.

[0108] The embodiment of the present application uses a trained speech recognition network to translate the acquired speech data into text data, calculates the confidence of the text data, obtains the acoustic feature vector of the speech data in the process of translating the speech data into text data, extracts the linguistic feature vector of the text data, and uses the confidence to fuse the acoustic feature vector with the linguistic feature vector to obtain a feature fusion result, inputs the feature fusion result into a trained emotion classification network, outputs the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, determines the preset emotion classification whose probability meets the preset conditions as the emotion recognition result of the speech data, introduces the recognition result confidence in the process to fuse the linguistic features and acoustic features of the text, so that the method has a certain fault tolerance for speech recognition translation, has a certain robustness for high-noise speech, and improves the recognition accuracy of the emotion recognition model.

[0109] See also Figure 3 , is a flow chart of an emotion recognition method based on artificial intelligence provided in Example 3 of the present application, such as Figure 3 As shown in FIG, it specifically shows the training process of the emotion recognition method under certain limited conditions, targeting the parts such as the network and model that need to be trained.

[0110] In this application, the emotion classification network is defined as comprising a two-layer feedforward neural network layer and a one-layer softmax layer, and adopting a cross entropy function as a loss function; the speech recognition network is defined as comprising a transformer model and a one-layer feedforward neural network layer, and adopting CTC as a loss function; the emotion classification network and the speech recognition network are jointly trained, and the joint training may include the following steps:

[0111] Step S301: Use the speech recognition network to be trained to translate the training speech into training text and calculate the CTC loss.

[0112] In this application, training speech and annotations of the corresponding training language are obtained from the corresponding database, the training speech is translated into training text using the speech recognition network to be trained, and the CTC loss is calculated.

[0113] The input speech sequence is assumed to be X = [x1, x2, ..., x i ,…,x T ], where T is the length of speech, x iThe feature vector representing the i-th frame of speech, the output text sequence is assumed to be Y = [y1, y2, ..., y j ,…,y N ], where N is the length of the output sequence, y j represents the jth word, L represents all output unit spaces, and defines the expansion space L * =L∪{blank}.

[0114] This defines the following formula:

[0115] P(Y|X)=∑ c∈A(Y) P(C|X)

[0116]

[0117] In the formula, P(C t ,t) indicates that label C is observed at time t t The probability of C t ∈L * , P(C|X) represents the probability of the network output sequence C under the condition of a given input sequence X, A(Y) represents the probability combination of all text sequences Y and the special label blank, C is any subsequence, and P(YX) represents the probability of the network output text sequence Y under the condition of a given input feature sequence X.

[0118] The CTC loss function is thus defined as:

[0119] loss ctc =∑ (X,Y) -logP(Y|X).

[0120] In this application, a feedforward neural network layer can use a greedy-search decoding algorithm to obtain a speech recognition translation result for the decoding result output by the transformer model.

[0121] Step S302 , calculating the confidence of the training text, and obtaining the acoustic feature vector of the training speech during the process of translating the training speech into the training text.

[0122] Step S303 : extracting the linguistic feature vector of the training text, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a training feature fusion result.

[0123] Among them, the contents of steps S302 to S303 are the same as those of steps S202 and S203 in the above embodiment, and reference may be made to the above description, which will not be repeated here.

[0124] Step S304: Input the trained feature fusion result into the emotion classification network to be trained, output the trained emotion recognition result and the annotation result of the training speech to calculate the cross entropy loss.

[0125] In this application, the cross entropy loss function can be expressed as:

[0126]

[0127] Where M represents the total number of emotion categories in the network, e c represents the c-th emotion category, and feat represents the feature fusion result.

[0128] Step S305: Use the gradient descent method to reversely update the parameters of the speech recognition network to be trained and the parameters of the emotion classification network to be trained, and iterate until the sum of the cross entropy loss and the CTC loss converges to obtain the parameters of the trained speech recognition network and the trained emotion classification network.

[0129] In this application, joint training is used, and the total loss of the two networks needs to be calculated. The total loss function is expressed as:

[0130] loss=loss ser +λloss ctc

[0131] Where λ is an adjustable parameter, which is generally set to 0.1.

[0132] Using the gradient descent method to reversely update the parameters can promote loss convergence more quickly, thereby maximizing the training efficiency of the network.

[0133] The embodiment of the present application adopts a joint training method to simultaneously train the speech recognition network and the emotion classification network, so that the above-mentioned emotion recognition method can better fit the network, and finally uses the trained speech recognition network to translate the acquired speech data into text data, calculate the confidence of the text data, and obtain the acoustic feature vector of the speech data in the process of translating the speech data into text data, extract the linguistic feature vector of the text data, and use the confidence to perform feature fusion on the acoustic feature vector and the linguistic feature vector to obtain the feature fusion result, input the feature fusion result into the trained emotion classification network, output the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, determine the preset emotion classification whose probability meets the preset conditions as the emotion recognition result of the speech data, introduce the recognition result confidence in the process to fuse the linguistic features and acoustic features of the text, so that the method has a certain fault tolerance for speech recognition translation, a certain robustness for high-noise speech, and improves the recognition accuracy of the emotion recognition model.

[0134] Corresponding to the emotion recognition method based on artificial intelligence in the above embodiment, Figure 4 The structure diagram of the emotion recognition device based on artificial intelligence provided by the fourth embodiment of the present application is shown. The emotion recognition device is applied to Figure 1 The server in the embodiment of the present invention is connected to a corresponding database by a computer device corresponding to the server to obtain the corresponding voice data in the database. The computer device can also be connected to a corresponding client, and the client sends the voice data to the server, thereby realizing the function of the server obtaining the voice data. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0135] See also Figure 4 , the emotion recognition device comprises:

[0136] The speech recognition module 41 is used to translate the acquired speech data into text data using the trained speech recognition network;

[0137] A confidence calculation module 42, used to calculate the confidence of text data;

[0138] A vector acquisition module 43 is used to obtain acoustic feature vectors of speech data during the process of translating speech data into text data;

[0139] A feature fusion module 44 is used to extract linguistic feature vectors from text data and fuse the acoustic feature vectors with the linguistic feature vectors using confidence to obtain a feature fusion result;

[0140] The emotion recognition module 45 is used to input the feature fusion result into the trained emotion classification network, output the probability that the speech data is divided into each preset emotion category in the trained emotion classification network, and determine the preset emotion classification whose probability meets the preset conditions as the emotion recognition result of the speech data.

[0141] Optionally, if the confidence level includes a confidence level for each word in the text data, the feature fusion module 44 includes:

[0142] The feature fusion unit is used to multiply the acoustic feature vector by the confidence of each word, merge it with the linguistic feature vector into an array, and determine the array as the feature fusion result.

[0143] Optionally, the speech recognition module 41 includes:

[0144] An acoustic vector extraction unit is used to extract Fbank features from the acquired speech data and calculate the acoustic feature vector of each frame of speech based on the Fbank features;

[0145] The text matching unit is used to match the acoustic feature vector of each frame of speech with the words in the word list, and serialize the matched words to obtain text data.

[0146] Optionally, the confidence calculation module 42 includes:

[0147] a frame number determining unit, configured to determine, for any target word in the text data, a start frame number and an end frame number corresponding to the target word in the acoustic feature vector;

[0148] The confidence determination unit is used to calculate the probability of outputting the above target word under the condition of the acoustic feature vector corresponding to each frame number between the above starting frame number and the above ending frame number, and determine the average value of all probabilities as the confidence of the above target word.

[0149] Optionally, the feature fusion module 44 includes:

[0150] A semantic vector determination unit, configured to extract a character feature vector and a position feature vector of each character in the text data to obtain a semantic feature vector of each character;

[0151] The linguistic vector output unit is used to input the semantic feature vector of each word into the transformer model, and the output feature is the linguistic feature vector corresponding to the above text data.

[0152] Optionally, the emotion classification network includes a two-layer feedforward neural network layer and a one-layer Softmax layer, and uses a cross-entropy function as a loss function; the speech recognition network includes a transformer model and a one-layer feedforward neural network layer, and uses CTC as a loss function, and the emotion classification network and the speech recognition network are jointly trained;

[0153] The above joint training process is:

[0154] Use the speech recognition network to be trained to translate the training speech into training text and calculate the CTC loss;

[0155] Calculating the confidence of the training text and obtaining the acoustic feature vector of the training speech during the process of translating the training speech into the training text;

[0156] Extracting the linguistic feature vector of the training text, and using the confidence level to perform feature fusion on the acoustic feature vector and the linguistic feature vector to obtain a training feature fusion result;

[0157] The feature fusion results of the above training are input into the emotion classification network to be trained, the emotion recognition results of the output training are calculated with the annotation results of the above training speech to calculate the cross entropy loss, and the gradient descent method is used to reversely update the parameters of the above speech recognition network to be trained and the parameters of the above emotion classification network to be trained. It is iterated until the sum of the above cross entropy loss and the above CTC loss converges, and the parameters of the trained speech recognition network and the trained emotion classification network are obtained.

[0158] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0159] Figure 5 This is a schematic diagram of the structure of a computer device provided in Example 5 of this application. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned embodiments of the emotion recognition method based on artificial intelligence are implemented.

[0160] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 5 The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.

[0161] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.

[0162] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.

[0163] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0164] The present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing it.

[0165] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0166] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0167] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which can be electrical, mechanical or other forms.

[0168] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0169] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. An emotion recognition method based on artificial intelligence, characterized in that: The emotion recognition method comprises: Use the trained speech recognition network to translate the acquired speech data into text data; Calculating the confidence of the text data, and obtaining the acoustic feature vector of the speech data in the process of translating the speech data into the text data; Extracting a linguistic feature vector of the text data, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a feature fusion result; Inputting the feature fusion result into a trained emotion classification network, outputting the probability that the speech data is classified into each preset emotion category in the trained emotion classification network, and determining the preset emotion category whose probability meets the preset conditions as the emotion recognition result of the speech data; If the confidence level includes a confidence level for each word in the text data, the acoustic feature vector and the linguistic feature vector are subjected to feature fusion using the confidence level, and a feature fusion result is obtained, including: The confidence of each word is dot-multiplied by the acoustic feature vector, and the result is combined with the linguistic feature vector into an array, and the array is determined to be a feature fusion result.

2. The emotion recognition method according to claim 1, characterized in that Using the trained speech recognition network, the acquired speech data is translated into text data, including: Extracting Fbank features from the acquired speech data, and calculating the acoustic feature vector of each frame of speech based on the Fbank features; The acoustic feature vector of each frame of speech is matched with the words in the vocabulary, and the matched words are serialized to obtain text data.

3. The emotion recognition method according to claim 2, characterized in that Calculating the confidence of the text data includes: For any target word in the text data, determining a start frame number and an end frame number corresponding to the target word in the acoustic feature vector; The probability of outputting the target word under the condition of the acoustic feature vector corresponding to each frame number between the starting frame number and the ending frame number is calculated, and the average value of all probabilities is determined as the confidence of the target word.

4. The emotion recognition method according to claim 3, characterized in that Extracting the linguistic feature vector of the text data includes: Extracting a character feature vector and a position feature vector of each character in the text data to obtain a semantic feature vector of each character; The semantic feature vector of each word is input into the transformer model, and the output feature is the linguistic feature vector corresponding to the text data.

5. The emotion recognition method according to any one of claims 1 to 4, characterized in that: The emotion classification network includes a two-layer feedforward neural network layer and a one-layer Softmax layer, and uses a cross-entropy function as a loss function. The speech recognition network includes a transformer model and a one-layer feedforward neural network layer, and uses CTC as a loss function. The emotion classification network and the speech recognition network are jointly trained; The joint training process is: Use the speech recognition network to be trained to translate the training speech into training text and calculate the CTC loss; Calculating the confidence of the training text and obtaining the acoustic feature vector of the training speech during the process of translating the training speech into the training text; Extracting the linguistic feature vector of the training text, and performing feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a training feature fusion result; The feature fusion result of the training is input into the emotion classification network to be trained, the emotion recognition result of the training is output and the cross entropy loss is calculated with the annotation result of the training speech, and the parameters of the speech recognition network to be trained and the parameters of the emotion classification network to be trained are reversely updated using the gradient descent method. It is iterated until the sum of the cross entropy loss and the CTC loss converges, thereby obtaining the parameters of the trained speech recognition network and the trained emotion classification network.

6. An emotion recognition device based on artificial intelligence, characterized in that: The emotion recognition device comprises: A speech recognition module is used to translate the acquired speech data into text data using a trained speech recognition network; A confidence calculation module, used to calculate the confidence of the text data; A vector acquisition module, configured to acquire an acoustic feature vector of the speech data during the process of translating the speech data into the text data; a feature fusion module, configured to extract the linguistic feature vector of the text data, and perform feature fusion on the acoustic feature vector and the linguistic feature vector using the confidence level to obtain a feature fusion result; An emotion recognition module is configured to input the feature fusion result into a trained emotion classification network, output the probability that the speech data is classified into each preset emotion category in the trained emotion classification network, and determine the preset emotion category whose probability meets the preset conditions as the emotion recognition result of the speech data; If the confidence level includes a confidence level for each word in the text data, the feature fusion module includes: The feature fusion unit is used to perform dot multiplication of the acoustic feature vector with the confidence of each word, and merge the acoustic feature vector with the linguistic feature vector into an array, and determine the array as the feature fusion result.

7. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the emotion recognition method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the emotion recognition method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Speech recognition model training method and speech recognition method and device

    CN110827805A

  • Illegal information identification method and device based on real-time voice emotion analysis

    CN113314103A

  • Customer service intelligent quality inspection method and device and storage medium

    CN114051076A

  • Emotion recognition method and device and robot

    CN114420169A