Speech emotion recognition method and device
By combining speech frames and multi-round dialogue context information with a two-level neural network model, the problem of insufficient utilization of speech context in existing technologies is solved, achieving more accurate speech emotion recognition.
Patent Information
- Application Number
- CN202080029552.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-16
- Filing Date
- 2020-04-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2040-04-16
AI Technical Summary
Existing speech emotion recognition technology is mainly based on single-sentence speech analysis and fails to fully consider the speech context, resulting in inaccurate emotion recognition.
A two-level neural network architecture based on the first neural network model and the second neural network model is adopted to identify the emotional state of the current speech segment by performing statistical operations on multiple speech frames and multi-round dialogue context information.
It achieves more accurate speech emotion recognition and can better learn and utilize the impact of conversation context on the emotional state of the current utterance segment.
Smart Images

Figure CN114127849B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing, and more specifically, to a method and apparatus for speech emotion recognition. Background Art
[0002] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0003] With the continuous development of artificial intelligence (AI), emotional interaction plays a vital role in human communication. Emotion recognition is a fundamental technology for human-computer interaction. Currently, researchers are working on enabling AI assistants to understand people's emotions through their voices. By learning and identifying emotions like anxiety, excitement, and anger in voices, they can achieve more personalized communication.
[0004] Most existing speech emotion recognition technologies are mainly based on speech analysis of a single sentence (current utterance) to identify the speaker's emotions, without considering the speech context, resulting in inaccurate emotion recognition. Summary of the Invention
[0005] The present application provides a method and device for speech emotion recognition, which can achieve more accurate speech emotion recognition effect.
[0006] In a first aspect, a method for speech emotion recognition is provided, the method comprising: determining, based on a first neural network model, multiple emotion state information corresponding to multiple speech frames included in a current utterance in a target dialogue, wherein one speech frame corresponds to one emotion state information, and the emotion state information represents the emotion state corresponding to the speech frame; performing statistical operations on the multiple emotion state information to obtain statistical results, wherein the statistical results are the statistical results corresponding to the current utterance; determining, based on a second neural network model, the emotion state information corresponding to the current utterance according to the statistical results corresponding to the current utterance and the n-1 statistical results corresponding to the n-1 utterances before the current utterance, wherein the n-1 utterances correspond one-to-one to the n-1 statistical results, and the statistical result corresponding to any utterance in the n-1 utterances is obtained by performing statistical operations on the multiple emotion state information corresponding to the multiple speech frames included in the utterance, and the n-1 utterances belong to the target dialogue, and n is an integer greater than 1.
[0007] The method provided in this application, based on a first neural network model, can obtain multiple emotion state information corresponding to multiple speech frames in the current utterance segment. Then, based on the statistical results corresponding to the current utterance segment and the statistical results corresponding to multiple utterance segments before the current utterance segment, the emotion state information corresponding to the current utterance segment can be obtained. Therefore, by using two neural network models, the first neural network model and the second neural network model, the influence of the context of the current utterance segment on the emotion state information corresponding to the current utterance segment can be more fully studied, thereby achieving more accurate speech emotion recognition results.
[0008] Optionally, the statistical operation includes, but is not limited to, calculating one or more of the following: calculating the mean, calculating the variance, calculating the extreme value, calculating the coefficient of a linear fit, and calculating the coefficient of a higher-order fit. Accordingly, the statistical result includes, but is not limited to, one or more of the following: calculating the mean, calculating the variance, calculating the extreme value, calculating the coefficient of a linear fit, and calculating the coefficient of a higher-order fit.
[0009] In conjunction with the first aspect, in certain implementations of the first aspect, the n-1 speech segments include speech data of multiple speakers, that is, the n-1 speech segments are conversations between multiple speakers.
[0010] Based on this solution, by performing speech recognition based on the conversation of multiple speakers, a more accurate speech emotion recognition effect can be achieved compared to the existing technology of performing speech emotion recognition based on a single sentence of a speaker.
[0011] In combination with the first aspect, in some implementations of the first aspect, the multiple speakers include a speaker corresponding to the current speech segment;
[0012] Furthermore, the determining, based on the second neural network model, the emotional state information corresponding to the current utterance segment according to the statistical results corresponding to the current utterance segment and the n-1 statistical results corresponding to the n-1 utterance segments before the current utterance segment, includes:
[0013] Based on the second neural network model, the emotional state information corresponding to the current speech segment is determined according to the statistical results corresponding to the current speech segment, the n-1 statistical results, and the genders of the multiple speakers.
[0014] By combining the speaker's gender for speech emotion recognition, more accurate speech emotion recognition results can be obtained.
[0015] In conjunction with the first aspect, in certain implementations of the first aspect, the n-1 speech segments are temporally adjacent, that is, there is no other speech data between any two of the n-1 speech segments.
[0016] In combination with the first aspect, in certain implementations of the first aspect, the emotional state information corresponding to the current speech segment is determined based on the second neural network model according to the statistical results corresponding to the current speech segment and the n-1 statistical results corresponding to the n-1 speech segments before the current speech segment, including: determining the round features corresponding to the w rounds corresponding to the current speech segment and the n-1 speech segments according to the statistical results corresponding to the current speech segment and the n-1 statistical results, wherein the round features corresponding to any round are determined by the statistical results corresponding to the speech segments of all speakers in the round, and w is an integer greater than or equal to 1; based on the second neural network model, the emotional state information corresponding to the current speech segment is determined according to the round features corresponding to the w rounds.
[0017] Specifically, taking the speech data of two speakers A and B in each round as an example, the round features corresponding to any round are determined based on the statistical results corresponding to A and the statistical results corresponding to B in the round of dialogue. For example, the round features corresponding to the current round corresponding to the current speech segment are the vector splicing of the statistical results corresponding to each speech segment included in the current round. Furthermore, the round features can also be determined in combination with the genders of A and B. For example, the round features corresponding to the current round corresponding to the current speech segment are the vector splicing of the statistical results corresponding to each speech segment included in the current round and the genders of each speaker corresponding to the current round. In the present application, w round features can be input into the second neural network model, and the output of the second neural network model is the emotional state information corresponding to the current speech segment.
[0018] Optionally, an utterance segment is a sentence, so an utterance segment corresponds to one speaker.
[0019] Therefore, the method provided in this application performs speech emotion recognition based on the speech data of multiple speakers before the current speech segment, that is, based on the context information of multiple rounds of dialogue. Compared with the existing technology of speech emotion recognition based on a single sentence, it can achieve more accurate speech emotion recognition effect.
[0020] In combination with the first aspect, in some implementations of the first aspect, w is a value input by a user.
[0021] In combination with the first aspect, in some implementations of the first aspect, the method further includes: presenting emotional state information corresponding to the current speech segment to the user.
[0022] In combination with the first aspect, in some implementations of the first aspect, the method further includes: obtaining a correction operation performed by the user on the emotional state information corresponding to the current speech segment.
[0023] Furthermore, the method further includes: updating the value of w.
[0024] That is, if the prediction result is not what the user expects, the user can modify the prediction result. After the speech emotion recognition device recognizes the user's modification operation, it can update the value of w to achieve a more accurate prediction result.
[0025] In combination with the first aspect, in some implementations of the first aspect, the first neural network model is a long short-term memory model (LSTM); and / or the second neural network model is LSTM.
[0026] Because the LSTM model has good memory capacity, it can more fully learn the impact of the conversation context on the emotional state information corresponding to the current utterance segment, thereby achieving more accurate speech emotion recognition effects.
[0027] It should be understood that the first neural network model and the second neural network model may be the same or different, and this application does not limit this.
[0028] In conjunction with the first aspect, in certain implementations of the first aspect, determining, based on the first neural network model, multiple emotion state information corresponding to multiple speech frames included in the current utterance segment in the target conversation includes:
[0029] Based on the first neural network model, for each speech frame among the multiple speech frames, the emotional state information corresponding to the speech frame is determined according to the feature vector corresponding to the speech frame and the feature vectors corresponding to the q-1 speech frames before the speech frame, wherein the q-1 speech frames are the speech frames of the speaker corresponding to the current speech segment, q is an integer greater than 1, and the feature vector of the speech frame k represents the acoustic features of the speech frame k.
[0030] Optionally, the acoustic features include but are not limited to one or more of energy, fundamental frequency, zero-crossing rate, Mel-frequency cepstral coefficient (MFCC), etc. Exemplarily, the feature vector of each speech frame can be obtained by concatenating the aforementioned acoustic features.
[0031] In combination with the first aspect, in some implementations of the first aspect, there are m speech frames between any two speech frames in the q speech frames, where m is an integer greater than or equal to 0.
[0032] Based on this technical solution, when m is not 0, the context contained in the window corresponding to the speech frame can be expanded while avoiding the window sequence being too long, thereby further improving the accuracy of the prediction result.
[0033] In a second aspect, a speech emotion recognition device is provided, which includes a module for executing the method in the first aspect.
[0034] In a third aspect, a speech emotion recognition device is provided, which includes: a memory for storing programs; a processor for executing the programs stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method in the first aspect.
[0035] According to a fourth aspect, a computer-readable medium is provided, wherein the computer-readable medium stores a program code for execution by a device, wherein the program code includes instructions for executing the method according to the first aspect.
[0036] In a fifth aspect, a computer program product comprising instructions is provided, which, when run on a computer, enables the computer to execute the method in the first aspect.
[0037] In a sixth aspect, a chip is provided, comprising a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface and executes the method in the first aspect.
[0038] Optionally, as an implementation method, the chip may further include a memory, in which instructions are stored, and the processor is used to execute the instructions stored on the memory. When the instructions are executed, the processor is used to execute the method in the first aspect.
[0039] In a seventh aspect, an electronic device is provided, which includes the motion recognition device of any one of the second to fourth aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of a natural language processing system;
[0041] Figure 2 This is another application scenario diagram of the natural language processing system;
[0042] Figure 3 This is a schematic diagram of the system architecture provided by the embodiment of the present application;
[0043] Figure 4 This is a schematic diagram of a chip hardware structure provided by an embodiment of the present application;
[0044] Figure 5 is a schematic flow chart of the speech emotion recognition method according to an embodiment of the present application;
[0045] Figure 6 It is a schematic diagram of dividing the speech stream into utterance segments;
[0046] Figure 7 It is a schematic diagram of dividing the speech segment into speech frames;
[0047] Figure 8 is a schematic block diagram of a speech emotion recognition device according to an embodiment of the present application;
[0048] Figure 9 Schematic diagram of the hardware structure of the neural network training device according to an embodiment of the present application;
[0049] Figure 10 Schematic diagram of the hardware structure of the speech emotion recognition device according to the embodiment of the present application. DETAILED DESCRIPTION
[0050] The technical solution in this application will be described below with reference to the accompanying drawings.
[0051] In order to make those skilled in the art better understand this application, first combine Figure 1 and Figure 2 Briefly introduce the scenarios in which this application can be used.
[0052] like Figure 1 A schematic diagram of a natural language processing system is shown. Figure 1 The system includes user devices and data processing equipment. The user devices include users and intelligent terminals such as mobile phones, personal computers, or information processing centers. The user devices are the initiators of natural language data processing, serving as the initiators of language questions and answers or queries, typically initiated by users through the user devices.
[0053] The data processing device can be a device or server with data processing capabilities, such as a cloud server, network server, application server, or management server. The data processing device receives query statements / voice / text queries from the smart terminal via the interactive interface, and then performs language data processing using a memory for storing data and a data processing processor, including machine learning, deep learning, search, reasoning, and decision-making. The memory can be a general term that includes local storage and a database for storing historical data. The database can be located on the data processing device or on another network server.
[0054] like Figure 2 The figure shows another application scenario of the natural language processing system. In this scenario, the smart terminal directly acts as a data processing device, directly receiving input from the user and directly processing it by the hardware of the smart terminal itself. The specific process is the same as Figure 1 Similarly, please refer to the above description and will not be repeated here.
[0055] Emotional interaction plays a crucial role in human communication. Research shows that 80% of communication is emotional. Therefore, affective computing is essential for achieving humanized human-computer interaction, and emotion recognition and understanding technology is a foundational technology for human-computer interaction.
[0056] Most existing intelligent assistants, such as Apple's Siri and Amazon's Alexa, are still primarily knowledge-based question-and-answer interactions. However, researchers are working to help AI assistants understand people's emotions through voice. By learning and identifying emotions such as anxiety, excitement, and anger in voice, they can achieve more personalized communication. For example, if an intelligent assistant detects a low tone in your voice, it may play a cheerful song for you, or tell you a white lie, saying that your friend sounds depressed and recommends that you go to an uplifting movie with them.
[0057] The emotion in speech is affected by the speech context, and most existing speech emotion recognition technologies are mainly based on speech analysis of a single sentence (current utterance) to identify the speaker's emotion, without considering the speech context, resulting in inaccurate emotion recognition.
[0058] Furthermore, in a conversation, not only does the speaker's previous words influence the current emotion, but the other person's speech and emotional state also influence the current emotion. In practice, the emotional state in a conversation is influenced by two different levels of speech context: one is the level of the speech stream, where the previous and subsequent frames affect the pronunciation of the current frame; the other is the influence of previous turns in the conversation on the current turn. In this case, identifying the speaker's emotion based solely on speech analysis of a single sentence (the current utterance) without considering the conversational context will result in inaccurate and inappropriate predictions in multi-turn interactions.
[0059] In response to the above problems, the present application provides a speech emotion recognition method and device, which can achieve more accurate emotion recognition effects by introducing speech context information into the speech emotion recognition process.
[0060] The speech emotion recognition method proposed in this application can be applied in fields with natural human-computer interaction requirements. Specifically, according to the method of this application, by performing speech emotion recognition on the input conversation voice data stream, the emotional state corresponding to the current speech, such as happiness, anger, etc., can be identified. Subsequently, depending on different application scenarios, the identified real-time emotional state can be used to formulate conversation response strategies, divert calls to human customer service, adjust teaching progress, etc. For example, a voice assistant can adjust the voice response strategy based on the emotional changes identified during the conversation with the user, thereby achieving more personalized human-computer interaction. In addition, the customer service system can be used to sort the urgency of call center users, thereby improving service quality. For example, by promptly identifying users with more intense negative emotions and transferring their calls to human customer service, the user experience can be optimized. Distance education systems can be used to monitor the emotional state of remote online classroom users during the learning process, thereby adjusting the teaching focus or progress in a timely manner. Hospitals can use it to track the emotional changes of patients with depression as a basis for disease diagnosis and treatment. It can also be used to assist and guide children with autism in learning to understand and express emotions.
[0061] The speech emotion recognition device provided in this application can be Figure 1 The data processing device shown or Figure 1 In addition, the speech emotion recognition device provided by this application can also be a unit or module in the data processing device shown. Figure 2 The user equipment shown or Figure 2 The units or modules in the data processing device shown. For example, Figure 1The data processing device shown may be a cloud server, and the speech emotion recognition device may be a conversational speech emotion recognition service application programming interface (API) on the cloud server; for example, the speech device may be Figure 2 The voice assistant application (APP) in the user device is shown.
[0062] For example, the speech emotion recognition device provided in this application can be a standalone conversational speech emotion recognition software product, a conversational speech emotion recognition service API on a public cloud, or a functional module embedded in a voice interaction product, such as a smart speaker, a voice assistant app on a mobile phone, intelligent customer service software, or an emotion recognition module in a distance education system. It should be understood that the product forms listed here are for illustrative purposes only and do not constitute any limitation on this application.
[0063] The following combination Figure 3 The system architecture shown illustrates the training process of the model used in this application.
[0064] like Figure 3 As shown, the embodiment of the present application provides a system architecture 100. Figure 3 In the present application, the data acquisition device 160 is used to collect training corpus. In the present application, the speech of a conversation between multiple people (for example, two or more people) can be used as training corpus. The training corpus contains two types of annotations: one is the annotation of the emotional state information for each frame, and the other is the annotation of the emotional state information for each speech segment. At the same time, the training corpus has marked the speaker of each speech segment. Furthermore, the gender of the speaker can also be marked. It should be understood that each speech segment can be divided into multiple speech frames. For example, each speech segment can be framed according to a frame length of 25ms and a frame shift of 10ms, so that multiple speech frames corresponding to each speech segment can be obtained.
[0065] It should be noted that the emotional state information in this application can adopt any emotional state representation method. At present, the emotional state representation commonly used in the industry includes two types: discrete representation, such as adjective label forms such as happy and angry; dimensional representation, that is, describing the emotional state as a point (x, y) in a multidimensional emotional space. For example, the emotional state information in the embodiment of the present application can be represented by an arousal-valence space model, that is, the emotional state information can be represented by (x, y). Among them, in the activation-valence space model, the vertical axis is the activation dimension, which is a description of the intensity of the emotion; the horizontal axis is the valence dimension, which is an evaluation of the positive or negative degree of the emotion. For example, x and y can be described by numerical values 1 to 5, respectively, but this application is not limited to this.
[0066] After collecting the training corpus, the data collection device 160 stores the training corpus in the database 130 , and the training device 120 obtains the target model / rule 101 through training based on the training corpus maintained in the database 130 .
[0067] The target model / rule 101 is a two-level neural network device model, namely, a first neural network model and a second neural network model. The following describes the process of the training device 120 obtaining the first neural network model and the second neural network model based on the training corpus.
[0068] For each speech segment of each speaker in the training corpus, the frame is divided according to a certain frame length and frame shift, so that multiple speech frames corresponding to each speech segment can be obtained. Then, the feature vector of each frame is obtained. Among them, the feature vector represents the acoustic feature of the speech frame. Acoustic features include but are not limited to one or more of energy, fundamental frequency, zero crossing rate, Mel frequency cepstral coefficient (MFCC), etc. Exemplarily, the feature vector of each speech frame can be obtained by splicing the aforementioned acoustic features together. For the feature vector of each current frame, the feature vectors of the previous q-1 frames are combined to form a window sequence of length q, where q is an integer greater than 1. Optionally, in order to expand the context contained in the window without making the window sequence too long, a downsampling method can be adopted, that is, one frame is taken every m frames and added to the sequence. Wherein, m is a positive integer. Each window sequence is used as a training sample, and all training samples are used as input to train a first neural network model. In this application, the first neural network model can be an LSTM model, but this application is not limited to this. Exemplarily, the first neural network model in the present application may adopt a two-layer structure, with 60 and 80 hidden layer neurons respectively, and the loss function is mean squared error (MSE).
[0069] Determine the statistical results corresponding to each speech segment. Specifically, for each speech frame, the first neural network model can output an emotional state information prediction result. By performing statistical operations on the emotional state information corresponding to all or part of the speech frames corresponding to each speech segment, the statistical results corresponding to each speech segment can be obtained. Exemplarily, the statistical operations include but are not limited to calculating one or more of the mean, variance, extreme value, linear fit and high-order fit coefficients. Accordingly, the statistical results include but are not limited to one or more of the mean, variance, extreme value, linear fit and high-order fit coefficients.
[0070] Then, the statistics corresponding to each utterance segment and the statistics corresponding to the previous utterance segments can be concatenated as input to train a second neural network model. Furthermore, the statistics corresponding to each utterance segment and the speaker, as well as the statistics corresponding to the previous utterance segments and the speaker, can be concatenated as input to train a second neural network model. Alternatively, the round features corresponding to each round and the round features corresponding to the previous rounds can be used as input to train a second neural network model. Exemplarily, the round features corresponding to any round can be obtained by concatenating the statistics corresponding to the utterance segments corresponding to all speakers in that round. Furthermore, the round features corresponding to any round can be obtained by concatenating the statistics corresponding to the utterance segments corresponding to all speakers in that round and the genders of all speakers. The concatenation can be vector concatenation or a weighted sum operation, and the specific method of concatenation is not limited in this application. In this application, the second neural network model can be an LSTM model, but this application is not limited to this. Exemplarily, the second neural network model can adopt a one-layer structure with 128 hidden layer neurons and a loss function of MSE.
[0071] Because the LSTM model has good memory capacity, it can more fully learn the impact of the conversation context on the emotional state information corresponding to the current utterance segment, thereby achieving more accurate speech emotion recognition effects.
[0072] It should be understood that in the present application, the first neural network model and the second neural network model can be recurrent neural network models, and the two can be the same or different, and the present application does not limit this.
[0073] After completing the training of the target model / rule 101, that is, obtaining the first neural network model and the second neural network model, the speech emotion recognition method of the embodiment of the present application can be implemented through the above-mentioned target model / rule 101. That is, by inputting the target dialogue into the target model / rule 101, the emotional state information of the current speech segment can be obtained. It should be understood that the model training process described above is only an exemplary implementation method of the present application and should not constitute any limitation on the present application.
[0074] It should be noted that, in actual applications, the training corpus maintained in the database 130 does not necessarily come from the data acquisition device 160, but may also be received from other devices. It should also be noted that the training device 120 does not necessarily train the target model / rule 101 entirely based on the training corpus maintained by the database 130, but may also obtain training corpus from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.
[0075] The target model / rule 101 obtained by training the training device 120 can be applied to different systems or devices, such as Figure 3 The execution device 110 shown in the figure can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) or virtual reality (VR), a vehicle terminal, etc. It can also be a server or a cloud. Figure 3 In the embodiment of the present application, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. The user can input data to the I / O interface 112 through the client device 140. The input data may include: a target dialogue input by the client device.
[0076] The preprocessing module 113 and the preprocessing module 114 are used to perform preprocessing based on the input data (such as the target dialogue) received by the I / O interface 112. In an embodiment of the present application, the preprocessing module 113 and the preprocessing module 114 may be omitted (or only one of the preprocessing modules may be present), and the computing module 111 may be used directly to process the input data.
[0077] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs calculations and other related processing, the execution device 110 can call the data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.
[0078] Finally, I / O interface 112 returns the processing result, such as the emotional state information of the current utterance segment obtained above, to client device 140, thereby providing it to the user. It should be understood that I / O interface 112 may not return the emotional state information of the current utterance segment to client device 140, and this application is not limited to this.
[0079] It is worth noting that the training device 120 can generate corresponding target models / rules 101 based on different training data for different goals or different tasks. The corresponding target models / rules 101 can be used to achieve the above goals or complete the above tasks, thereby providing users with the desired results.
[0080] In the attached Figure 3In the case shown in FIG, the user can manually select input data, and this manual selection can be performed through the interface provided by I / O interface 112. In another case, client device 140 can automatically send input data to I / O interface 112. If the automatic sending of input data by client device 140 requires user authorization, the user can set the corresponding permissions in client device 140. The user can view the results output by execution device 110 on client device 140, and the specific presentation form can be a display, sound, action, etc. Client device 140 can also serve as a data acquisition terminal, collecting input data input into I / O interface 112 and output results output from I / O interface 112 as new training corpus, and storing them in database 130. Of course, it is also possible to bypass client device 140 for collection, and instead have I / O interface 112 directly store the input data input into I / O interface 112 and output results output from I / O interface 112 as new training corpus in database 130.
[0081] It is worth noting that the Figure 3 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in the attached Figure 3 In the embodiment, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.
[0082] Figure 4 A chip hardware structure provided in an embodiment of the present application includes a neural network processor 20. The chip can be set as follows Figure 3 The execution device 110 shown in FIG. 1 is used to complete the calculation work of the calculation module 111. The chip can also be set in Figure 3 The training device 120 shown is used to complete the training work of the training device 120 and output the target model / rule 101.
[0083] The neural network processor NPU 20 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core part of the NPU is the arithmetic circuit 20. The controller 204 controls the arithmetic circuit 203 to extract data from the memory (weight memory or input memory) and perform calculations.
[0084] In some implementations, arithmetic circuit 203 includes multiple processing engines (PEs). In some implementations, arithmetic circuit 203 is a two-dimensional systolic array. Arithmetic circuit 203 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 203 is a general-purpose matrix processor.
[0085] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 201 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 208.
[0086] The vector calculation unit 207 can further process the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit 207 can be used for network calculations of non-convolutional / non-FC layers in a neural network, such as pooling, batch normalization, local response normalization, etc.
[0087] In some implementations, the vector calculation unit 207 can store the processed output vector to the unified buffer 206. For example, the vector calculation unit 207 can apply a nonlinear function to the output of the operation circuit 203, such as a vector of accumulated values, to generate an activation value. In some implementations, the vector calculation unit 207 generates a normalized value, a merged value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 203, for example, for use in a subsequent layer in a neural network.
[0088] The unified memory 206 is used to store input data and output data.
[0089] The weight data is directly transferred from the external memory to the input memory 201 and / or the unified memory 206 through the direct memory access controller 205 (DMAC), the weight data in the external memory is stored in the weight memory 202, and the data in the unified memory 206 is stored in the external memory.
[0090] The bus interface unit (BIU) 210 is used to implement interaction between the main CPU, DMAC and instruction fetch memory 209 through the bus.
[0091] An instruction fetch buffer 209 connected to the controller 204 and used to store instructions used by the controller 204;
[0092] The controller 204 is used to call the instructions cached in the memory 209 to control the working process of the computing accelerator.
[0093] Entrance: According to the actual invention, the data here is descriptive data, such as the detected vehicle speed, obstacle distance, etc.
[0094] Generally, the unified memory 206, the input memory 201, the weight memory 202 and the instruction fetch memory 209 are all on-chip memories, and the external memory is a memory outside the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM) or other readable and writable memory.
[0095] Introduced above Figure 3 The execution device 110 in the embodiment of the present application can execute each step of the speech emotion recognition method. Figure 3 The chip shown can also be used to execute the various steps of the speech emotion recognition method of the embodiment of the present application. The speech emotion recognition method of the embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0096] Figure 5 This is a schematic flow chart of the speech emotion recognition method provided by this application. The following describes each step in the method. It should be understood that the method can be performed by a speech emotion recognition device.
[0097] S310, based on the first neural network model, determining multiple emotion state information corresponding to multiple speech frames included in the current utterance in the target dialogue.
[0098] The emotional state information represents the emotional state corresponding to the speech frame. The method for representing the emotional state information can be found in the above description and will not be repeated here.
[0099] The target conversation refers to the speech data stream input to the speech emotion recognition device. It can be a real-time speech data stream of a user, but this application is not limited to this. The speech emotion recognition device can divide the target conversation into small segments based on existing or new speech recognition methods that may emerge as technology develops, and mark the speaker of each small segment.
[0100] For example, a speech recognition device can use voiceprint recognition technology to segment the target conversation into small segments according to the switching of speakers, and one of the small segments can be considered as a speech segment. Figure 6 ,Based on the switching of speakers, the target dialogue can be divided into segments A1, B1, A2, B2, ..., A t-1 , B t-1 , A t , B t , where A represents one speaker and B represents the other speaker.
[0101] For example, the speech recognition device may consider a segment of speech data with a pause time exceeding a preset time (for example, 200ms) as a speech segment based on the temporal continuity of the speech data. Figure 6 In A2, if speaker A pauses for a period of time, such as 230ms, when speaking this paragraph, then the speech data before the pause in A2 can be considered as a speech segment A. 2-0 The speech data from the pause to the end of A2 is another speech segment A. 2-1 .
[0102] It is understandable that different speech recognition methods determine different speech segments. It is generally believed that a speech segment is a sentence, or a speech segment can be the voice data from the beginning to the end of a speaker's speech without being interrupted by others, but the embodiments of the present application are not limited to this.
[0103] Each speech segment can be divided into multiple speech frames. For example, each speech segment can be divided into frames according to a frame length of 25ms and a frame shift of 10ms, thereby obtaining multiple speech frames corresponding to each speech segment.
[0104] Taking the current speech segment as an example, multiple speech frames can be obtained by framing the current speech segment. The "multiple speech frames included in the current speech segment" mentioned herein may be some or all (a number g) of the speech frames obtained by framing the current speech segment. For example, the current speech segment can be framed with a frame length of 25ms and a frame shift of 10ms, and then a frame is taken every h frames, for a total of g frames as the multiple speech frames, where h is a positive integer and g is an integer greater than 1.
[0105] After obtaining g speech frames in the current speech segment, the emotional state information corresponding to the g speech frames can be obtained based on the first neural network model.
[0106] In one implementation, based on a first neural network model, determining the emotional state information corresponding to g speech frames includes: based on the first neural network model, for each of the g speech frames, determining the emotional state information corresponding to the speech frame based on a feature vector corresponding to the speech frame and feature vectors corresponding to the q-1 speech frames preceding the speech frame. The q-1 speech frames are speech frames of the speaker corresponding to the current utterance, the value of q can be found in the description above, and the feature vector of speech frame k represents the acoustic features of speech frame k.
[0107] Specifically, for any speech frame k, the feature vectors corresponding to the q speech frames can be combined into a window sequence of length q and input into the first neural network model. The output of the first neural network model is the emotional state information corresponding to the speech frame k. As mentioned above, the acoustic features of the speech frame k include but are not limited to one or more of energy, fundamental frequency, zero-crossing rate, Mel frequency cepstral coefficient (MFCC), etc. The feature vector of the speech frame k can be obtained by splicing the aforementioned acoustic features.
[0108] It should be understood that the q speech frames corresponding to speech frame k may only include speech frames in the speech segment to which speech frame k belongs, or may also include speech frames in the speech segment to which speech frame k belongs and speech frames in other speech segments. The specific situation depends on the number of speech frames in the speech segment to which the speech frame belongs.
[0109] It should also be understood that q can be a fixed value set when the voice emotion recognition device leaves the factory, or it can be a non-fixed value. For example, q can be set by the user, and this application does not limit this.
[0110] Optionally, among the q speech frames mentioned above, any two speech frames may be separated by m speech frames, where m is defined as described above.
[0111] by Figure 7 Segment B shown t and B t-1 For example, see Figure 7 , Segment B t is divided into speech frames F t,0 , F t,1 , F t,2 , ..., Segment B t-1 is divided into speech frames F t-1,0 , F t-1,1 , F t-1,2, F t-1,3 , F t-1,4 , F t-1,5 Assume q = 4, utterance B t The g speech frames include speech frame F t,0 Then, if m=0, the speech frame F t,0 The corresponding q speech frames can be F t,0 , F t-1,5 , F t-1,4 , F t-1,3 If m=1, then the speech frame F t,0 The corresponding q speech frames can be F t,0 , F t-1,4 , F t-1,2 , F t-1,0 .
[0112] Based on this technical solution, when m is not 0, the context contained in the window corresponding to the speech frame k can be expanded while avoiding the window sequence being too long, thereby further improving the accuracy of the prediction result.
[0113] S320: Perform statistical operations on the g pieces of emotional state information to obtain a statistical result, which is the statistical result corresponding to the current utterance segment.
[0114] For example, the statistics in this application include but are not limited to mean, variance, extreme value, coefficients of linear fit and high-order fit.
[0115] S330, based on the second recurrent neural network model, determine the emotional state information corresponding to the current utterance segment according to the statistical results corresponding to the current utterance segment and the n-1 statistical results corresponding to the n-1 utterance segments before the current utterance segment.
[0116] The n-1 utterances correspond one-to-one to the n-1 second statistical results, i.e., one utterance corresponds to one statistical result. Furthermore, the statistical result corresponding to any one of the n-1 utterances is obtained by performing a statistical operation on the g pieces of emotional state information corresponding to the g speech frames included in the utterance. The n-1 utterances belong to the target conversation, and n is an integer greater than 1.
[0117] The above description uses the current utterance as an example to describe in detail how to determine the emotional state information corresponding to the g speech frames included in the current utterance, thereby further determining the statistical results corresponding to the current utterance. For any of the n-1 utterances, the method for determining the emotional state information corresponding to the multiple speech frames included in the utterance is similar to the method for determining the emotional state information corresponding to the g speech frames included in the current utterance, and will not be repeated here. Therefore, n-1 statistical results corresponding to the n-1 utterances can be further determined.
[0118] It should be understood that in practice, the speech emotion recognition device can determine the corresponding statistical results for each speech segment received in chronological order. Figure 6 B shown t , then the speech emotion recognition device can determine B t The statistical results corresponding to the previous discourse segment.
[0119] It should be noted that the method provided in this application can be applied to two scenarios: (1) For each utterance segment input to the speech emotion recognition device, there is corresponding emotional state information output. In this scenario, if the number of utterance segments before the current utterance segment is less than n-1, for example, the current utterance segment is Figure 6 As shown in A1, n-1 speech segments can be padded with the first default value (such as 0), and the statistical results corresponding to these speech segments with the default value are considered to be the second default value. It should be understood that the first default value and the second default value can be the same or different. (2) When the speech segments input to the speech emotion recognition device reach n, the input speech emotion recognition device outputs the emotional state information. In other words, for the 1st to n-1th speech segments, there is no corresponding emotional state information output. That is, there is no need to consider the problem in the first scenario described above.
[0120] The method provided in this application, based on a first neural network model, can obtain multiple emotion state information corresponding to multiple speech frames in the current utterance segment. Then, based on the statistical results corresponding to the current utterance segment and the statistical results corresponding to multiple utterance segments before the current utterance segment, the emotion state information corresponding to the current utterance segment can be obtained. Therefore, by using two neural network models, the first neural network model and the second neural network model, the influence of the context of the current utterance segment on the emotion state information corresponding to the current utterance segment can be more fully studied, thereby achieving more accurate speech emotion recognition results.
[0121] Optionally, the method may further include: presenting the emotional state information corresponding to the current speech segment to the user.
[0122] That is, after determining the emotional state information corresponding to the current speech segment, the prediction result can be presented to the user.
[0123] Furthermore, the method may include: obtaining a user's correction operation on the emotional state information corresponding to the current speech segment.
[0124] Specifically, if the presented prediction result is inaccurate, the user can also correct the prediction result.
[0125] Optionally, the n-1 utterance segments are adjacent in time, that is, there is no other speech data between any two utterance segments in the n-1 utterance segments.
[0126] by Figure 6 For example, if the current utterance is B t , the n-1 utterance segments can be A1, B1, A2, B2, ..., A t-1 , B t-1 , A t , or, it can be A2, B2, ..., A t-1 , B t-1 It should be understood that two utterance segments such as A1 and A2 cannot be called adjacent, as there is one utterance segment between the two utterance segments, namely B1.
[0127] In addition, any two utterances in the n-1 utterances may not be adjacent. For example, if the target conversation is a conversation between two speakers, any two utterances in the n-1 utterances may be separated by one or two utterances.
[0128] Furthermore, the n-1 speech segments include speech data of multiple speakers, that is, the n-1 speech segments are conversations between multiple speakers.
[0129] For example, if the current utterance is Figure 6 As shown in B2, the n-1 speech segments may include A1, B1, A2, that is, the n-1 speech segments include the voice data of A and B.
[0130] In addition, the n-1 speech segments may also only include the speech segments of the speaker corresponding to the current speech segment. For example, the current speech segment corresponds to speaker A, and speaker A has been speaking before the current speech segment. In this case, the n-1 speech segments do not include the speech segments corresponding to other people.
[0131] Based on this solution, by performing speech recognition based on the speaker's context, a more accurate speech emotion recognition effect can be achieved compared to the existing technology of performing speech emotion recognition based on a sentence of the speaker.
[0132] As an implementation method of S330, the statistical results corresponding to the current speech segment and the n-1 statistical results, which are n statistical results, can be input into a second neural network model, and the output of the second neural network model is the emotional state information corresponding to the current speech segment.
[0133] That is, the input of the second neural network model is the n statistical results. Based on this solution, no processing of the statistical results is required, and the implementation is relatively simple.
[0134] Furthermore, the emotional state information corresponding to the current utterance segment may be determined by combining the gender of the speaker corresponding to the current utterance segment and the n-1 utterance segments.
[0135] For example, the n statistical results and the genders of the speakers corresponding to the current utterance segment and the n-1 utterance segments can be input into the second neural network model, and the second neural network model outputs the emotional state information corresponding to the current utterance segment.
[0136] By combining the speaker's gender for speech emotion recognition, more accurate recognition results can be obtained.
[0137] As another implementation of S330, the emotional state information corresponding to the current speech segment may be determined based on the second neural network model and according to the w turn features.
[0138] The n-1 utterances and the current utterance correspond to w rounds of dialogue, or the n utterances correspond to w rounds, where w is an integer greater than 1. Optionally, the rounds can be divided based on the speaker. Figure 6 For example, if the speaker corresponding to A2 is A, and the most recent speech segment corresponding to A is A1, then the speech segments from A1 to A2 are classified as one round, that is, A1 and B1 are one round of dialogue.
[0139] As mentioned above, the n-1 utterances and the current utterance correspond to w turns, wherein the turn feature corresponding to any turn can be determined by the statistical results corresponding to the utterances of all speakers in that turn.
[0140] Specifically, taking the speech data of two speakers A and B in each round as an example, the round features corresponding to any round are determined based on the statistical results corresponding to A and the statistical results corresponding to B in the round of dialogue. For example, the round features corresponding to the current round corresponding to the current speech segment are the vector splicing of the statistical results corresponding to each speech segment included in the current round. Furthermore, the round features can also be determined in combination with the genders of A and B. For example, the round features corresponding to the current round corresponding to the current speech segment are the vector splicing of the statistical results corresponding to each speech segment included in the current round and the genders of each speaker corresponding to the current round. In the present application, w round features can be input into the second neural network model, and the output of the second neural network model is the emotional state information corresponding to the current speech segment.
[0141] Therefore, the method provided in this application performs speech emotion recognition based on the speech data of multiple speakers before the current speech segment, that is, based on the context information of multiple rounds of dialogue. Compared with the existing technology of speech emotion recognition based on a single sentence, it can achieve more accurate speech emotion recognition effect.
[0142] Different from the previous implementation of S330, this implementation processes the statistical results corresponding to each speech segment in each round and then inputs them into the second neural network model.
[0143] It should be understood that when the input of the second neural network model is the turn feature, it is necessary to set the value of w, but not the value of n; when the input of the second neural network model is the statistical result corresponding to the discourse segment, it is necessary to set the value of n, but not the value of w.
[0144] Alternatively, w can be input from the user. For example, an interface may be presented to the user requiring the user to input a value for w, and the user may determine the value for w. In another example, multiple values for w may be presented to the user, and the user may select one of the values for w.
[0145] Optionally, after determining the emotional state information corresponding to the current utterance, the emotional state information corresponding to the current utterance can also be presented to the user. The user can modify the prediction result, and if the user's modification operation on the emotional state information corresponding to the current utterance is obtained, the value of w can be updated.
[0146] Furthermore, the process of updating the value of w can be to set the value of w and re-predict the emotional state information corresponding to the current speech segment. If the prediction result matches the result input by the user, the value of w set at this time is used as the updated value of w. Otherwise, the value of w is reset again and the emotional state information corresponding to the current speech segment is predicted until the prediction result matches the result input by the user, and the value of w that matches the result input by the user is used as the updated value of w.
[0147] That is, if the prediction result is not what the user expects, the user can modify the prediction result. After the speech emotion recognition device recognizes the user's modification operation, it can update the value of w to achieve a more accurate prediction result.
[0148] It should be noted that if the current round corresponding to the current utterance segment only includes the speech of the speaker corresponding to the current utterance segment, then the statistical results corresponding to other speakers in the current round can be set as default values.
[0149] Table 1 below shows an experimental result of speech emotion recognition according to the method of the present application. It should be understood that the numbers in the first row of Table 1 are the values of w.
[0150] Specifically, we used the public IEMOCAP database for experiments. This database contains five dialogues, each with several segments. Each dialogue contains 10-90 utterances, with an average length of 2-5 seconds.
[0151] The experiment used a cross-validation approach, looping through four conversations for training and the remaining one for testing, ultimately generating prediction results for all five conversations. The average recall (UAR) was used as the evaluation metric to predict affective states based on both valence and activation. The experimental results are as follows:
[0152] Table 1
[0153]
[0154] Table 1 above shows the UAR obtained by using context windows of different lengths in the conversation-level LSTM model.
[0155] Compared with a single-stage LSTM, we can see that the prediction results are significantly improved after considering historical conversation turns. The highest UAR for valence is 72.32%, and the highest UAR for activation is 68.20%.
[0156] Therefore, the method provided in this application performs speech emotion recognition based on multi-round dialogue context information, which can achieve more accurate speech emotion recognition effect.
[0157] Combined with the above Figures 5 to 7 The speech emotion recognition method of the embodiment of the present application is described in detail. Figure 8 The speech emotion recognition device of the embodiment of the present application is described. It should be understood that the above Figure 5 The various steps in the method shown can be Figure 8 The speech emotion recognition device shown in the figure is used to perform the above description and limitation of the speech emotion recognition method also apply to Figure 8 The speech emotion recognition device shown in the following description Figure 8 The speech emotion recognition device shown is appropriately omitted from repeated descriptions.
[0158] Figure 8 It is a schematic block diagram of the speech emotion recognition device according to an embodiment of the present application. Figure 8 The illustrated speech emotion recognition apparatus 400 includes a determination module 410 and a statistics module 420 .
[0159] Determination module 410, configured to determine, based on the first neural network model, a plurality of emotion state information corresponding to a plurality of speech frames included in a current utterance segment in the target dialogue, wherein each speech frame corresponds to a piece of emotion state information, and the emotion state information represents the emotion state corresponding to the speech frame;
[0160] A statistics module 420 is configured to perform statistical operations on the plurality of emotion state information to obtain a statistical result, wherein the statistical result is a statistical result corresponding to the current speech segment;
[0161] The determination module 410 is further configured to determine, based on the second neural network model, the emotional state information corresponding to the current utterance segment according to the statistical results corresponding to the current utterance segment and the n-1 statistical results corresponding to the n-1 utterance segments before the current utterance segment,
[0162] Among them, the n-1 speech segments correspond one-to-one to the n-1 statistical results, and the statistical result corresponding to any speech segment in the n-1 speech segments is obtained by performing statistical operations on multiple emotional state information corresponding to multiple speech frames included in the speech segment. The n-1 speech segments belong to the target dialogue, and n is an integer greater than 1.
[0163] The device provided in this application, based on a first neural network model, can obtain multiple emotion state information corresponding to multiple speech frames in the current utterance segment. Then, based on the statistical results corresponding to the current utterance segment and the statistical results corresponding to multiple utterance segments before the current utterance segment, it can obtain the emotion state information corresponding to the current utterance segment. Therefore, by using two neural network models, the first neural network model and the second neural network model, it is possible to more fully learn the impact of the context of the current utterance segment on the emotion state information corresponding to the current utterance segment, thereby achieving more accurate speech emotion recognition results.
[0164] It should be understood that the division of the above modules is merely a functional division, and there may be other division methods in actual implementation.
[0165] Figure 9 Schematic diagram of the hardware structure of the neural network training device provided in the embodiment of the present application. Figure 9 The neural network training device 500 shown (the device 500 may be a computer device) includes a memory 501, a processor 502, a communication interface 503, and a bus 504. The memory 501, the processor 502, and the communication interface 503 are connected to each other via the bus 504.
[0166] Memory 501 can be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). Memory 501 can store programs. When the program stored in memory 501 is executed by processor 502, processor 502 and communication interface 503 are used to perform the various steps of the neural network training method of the embodiment of the present application.
[0167] The processor 502 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits, and is used to execute relevant programs to implement the functions required to be performed by the units in the neural network training device of the embodiment of the present application, or to execute the neural network training method of the method embodiment of the present application.
[0168] The processor 502 may also be an integrated circuit chip with signal processing capabilities. During implementation, the various steps of the neural network training method of the present application may be completed by hardware integrated logic circuits or software instructions in the processor 502. The aforementioned processor 502 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory 501, and the processor 502 reads the information in the memory 501 and combines its hardware to complete the functions required to be performed by the units included in the neural network training device of the embodiment of the present application, or executes the neural network training method of the method embodiment of the present application.
[0169] The communication interface 503 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the apparatus 500 and other devices or a communication network. For example, training corpus can be obtained through the communication interface 503.
[0170] The bus 504 may include a path for transmitting information between various components of the device 500 (eg, the memory 501 , the processor 502 , and the communication interface 503 ).
[0171] Figure 10 Schematic diagram of the hardware structure of the speech emotion recognition device according to the embodiment of the present application. Figure 10 The illustrated speech emotion recognition apparatus 600 (the apparatus 600 may be a computer device) includes a memory 601, a processor 602, a communication interface 603, and a bus 604. The memory 601, the processor 602, and the communication interface 603 are connected to each other via the bus 604.
[0172] The memory 601 may be a ROM, a static storage device, or a RAM. The memory 601 may store a program. When the program stored in the memory 601 is executed by the processor 602, the processor 602 and the communication interface 603 are used to perform the various steps of the speech emotion recognition method of the embodiment of the present application.
[0173] The processor 602 can be a general-purpose CPU, microprocessor, ASIC, GPU or one or more integrated circuits to execute relevant programs to implement the functions required to be performed by the module in the speech emotion recognition embodiment of the present application, or to execute the speech emotion recognition method of the method embodiment of the present application.
[0174] The processor 602 may also be an integrated circuit chip with signal processing capabilities. During implementation, the various steps of the speech emotion recognition method of the embodiments of the present application may be completed by hardware integrated logic circuits or software instructions in the processor 602. The aforementioned processor 602 may also be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 601. The processor 602 reads the information in the memory 601 and, in conjunction with its hardware, completes the functions required to be performed by the modules included in the speech emotion recognition device of the embodiments of the present application, or executes the speech emotion recognition method of the method embodiments of the present application.
[0175] The communication interface 603 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the apparatus 600 and other devices or a communication network. For example, training corpus can be obtained through the communication interface 603.
[0176] The bus 604 may include a path for transmitting information between various components of the device 600 (eg, the memory 601 , the processor 602 , and the communication interface 603 ).
[0177] It should be understood that the determination module 410 and the statistics module 420 in the speech emotion recognition device 400 are equivalent to the processor 602 .
[0178] It should be noted that although Figure 9 and Figure 10 The devices 500 and 600 shown in the figure only show a memory, a processor, and a communication interface. However, in the specific implementation process, those skilled in the art should understand that the devices 500 and 600 also include other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the devices 500 and 600 may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the devices 500 and 600 may also only include the devices necessary to implement the embodiments of the present application, and do not necessarily include Figure 9 or Figure 10 All devices shown in .
[0179] It can be understood that the apparatus 500 is equivalent to the training device 120 in 1, and the apparatus 600 is equivalent to Figure 1 The execution device 110 in.
[0180] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which includes: computer program code, which, when running on a computer, enables the computer to execute the methods described above.
[0181] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable medium, which stores program code. When the program code is run on a computer, the computer executes the parties described above.
[0182] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in accordance with the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0183] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0184] It should also be understood that in this application, "when", "if" and "if" all mean that the terminal device or network device will make corresponding processing under certain objective circumstances, and it does not limit the time, nor does it require the terminal device or network device to make a judgment action when implementing it, nor does it mean that there are other limitations.
[0185] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0186] In this document, the term "at least one of..." or "at least one of..." or "at least one item of..." means all or any combination of the listed items. For example, "at least one of A, B and C" may mean: A exists alone, B exists alone, C exists alone, A and B exist at the same time, B and C exist at the same time, and A, B and C exist at the same time.
[0187] It should be understood that in each embodiment of the present application, "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A, but B can also be determined based on A and / or other information.
[0188] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0189] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0190] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0191] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0192] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0193] In the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0194] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0195] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A speech emotion recognition method, characterized in that: include: Determining, based on the first neural network model, a plurality of emotion state information corresponding to a plurality of speech frames included in a current utterance segment in the target conversation, wherein each speech frame corresponds to a piece of emotion state information, and the emotion state information represents the emotion state corresponding to the speech frame; Performing statistical operations on the plurality of emotional state information to obtain statistical results, wherein the statistical results are statistical results corresponding to the current speech segment; Based on the second neural network model, the emotional state information corresponding to the current utterance segment is determined according to the statistical results corresponding to the current utterance segment and the n-1 statistical results corresponding to the n-1 utterance segments before the current utterance segment. Among them, the n-1 speech segments correspond one-to-one to the n-1 statistical results, and the statistical result corresponding to any speech segment in the n-1 speech segments is obtained by performing statistical operations on multiple emotional state information corresponding to multiple speech frames included in the speech segment. The n-1 speech segments belong to the target dialogue, and n is an integer greater than 1.
2. The method according to claim 1, wherein The n-1 speech segments correspond to multiple speakers.
3. The method according to claim 2, wherein The multiple speakers include the speaker corresponding to the current speech segment; Furthermore, the determining, based on the second neural network model, the emotional state information corresponding to the current utterance segment according to the statistical results corresponding to the current utterance segment and the n-1 statistical results corresponding to the n-1 utterance segments before the current utterance segment, includes: Based on the second neural network model, the emotional state information corresponding to the current speech segment is determined according to the statistical results corresponding to the current speech segment, the n-1 statistical results, and the genders of the multiple speakers.
4. The method according to claim 1, wherein The method further comprises: Presenting the emotional state information corresponding to the current speech segment to the user; Obtaining the user's correction operation on the emotional state information corresponding to the current speech segment.
5. The method according to claim 1, wherein The n-1 speech segments are adjacent in time, and n is an integer greater than 2.
6. The method according to claim 5, wherein The determining, based on the second neural network model, the emotional state information corresponding to the current utterance segment according to the statistical results corresponding to the current utterance segment and n-1 statistical results corresponding to n-1 utterance segments before the current utterance segment, includes: Determine, based on the statistical result corresponding to the current utterance segment and the n-1 statistical results, the turn features corresponding to the w turns corresponding to the n utterance segments, the current utterance segment and the n-1 utterance segments, respectively, wherein the turn feature corresponding to any turn is determined by the statistical results corresponding to the utterance segments of all speakers in the turn, and w is an integer greater than or equal to 1; Based on the second neural network model, the emotional state information corresponding to the current speech segment is determined according to the round features corresponding to the w rounds.
7. The method according to claim 6, wherein w is the value entered by the user.
8. The method according to claim 6, wherein The method further comprises: Update the value of w.
9. The method according to claim 1, wherein The first neural network model is a long short-term memory model LSTM; and / or the second neural network model is LSTM.
10. The method according to any one of claims 1 to 9, characterized in that The determining, based on the first neural network model, multiple emotion state information corresponding to multiple speech frames included in the current speech segment in the target dialogue includes: Based on the first neural network model, for each speech frame among the multiple speech frames, the emotional state information corresponding to the speech frame is determined according to the feature vector corresponding to the speech frame and the feature vectors corresponding to the q-1 speech frames before the speech frame, wherein the q-1 speech frames are the speech frames of the speaker corresponding to the current speech segment, q is an integer greater than 1, and the feature vector of the speech frame k represents the acoustic features of the speech frame k.
11. The method according to claim 10, wherein There are m speech frames between any two speech frames in the q speech frames, where m is an integer greater than or equal to 0.
12. A speech emotion recognition device, characterized in that: include: a determination module configured to determine, based on the first neural network model, a plurality of emotion state information corresponding to a plurality of speech frames included in a current utterance segment in the target dialogue, wherein each speech frame corresponds to a piece of emotion state information, and the emotion state information represents the emotion state corresponding to the speech frame; A statistical module, configured to perform statistical operations on the plurality of emotional state information to obtain statistical results, wherein the statistical results are statistical results corresponding to the current speech segment; The determination module is further configured to determine, based on a second neural network model, the emotional state information corresponding to the current utterance segment according to the statistical results corresponding to the current utterance segment and the n-1 statistical results corresponding to the n-1 utterance segments preceding the current utterance segment, Among them, the n-1 speech segments correspond one-to-one to the n-1 statistical results, and the statistical result corresponding to any speech segment in the n-1 speech segments is obtained by performing statistical operations on multiple emotional state information corresponding to multiple speech frames included in the speech segment. The n-1 speech segments belong to the target dialogue, and n is an integer greater than 1.
13. The device according to claim 12, wherein The n-1 utterance segments include speech data of multiple speakers.
14. The device according to claim 13, wherein The multiple speakers include the speaker corresponding to the current speech segment; And, the determination module is specifically used for: Based on the second neural network model, the emotional state information corresponding to the current speech segment is determined according to the statistical results corresponding to the current speech segment, the n-1 statistical results, and the genders of the multiple speakers.
15. The device according to claim 12, wherein The device further comprises: A presentation module, configured to present the emotional state information corresponding to the current speech segment to the user; The acquisition module is used to obtain the user's correction operation on the emotional state information corresponding to the current speech segment.
16. The device according to claim 12, wherein The n-1 speech segments are adjacent in time, and n is an integer greater than 2.
17. The device according to claim 16, wherein The determining module is specifically configured to: Determine, based on the statistical result corresponding to the current utterance segment and the n-1 statistical results, the turn features corresponding to the w turns corresponding to the n utterance segments, the current utterance segment and the n-1 utterance segments, respectively, wherein the turn feature corresponding to any turn is determined by the statistical results corresponding to the utterance segments of all speakers in the turn, and w is an integer greater than or equal to 1; Based on the second neural network model, the emotional state information corresponding to the current speech segment is determined according to the round features corresponding to the w rounds.
18. The device according to claim 17, wherein w is the value entered by the user.
19. The device according to claim 17, wherein The device further comprises: Update module, used to update the value of w.
20. The device according to claim 12, wherein The first neural network model is a long short-term memory model LSTM; and / or the second neural network model is LSTM.
21. The device according to any one of claims 12 to 20, characterized in that The determining module is specifically configured to: Based on the first neural network model, for each speech frame among the multiple speech frames, the emotional state information corresponding to the speech frame is determined according to the feature vector corresponding to the speech frame and the feature vectors corresponding to the q-1 speech frames before the speech frame, wherein the q-1 speech frames are the speech frames of the speaker corresponding to the current speech segment, q is an integer greater than 1, and the feature vector of the speech frame k represents the acoustic features of the speech frame k.
22. The device according to claim 21, wherein There are m speech frames between any two speech frames in the q speech frames, where m is an integer greater than or equal to 0.
23. A speech emotion recognition device, characterized in that: include: Memory, used to store programs; A processor, configured to execute the program stored in the memory; when the program stored in the memory is executed, the processor is configured to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Processing method and system for voice data
CN102831891A
Method and system for voice emotion inference on basis of emotion context
CN103810994A