Training Method, Device, Equipment and Storage Medium of Emotion Recognition Model

By performing feature extraction, audio filtering and data enhancement on sample audio, training data and training neural network models, the problem of insufficient accuracy of emotion recognition in the prior art is solved, and the accuracy of negative emotion recognition is improved.

CN112382309BActive Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011446542.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-11
Publication Date
2025-06-27
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify user emotions, especially negative emotions, which leads to the inability to effectively evaluate user satisfaction with the service process.

Method used

By obtaining sample audio, extracting speech features, and filtering positive emotions audio, data enhancement of negative emotions audio, forming filtered sample audio and newly added negative emotions audio as training data, inputting the preset neural network for model training, obtaining the emotion recognition model.

Benefits of technology

The accuracy of the emotional recognition model for negative emotions audio is improved, and the data imbalance problem during model training is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112382309B_ABST
    Figure CN112382309B_ABST
Patent Text Reader

Abstract

This application relates to the field of classification models, and discloses a training method, device, equipment and storage medium for an emotion recognition model. The method includes: obtaining sample audio, where the sample audio includes positive emotion audio and negative emotion audio, and respectively extracting features from the positive emotion audio and the negative emotion audio to obtain speech features; filtering the positive emotion audio in the sample audio according to the speech features to obtain filtered sample audio; performing data augmentation on the negative emotion audio in the sample audio to obtain additional negative emotion audio; inputting the filtered sample audio and the additional negative emotion audio into a preset neural network for model training to obtain an emotion recognition model, so that the emotion recognition model can accurately recognize the emotions of users. In addition, the present invention also relates to blockchain technology, and the sample audio can be stored in the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of model training, and particularly to a method, device, equipment and storage medium for training an emotion recognition model. Background Art

[0002] With the rapid development of the Internet, a large number of services have started to be processed online. Currently, a large number of online customer services are set up to answer users' questions. To ensure the service quality of the customer service, most of them will record the call process and record the service process by converting the recording into text. However, the method of converting speech to text can only know the text content of the recording, and it is impossible to recognize the emotions of both parties in the conversation, especially the emotions of the customers. This leads to the inability to recognize the negative emotions of users, and thus it is impossible to know whether the users are satisfied with the current service process.

[0003] Therefore, how to train an emotion recognition model so that the emotion recognition model can accurately recognize the emotions of users has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a method, device, equipment and storage medium for training an emotion recognition model, so that the emotion recognition model can accurately recognize the emotions of users.

[0005] In a first aspect, this application provides a method for training an emotion recognition model, and the method includes:

[0006] Obtain sample audio, where the sample audio includes positive emotion audio and negative emotion audio, and respectively extract features from the positive emotion audio and the negative emotion audio to obtain speech features; filter the positive emotion audio in the sample audio according to the speech features to obtain filtered sample audio; perform data augmentation on the negative emotion audio in the sample audio to obtain additional negative emotion audio; use the filtered sample audio and the additional negative emotion audio as training data, and input the training data into a preset neural network for model training to obtain an emotion recognition model.

[0007] In a second aspect, this application also provides a device for training an emotion recognition model, and the device includes:

[0008] A feature extraction module for obtaining a sample audio, where the sample audio includes a positive emotion audio and a negative emotion audio, and respectively extracting features from the positive emotion audio and the negative emotion audio to obtain speech features; an audio filtering module for filtering the positive emotion audio in the sample audio according to the speech features to obtain a filtered sample audio; a data augmentation module for augmenting the negative emotion audio in the sample audio to obtain additional negative emotion audio; a model training module for using the filtered sample audio and the additional negative emotion audio as training data, and inputting the training data into a preset neural network for model training to obtain an emotion recognition model.

[0009] In a third aspect, the present application further provides a computer device, which includes a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program and implementing the training method of the emotion recognition model as described above when executing the computer program.

[0010] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the training method of the emotion recognition model as described above.

[0011] The present application discloses a training method, device, equipment and storage medium for an emotion recognition model. By obtaining a sample audio, where the sample audio includes a positive emotion audio and a negative emotion audio, then respectively extracting features from the positive emotion audio and the negative emotion audio to obtain speech features, then filtering the positive emotion audio in the sample audio according to the speech features to obtain a filtered sample audio, further augmenting the negative emotion audio in the sample audio to obtain additional negative emotion audio, and finally using the filtered sample audio and the additional negative emotion audio as training data to perform model training on a preset neural network to obtain an emotion recognition model. By filtering the positive emotion audio in the sample audio, the data proportion of the negative emotion audio in the filtered sample audio is increased, and then the negative emotion audio in the sample audio is augmented, and the obtained additional negative emotion audio and the filtered sample audio are jointly used as training data for model training, further increasing the data proportion of the negative emotion audio in the training data, thereby solving the data imbalance problem in the model training process and improving the recognition accuracy of the trained emotion recognition model for the negative emotion audio. Description of the Drawings

[0012] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0013] Figure 1 It is a schematic flowchart of a method for training an emotion recognition model provided by an embodiment of the present application;

[0014] Figure 2 It is a schematic flowchart of the steps for audio filtering of positive emotion audio provided by an embodiment of the present application;

[0015] Figure 3 It is a schematic block diagram of a device for training an emotion recognition model provided by an embodiment of the present application;

[0016] Figure 4 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0018] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0019] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0020] It should also be understood that the term " / and" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0021] Embodiments of the present application provide a method, apparatus, computer device, and storage medium for training an emotion recognition model. The method for training an emotion recognition model can be used to train an emotion recognition model for recognizing the negative emotions of a user based on the user's audio, and improve the accuracy of the trained emotion recognition model in recognizing negative emotions. The trained emotion recognition model can recognize the negative emotions of the user based on the audio.

[0022] For example, the emotion recognition model trained by the method for training an emotion recognition model provided by the embodiments of the present application can be applied to an online customer service system. By recognizing the emotions in the audio during the service process, the negative emotions of the user can be learned, thereby improving the quality of customer service.

[0023] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0024] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for training an emotion recognition model provided by an embodiment of the present application. The method for training an emotion recognition model improves the data ratio of negative emotion data in the training data by performing audio filtering on positive emotion audio and data augmentation on negative emotion audio, thereby solving the problem of data imbalance in the model training process and improving the recognition accuracy of the emotion recognition model for negative emotions.

[0025] As Figure 1 shown, the method for training an emotion recognition model specifically includes: steps S101 to S104.

[0026] S101. Obtain sample audio, where the sample audio includes positive emotion audio and negative emotion audio, and respectively extract features from the positive emotion audio and the negative emotion audio to obtain speech features.

[0027] Among them, the sample audio refers to the business data actually generated during the business process. For example, taking the customer service system as an example, the sample audio is the recorded audio in the historical customer service process. Since the proportion of negative emotion audio in the actually generated business data is extremely small, it is necessary to process the sample audio to reduce the data ratio in the sample audio.

[0028] The sample audio includes positive emotion audio and negative emotion audio. Positive emotion audio refers to the emotion of the speaker being positive, such as happy, excited, grateful, etc., while negative emotion audio refers to the emotion of the speaker being negative, such as dissatisfied, indignant, angry, etc.

[0029] Feature extraction is performed on the positive emotion audio and the negative emotion audio, specifically referring to extracting speech features from the positive emotion audio and the negative emotion audio. Among them, the speech features may include fundamental frequency, characteristic values corresponding to the fundamental frequency, sound intensity, characteristic values corresponding to the sound intensity, spectrum, and characteristic values corresponding to the spectrum.

[0030] In one embodiment, the method for training the emotion recognition model includes: removing noise from the sample audio to obtain the sample audio after noise removal.

[0031] Since the actually collected audio often has background sound with a certain intensity, and these background sounds are generally background noise. When the intensity of the background noise is relatively large, a large number of irrelevant features will be included in the speech features extracted from the audio. Therefore, noise can be removed from the sample audio to obtain the sample audio after noise removal, thereby improving the accuracy of the obtained speech features.

[0032] In the specific implementation process, the Fourier transform can be used to process the sample audio to obtain the spectrum of the sample audio, then extract the spectrum of the noise from the spectrum of the sample audio, and perform an inverse compensation operation on the sample audio according to the spectrum of the noise, so as to obtain the sample audio after noise removal.

[0033] In one embodiment, the method for training the emotion recognition model includes: performing audio analysis on the sample audio to obtain the change in the signal energy value of the sample audio; performing endpoint detection on the sample audio according to the change in the signal energy value of the sample audio, and cutting the sample audio based on the detected endpoints to obtain the audible audio segment in the sample audio.

[0034] Performing audio analysis on the sample audio specifically means obtaining the amplitude of each frame in the sample audio. Since the smaller the sound, the smaller the amplitude of the sound wave, and the amplitude of the sound wave also represents the magnitude of the signal energy value, the smaller the amplitude of the sound wave, the smaller the energy value of the signal. Therefore, endpoint detection can be performed on the sample audio according to the change in the signal energy value of the sample audio to obtain the front and rear endpoints of the audible audio segment in the sample audio, and then cut the sample audio according to the front and rear endpoints.

[0035] For example, when the signal energy values of the sample audio in a continuous number of frames (between the 0th frame and the Nth frame) are all lower than the energy value threshold E, and the signal energy values in the subsequent continuous number of frames (between the Nth frame and the Mth frame) are all higher than the energy value threshold E, then the place where the energy value of the sample audio increases (the Nth frame) is considered the front endpoint of the sample audio.

[0036] Similarly, when the signal energy values of the sample audio within a number of consecutive frames (between the Nth frame and the Mth frame) are all higher than the energy value threshold E, and the signal energy values within the next number of consecutive frames (between the Mth frame and the Pth frame) are all lower than the energy value threshold E, then the point where the sample audio energy value decreases (the Mth frame) is recognized as the end point of the sample audio.

[0037] Cut the sample audio according to the start point (the Nth frame) and the end point (the Mth frame) of the sample audio to obtain the audible audio segment (from the Nth frame to the Mth frame) in the sample audio.

[0038] In one embodiment, the training method of the emotion recognition model includes: performing speech recognition on the sample audio to determine whether the sample audio includes speech information; if the sample audio does not include speech information, then delete the sample audio; if the sample audio partially includes speech information, then cut the sample audio to obtain an audio segment including speech information.

[0039] For a sample audio, in addition to possibly including some silent segments, it may also include some invalid sound segments, such as snoring sounds, etc. Therefore, by performing speech recognition on the sample audio to determine whether the sample audio includes speech information, and screening and cutting the sample audio according to the speech information, an audio segment including speech information in the sample audio can be obtained, which can reduce the invalid audio in the sample audio and thus improve the recognition accuracy of the trained emotion recognition model.

[0040] Therefore, when performing speech recognition on the sample audio and obtaining that the sample audio does not include speech information, it is considered that the sample audio is an invalid sample audio, and the sample audio can be deleted.

[0041] When performing speech recognition on the sample audio and obtaining that at least a part of the sample audio includes speech information, the part of the sample audio including speech information can be cut to obtain an audio segment including speech information.

[0042] S102. Perform audio filtering on the positive emotion audio in the sample audio according to the speech feature to obtain a filtered sample audio.

[0043] Perform audio filtering on the positive emotion audio in the sample audio according to the speech feature, so as to reduce the number of positive emotion audio in the sample audio and achieve the purpose of reducing the positive-negative data ratio.

[0044] In one embodiment, please refer to Figure 2, the steps of audio filtering for positive emotion audio specifically include: S1021, analyzing the positive eigenvalue of the speech feature of the positive emotion audio and the negative eigenvalue of the speech feature of the negative emotion audio to obtain a regular curve of speech features, eigenvalues of speech features, and emotion categories; S1022, determining a screening threshold based on the regular curve, and performing audio filtering on the positive emotion audio in the sample audio according to the screening threshold and the speech feature.

[0045] For the extracted speech features, analyze the positive eigenvalue of each positive emotion audio in the sample audio for this speech feature, and then synthesize the positive eigenvalues of multiple positive emotion audios to obtain the changing trend of the positive eigenvalue of the positive emotion audio for this speech feature.

[0046] Similarly, analyze the negative eigenvalue of each negative emotion audio in the sample audio for this speech feature, and then synthesize the negative eigenvalues of multiple negative emotion audios to obtain the changing trend of the negative eigenvalue of the negative emotion audio for this speech feature.

[0047] For example, when an audio is a positive emotion audio, what is the positive eigenvalue corresponding to the fundamental frequency of the audio? Synthesize the positive eigenvalues corresponding to the fundamental frequencies of multiple positive emotion audios to obtain the changing trend of the positive eigenvalue corresponding to the fundamental frequency for the speech feature of the fundamental frequency in positive emotion audios.

[0048] When an audio is a negative emotion audio, what is the negative eigenvalue corresponding to the fundamental frequency of the audio? Synthesize the negative eigenvalues corresponding to the fundamental frequencies of multiple negative emotion audios to obtain the changing trend of the negative eigenvalue corresponding to the fundamental frequency for the speech feature of the fundamental frequency in negative emotion audios.

[0049] For the same speech feature, based on the changing trend of the eigenvalues of the speech feature in audios of different emotion types, it can be known that under this speech feature, the changing trend from positive emotion audio to negative emotion audio. Plot this changing trend as a regular curve of speech features, eigenvalues of speech features, and emotion categories.

[0050] Determine the screening threshold according to this regular curve, and then perform audio filtering on the positive emotion audio in the sample audio according to the screening threshold and the speech feature, so as to reduce the number of positive emotion audios in the sample audio.

[0051] It should be noted that the screening threshold determined according to the regular curve is not fixed and can be adaptively adjusted based on the regular curve according to the actual training situation.

[0052] When there are multiple voice features, a regular curve of the voice feature, the eigenvalue of the voice feature, and the emotion category can be constructed for each voice feature respectively, and then the screening threshold of each voice feature can be determined according to multiple regular curves.

[0053] S103. Perform data augmentation on the negative emotion audio in the sample audio to obtain additional negative emotion audio.

[0054] Among them, data augmentation means constructing a virtual sample based on the negative emotion audio in the sample audio, and using the constructed virtual sample as the additional negative emotion audio. By data augmentation, the data ratio of positive emotion audio and negative emotion audio is further reduced, and the generalization ability of the trained emotion recognition model is improved.

[0055] In one embodiment, the data augmentation includes at least one of speech rate perturbation, phase perturbation, and spectral masking.

[0056] Speech rate perturbation means linearly stretching or compressing the sample audio in the time domain. Since the perturbed sample audio also changes in proportion in the frequency domain, under the combined action of the changes in the time domain and the frequency domain, after the perturbed sample audio is subjected to feature extraction, it will show a certain difference from the original sample audio, thus achieving the purpose of constructing a new sample. For example, the speed of the sample audio in the time domain is adjusted to three levels of 0.9, 1.0, and 1.1 respectively.

[0057] The rule of phase perturbation is the same as that of speech rate perturbation. Spectral masking mainly performs a zero assignment process on some frequency domain values in the sample audio. The sample audio after the operation will not affect the actual use, and only masking processing is performed in some frequency bands.

[0058] S104. Use the filtered sample audio and the additional negative emotion audio as training data, and input the training data into a preset neural network for model training to obtain an emotion recognition model.

[0059] Use the filtered sample audio and the additional negative emotion audio together as training data to perform model training on a preset neural network. When the preset neural network is trained to convergence, the converged neural network is used as the emotion recognition model.

[0060] By filtering the positive emotion audio in the sample audio and performing data augmentation on the negative emotion audio, the ratio of positive emotion audio and negative emotion audio in the training data is reduced, so that when model training is performed based on real business data, the data imbalance problem between positive emotion audio and negative emotion audio in the model training process can be solved, and the recognition accuracy of the trained emotion recognition model for negative emotion audio can be improved.

[0061] In one embodiment, the preset neural network includes an input layer, a feature extraction layer, a hidden layer, a pooling layer, and an output layer; the step of inputting the training data into the preset neural network for model training includes: inputting the training data into the preset neural network through the input layer; performing feature extraction on the training data based on the feature extraction layer to obtain first training features; inputting the first training features into the hidden layer to obtain second training features corresponding to the first training features; performing feature dimensionality reduction on the second training features based on the pooling layer to obtain third training features; performing classification based on the third training features and outputting a classification result through the output layer; and performing iterative training on the preset neural network based on the classification result and the emotion type of the audio in the training data.

[0062] The training data is input through the input layer of the neural network. After the input training data reaches the feature extraction layer, the feature extraction layer performs feature extraction on the training data to obtain first training features. The first training features are input into the hidden layer of the neural network. The hidden layer can be an RNN network. The high-dimensional features of the input training data, that is, the second training features, are obtained through the hidden layer. Then, the pooling layer performs dimensionality reduction on the second training features to map the high-dimensional features to a low dimension and obtain third training features. Finally, binary classification is performed based on the third training features, and the classification result is output through the output layer.

[0063] The loss function of the preset neural network is calculated based on the classification result and the emotion type of the audio in the training data, and the parameters of the neural network are iteratively updated until the value of the loss function reaches a preset value. It is considered that the neural network converges, and the trained neural network is used as an emotion recognition model to complete the model training.

[0064] The training method of the emotion recognition model provided in the above embodiments includes obtaining sample audio, where the sample audio includes positive emotion audio and negative emotion audio, then respectively extracting features from the positive emotion audio and the negative emotion audio to obtain speech features, then filtering the positive emotion audio in the sample audio according to the speech features to obtain the filtered sample audio, and then performing data augmentation on the negative emotion audio in the sample audio to obtain additional negative emotion audio. Finally, the filtered sample audio and the additional negative emotion audio are used as training data to train a preset neural network to obtain an emotion recognition model. By filtering the positive emotion audio in the sample audio, the data proportion of the negative emotion audio in the filtered sample audio is increased. Then, data augmentation is performed on the negative emotion audio in the sample audio, and the obtained additional negative emotion audio and the filtered sample audio are jointly used as training data for model training to further increase the data proportion of the negative emotion audio in the training data, thereby solving the problem of data imbalance in the model training process and improving the recognition accuracy of the trained emotion recognition model for negative emotion audio.

[0065] Please refer to Figure 3 , Figure 3 FIG. is a schematic block diagram of a training device for an emotion recognition model provided in an embodiment of the present application. The training device for the emotion recognition model is used to execute the foregoing training method of the emotion recognition model. Among them, the training device for the emotion recognition model can be configured in a server or a terminal.

[0066] Among them, the server can be an independent server or a server cluster. The terminal can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device.

[0067] As Figure 3 shown, the training device 200 for the emotion recognition model includes: a feature extraction module 201, an audio filtering module 202, a data augmentation module 203, and a model training module 204.

[0068] The feature extraction module 201 is used to obtain sample audio, where the sample audio includes positive emotion audio and negative emotion audio, and respectively extract features from the positive emotion audio and the negative emotion audio to obtain speech features.

[0069] The audio filtering module 202 is used to filter the positive emotion audio in the sample audio according to the speech features to obtain the filtered sample audio.

[0070] In one embodiment, the audio filtering module 202 includes a curve construction sub-module 2021 and a threshold filtering sub-module 2022.

[0071] Among them, the curve construction sub-module 2021 is used to analyze the positive eigenvalue of the speech feature of the positive emotion audio and the negative eigenvalue of the speech feature of the negative emotion audio, and obtain the regular curve of the speech feature, the eigenvalue of the speech feature, and the emotion category. The threshold filtering sub-module 2022 is used to determine the screening threshold based on the regular curve, and filter the positive emotion audio in the sample audio according to the screening threshold and the speech feature.

[0072] The data augmentation module 203 is used to perform data augmentation on the negative emotion audio in the sample audio to obtain new negative emotion audio.

[0073] The model training module 204 is used to use the filtered sample audio and the new negative emotion audio as training data, and input the training data into a preset neural network for model training to obtain an emotion recognition model.

[0074] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the training device of the emotion recognition model and each module described above can refer to the corresponding processes in the foregoing embodiments of the training method of the emotion recognition model, and will not be elaborated herein.

[0075] The above-mentioned training device of the emotion recognition model can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 4 the figure.

[0076] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal.

[0077] Refer to Figure 4 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0078] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the training methods of the emotion recognition model.

[0079] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0080] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the training methods of the emotion recognition model.

[0081] The network interface is used for network communication, such as sending the assigned tasks, etc. Those skilled in the art can understand that Figure 4 the structure shown in Figure 4 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0082] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0083] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:

[0084] Obtain a sample audio, where the sample audio includes positive emotion audio and negative emotion audio, and respectively extract features of the positive emotion audio and the negative emotion audio to obtain speech features; filter the positive emotion audio in the sample audio according to the speech features to obtain a filtered sample audio; perform data augmentation on the negative emotion audio in the sample audio to obtain additional negative emotion audio; use the filtered sample audio and the additional negative emotion audio as training data, and input the training data into a preset neural network for model training to obtain an emotion recognition model.

[0085] In one embodiment, when the processor implements filtering the positive emotion audio in the sample audio according to the speech features, it is used to implement:

[0086] Analyze the positive eigenvalue of the speech feature of the positive emotion audio and the negative eigenvalue of the speech feature of the negative emotion audio to obtain a regular curve of the speech feature, the eigenvalue of the speech feature, and the emotion category; determine a screening threshold based on the regular curve, and filter the positive emotion audio in the sample audio according to the screening threshold and the speech feature.

[0087] In one embodiment, the processor is configured to implement:

[0088] Perform audio analysis on the sample audio to obtain the change in the signal energy value of the sample audio; perform endpoint detection on the sample audio according to the change in the signal energy value of the sample audio, and cut the sample audio based on the detected endpoints to obtain the audible audio segment in the sample audio.

[0089] In one embodiment, the processor is configured to implement:

[0090] Perform speech recognition on the sample audio to determine whether the sample audio includes speech information; if the sample audio does not include speech information, delete the sample audio; if the sample audio partially includes speech information, cut the sample audio to obtain the audio segment including speech information.

[0091] In one embodiment, the preset neural network includes an input layer, a feature extraction layer, a hidden layer, a pooling layer, and an output layer; when the processor implements inputting the training data into the preset neural network for model training, it is configured to implement:

[0092] Input the training data into the preset neural network through the input layer; perform feature extraction on the training data based on the feature extraction layer to obtain the first training feature; input the first training feature into the hidden layer to obtain the second training feature corresponding to the first training feature; perform feature dimensionality reduction on the second training feature based on the pooling layer to obtain the third training feature; perform classification based on the third training feature, and output the classification result through the output layer; perform iterative training on the preset neural network based on the classification result and the emotion type of the audio in the training data.

[0093] In one embodiment, the processor is configured to implement:

[0094] Perform noise removal on the sample audio to obtain the sample audio after noise removal.

[0095] In one embodiment, the data augmentation includes at least one of speech rate perturbation, phase perturbation, and spectral masking.

[0096] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the emotion recognition model training methods provided by the embodiments of the present application.

[0097] Among them, the computer-readable storage medium may be the internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.

[0098] Furthermore, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.

[0099] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0100] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A training method for an emotion recognition model, characterized in that, Including: Obtain a sample audio, where the sample audio includes a positive emotion audio and a negative emotion audio, and respectively extract features from the positive emotion audio and the negative emotion audio to obtain speech features; Filter the positive emotion audio in the sample audio according to the speech features to obtain a filtered sample audio; Perform data augmentation on the negative emotion audio in the sample audio to obtain a new negative emotion audio; Use the filtered sample audio and the new negative emotion audio as training data, and input the training data into a preset neural network for model training to obtain an emotion recognition model.

2. The training method of the emotion recognition model according to claim 1, characterized in that The filtering the positive emotion audio in the sample audio according to the speech features includes: Analyze the positive eigenvalue of the speech features of the positive emotion audio and the negative eigenvalue of the speech features of the negative emotion audio to obtain a regular curve of the speech features, the eigenvalues of the speech features, and the emotion categories; Determine a screening threshold based on the regular curve, and filter the positive emotion audio in the sample audio according to the screening threshold and the speech features; Among them, the analyzing the positive eigenvalue of the speech features of the positive emotion audio and the negative eigenvalue of the speech features of the negative emotion audio to obtain a regular curve of the speech features, the eigenvalues of the speech features, and the emotion categories includes: Analyze the positive eigenvalue of the speech features for each positive emotion audio in the sample audio to obtain the change trend of the positive eigenvalue corresponding to the speech features; Analyze the negative eigenvalue of the speech features for each negative emotion audio in the sample audio to obtain the change trend of the negative eigenvalue corresponding to the speech features; Based on the change trend of the positive eigenvalue and the change trend of the negative eigenvalue corresponding to the same speech feature, draw the change trend from the positive emotion audio to the negative emotion audio corresponding to the speech feature to obtain the regular curve.

3. The training method of the emotion recognition model according to claim 1, characterized in that The method includes: Perform audio analysis on the sample audio to obtain the change of the signal energy value of the sample audio; Perform endpoint detection on the sample audio according to the change of the signal energy value of the sample audio, and cut the sample audio based on the detected endpoints to obtain the audible audio segment in the sample audio.

4. The training method of the emotion recognition model according to claim 1, wherein The method includes: Perform speech recognition on the sample audio to determine whether the sample audio includes speech information; If the sample audio does not include speech information, delete the sample audio; If the sample audio partially includes speech information, cut the sample audio to obtain an audio segment including speech information.

5. The training method of the emotion recognition model according to claim 1, characterized in that The preset neural network includes an input layer, a feature extraction layer, a hidden layer, a pooling layer, and an output layer; the inputting the training data into the preset neural network for model training includes: Input the training data into the preset neural network through the input layer; Perform feature extraction on the training data based on the feature extraction layer to obtain first training features; Input the first training feature into the hidden layer to obtain a second training feature corresponding to the first training feature; Based on the pooling layer, perform feature dimensionality reduction on the second training feature to obtain a third training feature; Classify based on the third training feature and output a classification result through the output layer; Based on the classification result and the emotion type of the audio in the training data, perform iterative training on the preset neural network.

6. The training method of the emotion recognition model according to claim 1, characterized in that The method includes: Remove noise from the sample audio to obtain the sample audio after noise removal.

7. The training method of the emotion recognition model according to claim 1, characterized in that The data augmentation includes at least one of speech rate perturbation, phase perturbation, and spectral masking.

8. A training device for an emotion recognition model, characterized in that It includes: A feature extraction module, configured to obtain a sample audio, where the sample audio includes positive emotion audio and negative emotion audio, and respectively extract features from the positive emotion audio and the negative emotion audio to obtain speech features; An audio filtering module, configured to filter the positive emotion audio in the sample audio according to the speech features to obtain a filtered sample audio; A data augmentation module, configured to perform data augmentation on the negative emotion audio in the sample audio to obtain additional negative emotion audio; A model training module, configured to use the filtered sample audio and the additional negative emotion audio as training data, and input the training data into a preset neural network for model training to obtain an emotion recognition model.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the training method of the emotion recognition model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the training method of the emotion recognition model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Emotion recognition model training method and device, emotion recognition method and device, equipment and storage medium

    CN109800720A

  • Training method of emotion recognition model, emotion recognition method, device, equipment, and storage medium

    CN109817246A