Method and apparatus for recognizing speech emotion, processor and electronic device
By combining deep learning GNN model and traditional SVM model, the problem of low accuracy of speech emotion recognition is solved, more efficient user emotion recognition is achieved, and the effect of human-computer interaction is improved.
Patent Information
- Application Number
- CN202210428305.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-22
AI Technical Summary
In the prior art, the accuracy of speech emotion recognition is low, making it difficult to effectively identify the user's emotional state.
Speech sentiment recognition is performed using a mixed model based on GNN model and SVM model. By extracting MFCC feature, adding labels and proportional division of the target data set, the model is trained and combined to improve the recognition accuracy.
It improves the accuracy of voice emotion recognition, can better identify the user's emotional state, and improves the effect of human-computer interaction.
Smart Images

Figure CN114822597B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular, to a method and device for recognizing speech emotions, a processor, and an electronic device. Background Art
[0002] With the rapid development and wide application of the field of artificial intelligence, many aspects of human life are being affected by AI. For example, AI technology is used in fields such as image recognition and classification, speech recognition, and target retrieval. Among them, speech recognition is the most basic AI technology in speech interaction, and we are familiar with Siri, smart speakers, self-service voice customer service, etc. It can be seen that speech recognition has imperceptibly affected all aspects of people's lives and work.
[0003] In addition, an important behavioral signal reflecting human emotions is the emotional signal in speech, that is, the speech information carried by the same text spoken with different emotions may be completely different. Moreover, recognizing the emotions of users in speech is an important link in realizing human-computer interaction. For example, in the scenario of bank artificial customer service, the recognition of customer emotions can enable customer service personnel to timely discover the current emotional state of customers and better serve and promote customers according to different emotional states of customers. However, the accuracy of recognizing the emotions of users in speech in the current related technologies is relatively low.
[0004] Aiming at the problem of relatively low accuracy in recognizing the speech emotions of users in the related technologies, no effective solution has been proposed yet. Summary of the Invention
[0005] The main purpose of the present application is to provide a method and device for recognizing speech emotions, a processor, and an electronic device to solve the problem of relatively low accuracy in recognizing the speech emotions of users in the related technologies.
[0006] To achieve the above object, according to one aspect of the present application, a method for recognizing speech emotions is provided. The method includes: obtaining target speech information of a target object, where the target object is an object to be recognized for emotions; inputting the target speech information into a target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, where the target hybrid model is a model constructed based on a GNN model and an SVM model.
[0007] Further, before inputting the target voice information into the target hybrid model for emotion recognition processing, the method further includes: obtaining a target data set and obtaining a training set in the target data set; using the target data set to learn and train the GNN model to obtain a first recognition model; using the target data set to learn and train the SVM model to obtain a second recognition model; combining the first recognition model and the second recognition model according to preset requirements to obtain a first hybrid model; using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model.
[0008] Further, obtaining the training set in the target data set includes: performing MFCC feature extraction operation on the data in the target data set to obtain the data after feature extraction; obtaining the classification information of the preset speech emotion; adding labels to the data after feature extraction according to the classification information to obtain the data after adding labels; dividing the data after adding labels into the training set and the test set according to a preset ratio.
[0009] Further, after using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model, the method further includes: using the test set to test the target hybrid model to obtain a test result; determining the emotion recognition performance of the target hybrid model according to the test result.
[0010] Further, obtaining the target data set includes: obtaining a Chinese emotion corpus; using the Chinese emotion corpus as the target data set.
[0011] Further, after inputting the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, the method further includes: recommending a target product to the target object according to the emotion recognition result; or; adopting a target type of response manner for the target object according to the emotion recognition result.
[0012] To achieve the above object, according to another aspect of the present application, there is provided a device for recognizing speech emotion. The device includes: a first obtaining unit, configured to obtain target voice information of a target object, where the target object is an object to be subjected to emotion recognition; a first recognition unit, configured to input the target voice information into a target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, where the target hybrid model is a model constructed based on a GNN model and an SVM model.
[0013] Further, the device further includes: a second acquisition unit, configured to acquire a target data set and acquire a training set in the target data set before inputting the target voice information into a target hybrid model for emotion recognition processing; a first training unit, configured to perform learning and training on the GNN model by using the target data set to obtain a first recognition model; a second training unit, configured to perform learning and training on the SVM model by using the target data set to obtain a second recognition model; a first combination unit, configured to combine the first recognition model and the second recognition model according to a preset requirement to obtain a first hybrid model; a first construction unit, configured to perform regression verification on the first hybrid model by using the training set to obtain the target hybrid model.
[0014] Further, the second acquisition unit includes: a first extraction module, configured to perform MFCC feature extraction operation on the data in the target data set to obtain the data after feature extraction; a first acquisition module, configured to acquire classification information of preset speech emotions; a first addition module, configured to add labels to the data after feature extraction according to the classification information to obtain the data after adding labels; a first classification module, configured to divide the data after adding labels into the training set and the test set according to a preset ratio.
[0015] Further, the device further includes: a first test unit, configured to test the target hybrid model by using the test set after performing regression verification on the first hybrid model by using the training set to obtain the target hybrid model, so as to obtain a test result; a first determination unit, configured to determine the emotion recognition performance of the target hybrid model according to the test result.
[0016] Further, the second acquisition unit includes: a second acquisition module, configured to acquire a Chinese emotion corpus; a first determination module, configured to use the Chinese emotion corpus as the target data set.
[0017] Further, the device further includes: a first recommendation unit, configured to recommend a target product to the target object according to the emotion recognition result after inputting the target voice information into a target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object; or; a first response unit, configured to adopt a target type of response manner for the target object according to the emotion recognition result.
[0018] To achieve the above object, according to another aspect of the present application, there is provided a processor, where the processor is used to run a program, and when the program runs, it executes the method for recognizing speech emotion described in any one of the above.
[0019] To achieve the above object, according to another aspect of the present application, there is provided an electronic device, which includes one or more processors and a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for recognizing speech emotions described in any one of the above.
[0020] In the present application, the following steps are adopted: obtaining target speech information of a target object, where the target object is an object to be subjected to emotion recognition; inputting the target speech information into a target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, where the target hybrid model is a model constructed based on a GNN model and an SVM model, which solves the problem of low accuracy in recognizing the speech emotions of users in the related art. By inputting the obtained speech information of the user into the hybrid model constructed based on the GNN model and the SVM model for emotion recognition processing, an emotion recognition result of the user can be obtained, thereby improving the accuracy of recognizing the speech emotions of the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0022] Figure 1 is a flowchart of the method for recognizing speech emotions provided by an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of the training of the speech emotion recognition model in an embodiment of the present application;
[0024] Figure 3 is a schematic diagram of the extraction of speech MFCC features and data preparation in an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of the classification representation of speech emotions in an embodiment of the present application;
[0026] Figure 5 is a schematic diagram of the system for recognizing speech emotions provided by an embodiment of the present application;
[0027] Figure 6 is a schematic diagram of the device for recognizing speech emotions provided by an embodiment of the present application;
[0028] Figure 7 is a schematic diagram of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The following will describe the present application in detail with reference to the drawings and in conjunction with the embodiments.
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set between the present system and relevant users or institutions. Before obtaining relevant information, a request for acquisition needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, the relevant information is obtained.
[0033] The following describes the present invention in conjunction with preferred implementation steps. Figure 1 is a flowchart of a method for recognizing speech emotions provided according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0034] Step S101, obtain the target speech information of the target object, where the target object is the object to be recognized for emotion.
[0035] For example, the above-mentioned target object may be a customer to be recognized for emotion, and it is necessary to recognize the emotion of this customer through a piece of speech of this customer. Therefore, first, a piece of speech of this customer is obtained.
[0036] Step S102: Input the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, where the target hybrid model is a model constructed based on the GNN model and the SVM model.
[0037] For example, input a segment of the customer's voice obtained into the hybrid model for emotion recognition processing, and obtain the emotion recognition result of this user, that is, obtain what the emotion of this user is. Moreover, the hybrid model for voice input is a model obtained by combining the GNN model and the SVM model.
[0038] Through the above steps S101 to S102, by inputting the obtained voice information of the user into the hybrid model constructed based on the GNN model and the SVM model for emotion recognition processing, the emotion recognition result of the user can be obtained, thereby improving the accuracy of the user's voice emotion recognition.
[0039] Optionally, in the voice emotion recognition method provided in the embodiments of the present application, before inputting the target voice information into the target hybrid model for emotion recognition processing, the method further includes: obtaining a target data set and obtaining a training set in the target data set; using the target data set to learn and train the GNN model to obtain a first recognition model; using the target data set to learn and train the SVM model to obtain a second recognition model; combining the first recognition model and the second recognition model according to preset requirements to obtain a first hybrid model; using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model.
[0040] Figure 2 It is a schematic diagram of the training of the voice emotion recognition model in the embodiments of the present application. As Figure 2 shown, according to the data set, perform deep learning model training (GNN model training) and traditional SVM model training to obtain the trained models, and combine the two models with a ratio of 0% to 100% and a step (step size) of 10% to obtain a hybrid model. Then use this hybrid model to perform regression verification on the training set in the data set, and then obtain the hybrid model with the highest accuracy, which is the voice emotion recognition model that needs to be used finally. In addition, the above preset requirements can be to combine the trained GNN model and the trained SVM model with a step (step size) of 10% in a ratio of 0% to 100%. For example, the proportion of the trained GNN model that conforms to the happy emotion is 100%, and the proportion of the trained SVM model that conforms to the happy emotion is 0%. If the two models are combined with a step (step size) of 10%, then adjust the proportion of the trained GNN model that conforms to the happy emotion to 90%, and the proportion of the trained SVM model that conforms to the happy emotion to 10%, and so on.
[0041] In summary, by combining a machine learning model with a traditional model, a model for identifying user emotions can be obtained quickly and accurately.
[0042] Optionally, in the method for identifying speech emotions provided in the embodiments of the present application, obtaining the training set in the target dataset includes: performing MFCC feature extraction operations on the data in the target dataset to obtain the data after feature extraction; obtaining the classification information of the preset speech emotions; adding labels to the data after feature extraction according to the classification information to obtain the data after adding labels; and dividing the data after adding labels into a training set and a test set according to a preset ratio.
[0043] Figure 3 is a schematic diagram of speech MFCC feature extraction and data preparation in the embodiments of the present application. As Figure 3 shown, the Chinese Emotion Corpus CISIA recorded by the Institute of Automation, Chinese Academy of Sciences is used as the dataset to perform MFCC feature extraction (Mel-Frequency Cepstral Coefficients, first transformed to Mel frequency, and then cepstral analysis), and after extraction, labels are added to each data according to the classification of speech emotions as shown in Figure 4 shown, and then the entire data is divided into a training set and a test set according to the ratio of 80% and 20%. In addition, the above-mentioned preset classification information of speech emotions can be the classification of speech emotions as shown in Figure 4 shown, and the speech emotions are divided into six basic discrete emotions, including happy, sad, angry, afraid, surprised, and disgusted. The above-mentioned preset ratio can be the ratio of 80% and 20%.
[0044] In summary, by performing processing such as extraction and adding labels to the data in the dataset, and according to the ratio, the processed data can be divided into a training set and a test set.
[0045] Optionally, in the method for identifying speech emotions provided in the embodiments of the present application, after using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model, the method further includes: using the test set to test the target hybrid model to obtain a test result; and determining the emotion recognition performance of the target hybrid model according to the test result.
[0046] For example, the test set in the dataset is used to test the above-mentioned hybrid model with the highest accuracy to evaluate the performance of the hybrid model.
[0047] In summary, by using the data in the test set, the performance of the hybrid model can be conveniently evaluated.
[0048] Optionally, in the method for recognizing speech emotion provided in the embodiments of the present application, obtaining the target data set includes: obtaining a Chinese emotion corpus; using the Chinese emotion corpus as the target data set.
[0049] In this embodiment, the Chinese emotion corpus CISIA recorded by the Institute of Automation, Chinese Academy of Sciences is used as the data set. And the CASIA Chinese emotion corpus is recorded by the Institute of Automation, Chinese Academy of Sciences, and includes a total of four professional speakers, six emotions: angry, happy, fear, sad, surprise, and neutral, with a total of 9,600 different pronunciations. Among them, 300 sentences are of the same text, that is, different emotions are assigned to the same text for reading. These corpora can be used to compare and analyze the acoustic and prosodic performances under different emotional states; in addition, 100 sentences are of different texts, and the emotional attribution of these texts can be seen from the literal meaning, which is convenient for the recorder to express emotions more accurately.
[0050] Through the above solution, using the Chinese emotion corpus CISIA as the data set can increase the number of data in the data set, thereby improving the accuracy of the subsequent hybrid model obtained through training.
[0051] Optionally, in the method for recognizing speech emotion provided in the embodiments of the present application, after inputting the target speech information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, the method further includes: recommending a target product to the target object according to the emotion recognition result; or; adopting a target type of response method for the target object according to the emotion recognition result.
[0052] Figure 5 is a schematic diagram of the speech emotion recognition system provided in the embodiments of the present application, as Figure 5 shown, the speech emotion recognition system is divided into a customer input speech module, a customer service answering module, and a speech emotion recognition module. Moreover, the above-obtained hybrid model with the highest accuracy is used as Figure 5The model in the voice emotion recognition module, and using this model, real-time emotion analysis can be performed on a piece of voice input by a customer, and it can be visualized to the agent interface to provide real-time emotion analysis of the user when the agent answers the user's voice consultation. Then when the agent is serving, appropriate responses and service methods are adopted according to the customer's current emotional state, such as happy, sad, angry, afraid, surprised, disgusted. For example, in the scenario of a bank's artificial customer service promoting business, if the customer's emotional state is disgusted, the promotion to the customer should be stopped in time. If the customer's emotional state is happy, the product introduction can continue to be promoted to the user.
[0053] Through the above solution, the emotion information in the voice information can be identified by artificial intelligence, and in the scenario of artificial customer service, the customer service staff can timely identify the current emotional state of the customer, so as to better serve and promote for different emotional states of the customer, and then the warm service level of voice consultation can be improved and the customer complaint rate can be reduced.
[0054] In summary, the voice emotion recognition method provided by the embodiment of the present application obtains the target voice information of the target object, where the target object is the object to be subjected to emotion recognition; the target voice information is input into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, where the target hybrid model is a model constructed based on the GNN model and the SVM model, which solves the problem of low accuracy of voice emotion recognition of users in the related art. By inputting the obtained voice information of the user into the hybrid model constructed based on the GNN model and the SVM model for emotion recognition processing, the emotion recognition result of the user can be obtained, thereby improving the accuracy of voice emotion recognition of the user.
[0055] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0056] The embodiment of the present application also provides a voice emotion recognition device. It should be noted that the voice emotion recognition device of the embodiment of the present application can be used to execute the voice emotion recognition method provided by the embodiment of the present application. The following introduces the voice emotion recognition device provided by the embodiment of the present application.
[0057] Figure 6 is a schematic diagram of the voice emotion recognition device according to the embodiment of the present application. As Figure 6 shown, the device includes: a first acquisition unit 601 and a first recognition unit 602.
[0058] Specifically, the first acquisition unit 601 is configured to acquire the target voice information of the target object, where the target object is the object to be subjected to emotion recognition;
[0059] The first recognition unit 602 is configured to input the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, where the target hybrid model is a model constructed based on the GNN model and the SVM model.
[0060] In summary, for the voice emotion recognition device provided in the embodiments of the present application, the first acquisition unit 601 acquires the target voice information of the target object, where the target object is the object to be subjected to emotion recognition; the first recognition unit 602 inputs the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, where the target hybrid model is a model constructed based on the GNN model and the SVM model, which solves the problem of low accuracy of user voice emotion recognition in the related art. By inputting the acquired user voice information into the hybrid model constructed based on the GNN model and the SVM model for emotion recognition processing, the emotion recognition result of the user can be obtained, thereby improving the accuracy of user voice emotion recognition.
[0061] Optionally, in the voice emotion recognition device provided in the embodiments of the present application, the device further includes: a second acquisition unit, configured to acquire a target data set and acquire a training set in the target data set before inputting the target voice information into the target hybrid model for emotion recognition processing; a first training unit, configured to use the target data set to learn and train the GNN model to obtain a first recognition model; a second training unit, configured to use the target data set to learn and train the SVM model to obtain a second recognition model; a first combination unit, configured to combine the first recognition model and the second recognition model according to preset requirements to obtain a first hybrid model; a first construction unit, configured to perform regression verification on the first hybrid model using the training set to obtain the target hybrid model.
[0062] Optionally, in the voice emotion recognition device provided in the embodiments of the present application, the second acquisition unit includes: a first extraction module, configured to perform MFCC feature extraction operation on the data in the target data set to obtain the data after feature extraction; a first acquisition module, configured to acquire the classification information of the preset voice emotion; a first addition module, configured to add labels to the data after feature extraction according to the classification information to obtain the data after adding labels; a first classification module, configured to divide the data after adding labels into a training set and a test set according to a preset ratio.
[0063] Optionally, in the voice emotion recognition device provided in the embodiments of the present application, the device further includes: a first testing unit, configured to test the target hybrid model with a test set after performing regression verification on the first hybrid model with a training set to obtain the target hybrid model, so as to obtain a test result; a first determination unit, configured to determine the emotion recognition performance of the target hybrid model according to the test result.
[0064] Optionally, in the voice emotion recognition device provided in the embodiments of the present application, the second acquisition unit includes: a second acquisition module, configured to acquire a Chinese emotion corpus; a first determination module, configured to use the Chinese emotion corpus as the target data set.
[0065] Optionally, in the voice emotion recognition device provided in the embodiments of the present application, the device further includes: a first recommendation unit, configured to recommend a target product to the target object according to the emotion recognition result after inputting the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object; or; a first response unit, configured to adopt a target type of response manner to the target object according to the emotion recognition result.
[0066] The voice emotion recognition device includes a processor and a memory. The above-mentioned first acquisition unit 601, first recognition unit 602, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.
[0067] The processor includes a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and the accuracy of the user's voice emotion recognition can be improved by adjusting the kernel parameters.
[0068] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0069] An embodiment of the present invention provides a processor, and the processor is used to run a program, wherein the program executes the voice emotion recognition method when running.
[0070] Such as Figure 7As shown in the figure, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining target voice information of a target object, where the target object is an object to be subjected to emotion recognition; inputting the target voice information into a target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, where the target hybrid model is a model constructed based on a GNN model and an SVM model.
[0071] When the processor executes the program, the following steps are further implemented: before inputting the target voice information into the target hybrid model for emotion recognition processing, the method further includes: obtaining a target data set and obtaining a training set in the target data set; using the target data set to learn and train the GNN model to obtain a first recognition model; using the target data set to learn and train the SVM model to obtain a second recognition model; combining the first recognition model and the second recognition model according to preset requirements to obtain a first hybrid model; using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model.
[0072] When the processor executes the program, the following steps are further implemented: obtaining the training set in the target data set includes: performing MFCC feature extraction operation on the data in the target data set to obtain the data after feature extraction; obtaining preset classification information of speech emotions; adding labels to the data after feature extraction according to the classification information to obtain the data after adding labels; dividing the data after adding labels into the training set and the test set according to a preset ratio.
[0073] When the processor executes the program, the following steps are further implemented: after using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model, the method further includes: using the test set to test the target hybrid model to obtain a test result; determining the emotion recognition performance of the target hybrid model according to the test result.
[0074] When the processor executes the program, the following steps are further implemented: obtaining the target data set includes: obtaining a Chinese emotion corpus; using the Chinese emotion corpus as the target data set.
[0075] When the processor executes the program, the following steps are further implemented: after inputting the target voice information into the target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, the method further includes: recommending a target product to the target object according to the emotion recognition result; or; adopting a target type of response method for the target object according to the emotion recognition result. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0076] The present application also provides a computer program product which, when executed on a data processing device, is adapted to execute a program initialized with the following method steps: obtaining target voice information of a target object, where the target object is an object to be subjected to emotion recognition; inputting the target voice information into a target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, where the target hybrid model is a model constructed based on a GNN model and an SVM model.
[0077] When executed on a data processing device, it is also adapted to execute a program initialized with the following method steps: before inputting the target voice information into the target hybrid model for emotion recognition processing, the method further includes: obtaining a target data set and obtaining a training set in the target data set; using the target data set to learn and train the GNN model to obtain a first recognition model; using the target data set to learn and train the SVM model to obtain a second recognition model; combining the first recognition model and the second recognition model according to a preset requirement to obtain a first hybrid model; using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model.
[0078] When executed on a data processing device, it is also adapted to execute a program initialized with the following method steps: obtaining the training set in the target data set includes: performing MFCC feature extraction operation on the data in the target data set to obtain the data after feature extraction; obtaining preset classification information of speech emotions; adding labels to the data after feature extraction according to the classification information to obtain the data after adding labels; dividing the data after adding labels into the training set and the test set according to a preset ratio.
[0079] When executed on a data processing device, it is also adapted to execute a program initialized with the following method steps: after using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model, the method further includes: using the test set to test the target hybrid model to obtain a test result; determining the emotion recognition performance of the target hybrid model according to the test result.
[0080] When executed on a data processing device, it is also adapted to execute a program initialized with the following method steps: obtaining the target data set includes: obtaining a Chinese emotion corpus; using the Chinese emotion corpus as the target data set.
[0081] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: after inputting the target voice information into a target hybrid model for emotion recognition processing to obtain an emotion recognition result of the target object, the method further includes: recommending a target product to the target object according to the emotion recognition result; or; adopting a target type of response manner for the target object according to the emotion recognition result.
[0082] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0083] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0084] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0086] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0087] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0088] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0089] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0090] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0091] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for recognizing speech emotions, characterized in that, Including: Obtain the target voice information of the target object, where the target object is the object to be subjected to emotion recognition; Input the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, where the target hybrid model is a model constructed based on the GNN model and the SVM model; Among them, the construction process of the target hybrid model includes: respectively training the GNN model and the SVM model according to the training set in the target data set; gradually adjusting the weights of the GNN model and the SVM model in proportion to form a hybrid model; using the training set for regression verification, and determining the hybrid model with the highest accuracy as the target hybrid model; Among them, the target data set includes recording samples with the same text but different emotional expressions, and text recordings with obvious emotional attributions.
2. The method according to claim 1, characterized in that, Before inputting the target voice information into the target hybrid model for emotion recognition processing, the method further includes: Obtain the target data set and obtain the training set in the target data set; Use the target data set to learn and train the GNN model to obtain the first recognition model; Use the target data set to learn and train the SVM model to obtain the second recognition model; Combine the first recognition model and the second recognition model according to preset requirements to obtain the first hybrid model; Use the training set to perform regression verification on the first hybrid model to obtain the target hybrid model.
3. The method according to claim 2, wherein Obtaining the training set in the target data set includes: Perform MFCC feature extraction operation on the data in the target data set to obtain the data after feature extraction; Obtain the classification information of the preset speech emotion; According to the classification information, add labels to the data after feature extraction to obtain the data after adding labels; Divide the data after adding labels into the training set and the test set according to a preset ratio.
4. The method according to claim 3, characterized in that After using the training set to perform regression verification on the first hybrid model to obtain the target hybrid model, the method further includes: Use the test set to test the target hybrid model to obtain the test result; According to the test result, determine the emotion recognition performance of the target hybrid model.
5. The method according to claim 2, wherein Obtaining the target data set includes: Obtain the Chinese emotion corpus; Use the Chinese emotion corpus as the target data set.
6. The method according to any one of claims 1 to 5, characterized in that, After inputting the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, the method further includes: Recommend a target product to the target object according to the emotion recognition result; or; Adopt a target type of response method for the target object according to the emotion recognition result.
7. An apparatus for recognizing voice emotions, characterized in that, Including: The first acquisition unit is used to obtain the target voice information of the target object, where the target object is the object to be subjected to emotion recognition; The first recognition unit is used to input the target voice information into the target hybrid model for emotion recognition processing to obtain the emotion recognition result of the target object, where the target hybrid model is a model constructed based on the GNN model and the SVM model; Among them, the recognition device is further configured to train the GNN model and the SVM model respectively according to the training set in the target data set; gradually adjust the weights of the GNN model and the SVM model according to a ratio to form a hybrid model; use the training set for regression verification, and determine the hybrid model with the highest accuracy as the target hybrid model; Among them, the target data set includes recording samples with the same text but different emotional expressions, as well as text recordings with obvious emotional attributions.
8. The device according to claim 7, characterized in that, The device further includes: A second acquisition unit, configured to acquire the target data set and the training set in the target data set before inputting the target voice information into the target hybrid model for emotion recognition processing; A first training unit, configured to perform learning and training on the GNN model by using the target data set to obtain a first recognition model; A second training unit, configured to perform learning and training on the SVM model by using the target data set to obtain a second recognition model; A first combination unit, configured to combine the first recognition model and the second recognition model according to preset requirements to obtain a first hybrid model; A first construction unit, configured to perform regression verification on the first hybrid model by using the training set to obtain the target hybrid model.
9. A processor, characterized in that, The processor is configured to run a program, wherein when the program runs, it executes the method for recognizing speech emotions according to any one of claims 1 to 6.
10. An electronic device, characterized in that, It includes one or more processors and a memory, and the memory is configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for recognizing speech emotions according to any one of claims 1 to 6.
Citation Information
Patent Citations
Speech emotion recognition method and system based on multi-grade classification of support vector machine
CN108899046A
Voice emotion recognition method and device, computer equipment and storage medium
CN112735479A