Emotion recognition method, device, equipment and medium for audio data

The audio emotion recognition method based on sound source separation and encoding processing solves the problem of poor recognition effect in polyphonic music, achieves higher generalization and robustness, and improves recognition accuracy.

CN116612788BActive Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310576648.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2025-09-26
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

Existing technologies have poor effects in polyphonic music emotion recognition, lack of generalization and robustness, and it is difficult to guarantee the accuracy of recognition results.

Method used

By using the sound modality extraction model to separate the sound source, separated audio of different sound modalities is obtained, and each separated audio is encoded using a mapping table and encoding model, and the emotion category is output in combination with the decoding and classification model.

Benefits of technology

It improves the generalization and robustness of audio emotion recognition, effectively retains the various features of audio data, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612788B_ABST
    Figure CN116612788B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of artificial intelligence technology, and in particular relates to a method, device, equipment and medium for emotion recognition of audio data. The method uses a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality, compares the sound modality of the separated audio with the sound modality in the first mapping table, determines the encoding model corresponding to the separated audio, calls the corresponding encoding model to encode the separated audio, connects all the encoding results, inputs the connection result into a decoding model, inputs the decoding result into a classification model, outputs the emotion category of the audio data, and through audio separation, uses different models to perform different encoding processing on the separated audio, and then integrates the separated audio to output the emotion classification result, which can realize the processing of complex audio, improve generalization ability and robustness, and effectively retain various features of the audio data, which helps to improve the recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application is applicable to the field of artificial intelligence technology, and in particular relates to a method, apparatus, device, and medium for emotion recognition from audio data. Background Art

[0002] Music emotion recognition (MER) currently refers to the task of identifying the emotional information contained in a given musical clip. Its input is the original audio file, and its output is the emotion category or valence / activation value. Initial MER methods mostly use artificial features (such as Mel-frequency cepstral coefficients, Mel-frequency spectra, and zero-crossing rate) combined with traditional machine learning models (such as support vector machines, hidden Markov models, and decision trees) for emotion classification. Most existing methods use short-time Fourier transform spectra (STFTs) or Mel-frequency spectra as input, and employ deep learning-based models such as convolutional neural networks for feature extraction. These methods then use fully connected classifiers for prediction, achieving good results.

[0003] However, existing methods work better for monophonic music but less so for polyphonic music. Furthermore, the recognition results are difficult to directly apply to emotion-based music retrieval tasks. Furthermore, most existing methods have high requirements for training data, requiring relatively high annotation accuracy and data quantity. Their generalization and robustness are poor, often exhibiting significant performance gaps on unfamiliar test data. This is particularly true for music with complex composition and arrangement, such as pop songs. Therefore, how to improve the generalization and robustness of audio emotion recognition while ensuring the accuracy of the recognition results has become an urgent issue. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a method, apparatus, device, and medium for emotion recognition in audio data to solve the problem of how to improve the generalization and robustness of audio emotion recognition while ensuring the accuracy of the recognition results.

[0005] In a first aspect, an embodiment of the present application provides a method for emotion recognition from audio data, the method comprising:

[0006] Using a sound modality extraction model, performing sound source separation on the acquired audio data to obtain separated audio of at least one sound modality;

[0007] For any separated audio, compare the sound mode of the separated audio with the sound modes in a first mapping table to determine a coding model corresponding to the separated audio, wherein the first mapping table stores a mapping relationship between sound modes and coding models;

[0008] Calling the encoding model corresponding to each separated audio to encode each separated audio, determining the encoding result of each separated audio, and concatenating all the encoding results to obtain a concatenated result;

[0009] The connection result is input into a decoding model for decoding to obtain a decoding result, and the decoding result is input into a classification model to output the emotion category of the audio data.

[0010] In one embodiment, calling a coding model corresponding to each separated audio to encode each separated audio, and determining the coding result of each separated audio includes:

[0011] Use Mel spectrum to transform each separated audio to obtain the transformation result of the corresponding separated audio;

[0012] For any separated audio, the transformation result of the separated audio is input into the encoding model, the encoding result is output, and all separated audios are traversed to obtain the encoding result of each separated audio.

[0013] In one embodiment, using a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality includes:

[0014] Using the sound modality extraction model, extract the human voice modality from the acquired audio data, and use the extraction result as the separated audio corresponding to the human voice modality;

[0015] canceling the audio data using the separated audio corresponding to the vocal modality to obtain audio data from which the vocals have been removed;

[0016] A sound modality extraction model is used to extract audio representing other sound modalities from the audio data from which human voices have been removed, thereby obtaining separated audio of other sound modalities.

[0017] In one embodiment, after performing sound source separation on the acquired audio data using a sound modality extraction model to obtain separated audio of at least one sound modality, the method further includes:

[0018] comparing the composition information of the at least one sound modality with combination information in a second mapping table to determine a weight distribution result corresponding to the combination information matching the composition information, wherein the second mapping table stores a mapping relationship between the combination information and the weight distribution results;

[0019] Connect all the encoding results to get the connection results including:

[0020] According to the weight distribution result, a weighted sum is performed on the encoding result of each separated audio, and the weighted sum result is determined as the connection result.

[0021] In one embodiment, all encoding results are concatenated to obtain a concatenated result comprising:

[0022] According to the importance of the sound mode corresponding to each separated audio, the encoding results corresponding to all the separated audios are connected end to end, and the result of the connection is determined as the connection result. Among them, the higher the importance of the sound mode, the higher the position of the encoding result of the corresponding separated audio in the connection result.

[0023] In one embodiment, all encoding results are concatenated to obtain a concatenated result comprising:

[0024] Input all encoding results into the trained full connector and output the connection results.

[0025] In one embodiment, after performing sound source separation on the acquired audio data using the sound modality extraction model to obtain separated audio of at least one sound modality, the method further includes:

[0026] Performing noise reduction processing on each separated audio to obtain separated audio after noise reduction processing;

[0027] Performing feature dimensionality reduction on the separated audio after the noise reduction processing to obtain separated audio with a dimension less than N;

[0028] For any separated audio, comparing the sound mode of the separated audio with the sound mode in the first mapping table to determine the coding model corresponding to the separated audio includes:

[0029] For any separated audio with a dimension less than N, the sound mode of the separated audio is compared with the sound mode in the first mapping table to determine the encoding model corresponding to the separated audio.

[0030] In a second aspect, an embodiment of the present application provides an emotion recognition device for audio data, the emotion recognition device comprising:

[0031] A sound source separation module is used to use a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality;

[0032] a coding model determination module configured to compare, for any separated audio, the sound mode of the separated audio with the sound modes in a first mapping table to determine a coding model corresponding to the separated audio, wherein the first mapping table stores a mapping relationship between sound modes and coding models;

[0033] The encoding connection module is used to call the encoding model corresponding to each separated audio to encode, determine the encoding result of each separated audio, and connect all the encoding results to obtain a connection result;

[0034] The emotion classification module is used to input the connection result into a decoding model for decoding to obtain a decoding result, input the decoding result into a classification model, and output the emotion category of the audio data.

[0035] In one embodiment, the encoding connection module includes:

[0036] A data conversion unit, configured to convert each separated audio using the Mel spectrum to obtain a conversion result corresponding to the separated audio;

[0037] The audio encoding unit is used to input the transformation result of any separated audio into the encoding model, output the encoding result, traverse all separated audios, and obtain the encoding result of each separated audio.

[0038] In one embodiment, the sound source separation module includes:

[0039] A first separation unit is configured to extract a human voice modality from the acquired audio data using a sound modality extraction model, and use the extraction result as separated audio corresponding to the human voice modality;

[0040] an elimination unit, configured to cancel the audio data using the separated audio corresponding to the vocal modality to obtain audio data with the vocal modality removed;

[0041] The second separation unit is configured to extract audio representing other sound modes from the audio data from which the human voice has been removed by using a sound mode extraction model, thereby obtaining separated audio of other sound modes.

[0042] In one embodiment, the emotion recognition device further includes:

[0043] a weight distribution determination module, configured to use a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality, and then compare the composition information of the at least one sound modality with the combination information in a second mapping table to determine a weight distribution result corresponding to the combination information that matches the composition information, wherein the second mapping table stores a mapping relationship between the combination information and the weight distribution result;

[0044] The coding connection module includes:

[0045] The first connecting unit is configured to perform weighted summation on the encoding results of each separated audio according to the weight distribution result, and determine the weighted summation result as the connecting result.

[0046] In one embodiment, the encoding connection module includes:

[0047] The second connection unit is used to connect the encoding results corresponding to all separated audios end to end according to the importance of the sound mode corresponding to each separated audio, and determine the result of the connection end to end as the connection result, wherein the higher the importance of the sound mode, the higher the position of the encoding result of the corresponding separated audio in the connection result.

[0048] In one embodiment, the encoding connection module includes:

[0049] The third connection unit is used to input all encoding results into the trained full connector and output the connection result.

[0050] In one embodiment, the emotion recognition device further includes:

[0051] A noise reduction processing module is used to perform sound source separation on the acquired audio data using a sound modality extraction model to obtain separated audio of at least one sound modality, and then perform noise reduction processing on each separated audio to obtain separated audio after noise reduction processing;

[0052] A dimensionality reduction processing module is used to perform feature dimensionality reduction on the separated audio after the noise reduction processing to obtain separated audio with a dimension less than N;

[0053] The coding model determination module includes:

[0054] The coding model determining unit is used to compare the sound mode of the separated audio with the sound mode in the first mapping table for any separated audio with a dimension less than N, and determine the coding model corresponding to the separated audio.

[0055] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the emotion recognition method as described in the first aspect when executing the computer program.

[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the emotion recognition method as described in the first aspect is implemented.

[0057] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: the present application uses a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality, and for any separated audio, compares the sound modality of the separated audio with the sound modality in the first mapping table to determine the encoding model corresponding to the separated audio, and the first mapping table stores the mapping relationship between the sound modality and the encoding model, and calls the encoding model corresponding to each separated audio for encoding, determines the encoding result of each separated audio, connects all the encoding results to obtain a connection result, inputs the connection result into the decoding model for decoding, obtains a decoding result, inputs the decoding result into the classification model, and outputs the emotion category of the audio data. Through audio separation, different models are used to perform different encoding processing on the separated audio, and then the separated audio is integrated to output the emotion classification result, which can realize the processing of complex audio, improve the generalization ability and robustness, and effectively retain the various features of the audio data, which helps to improve the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0059] Figure 1 This is a schematic diagram of an application environment of an audio data emotion recognition method provided in Example 1 of the present application;

[0060] Figure 2 This is a flow chart of a method for emotion recognition of audio data provided in Example 2 of the present application;

[0061] Figure 3 This is a flowchart of a method for emotion recognition of audio data provided in Example 3 of the present application;

[0062] Figure 4 This is a schematic diagram of the structure of an audio data emotion recognition device provided in Example 4 of the present application;

[0063] Figure 5 This is a structural diagram of a computer device provided in Example 5 of the present application. DETAILED DESCRIPTION

[0064] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0065] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0066] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0067] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0068] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0069] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0070] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0071] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0072] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0073] In order to illustrate the technical solution of the present application, specific embodiments are provided below.

[0074] The first embodiment of the present application provides an audio data emotion recognition method, which can be applied to Figure 1 In an application environment, a client communicates with a server. Clients include, but are not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, server computers, personal digital assistants (PDAs), and other computer devices. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0075] See also Figure 2 , is a flow chart of a method for emotion recognition of audio data provided in Example 2 of the present application, wherein the method for emotion recognition of audio data is applied to Figure 1 The server in the server, the computer device corresponding to the server connects to the corresponding database to obtain the audio data in the database. In addition, if you need to train any model in the server, you can also obtain the training set in the database. The above computer device can also be connected to the corresponding client, which is operated by the user. The user can send audio data to the server through the client. Figure 2 As shown, the emotion recognition method of audio data may include the following steps:

[0076] In step S201 , a sound modality extraction model is used to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality.

[0077] In this application, sound modality may refer to the modes corresponding to the sounds of different parts in the audio. The sounds of different parts may refer to the sounds produced by different musical instruments, different people, etc. For example, the sounds produced by instruments such as bass and piano are different sound modes. For another example, the sounds produced by people and the sounds produced by musical instruments are different sound modes.

[0078] The sound modal extraction model may refer to a model that can extract sounds representing different parts of audio. The model may be a neural network model, a machine learning model, or a deep learning model trained based on a corresponding training set. For example, the sound modal extraction model is a Demucs model, which can separate human voice from music, thereby obtaining human voice data and music data.

[0079] After the audio data is separated from the sound source, separated audio of different sound modes is obtained, and each separated audio corresponds to a sound mode, that is, the sound of a voice part.

[0080] Since the audio of each sound mode in the audio data has different contributions and influences on the emotions that the audio data will ultimately express, separating the various voice parts in advance and then inputting them into the feature extraction module can help the model better learn the unique feature representations of each part and achieve an effect similar to the channel-dimensional attention mechanism, enabling the model to know which voice part contributes the most and is the most important.

[0081] Optionally, using a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality includes:

[0082] Using the sound modality extraction model, extract the human voice modality from the acquired audio data, and use the extraction result as the separated audio corresponding to the human voice modality;

[0083] The audio data is cancelled using the separated audio corresponding to the human voice mode to obtain audio data without the human voice;

[0084] The sound modality extraction model is used to extract the audio representing other sound modalities from the audio data without the human voice, thereby obtaining the separated audio of other sound modalities.

[0085] In step S202 , for any separated audio, the sound mode of the separated audio is compared with the sound modes in the first mapping table to determine a coding model corresponding to the separated audio.

[0086] In this application, the first mapping table stores a mapping relationship between sound modalities and coding models, where each sound modality corresponds to a coding model. Because different sound modalities ultimately express different emotional contributions and influences, different coding models are required to encode the audio in order to obtain the true characteristics of the audio corresponding to the sound modality.

[0087] The sound modality can be represented as an identification number (Identity Document, ID). In the first mapping table, the ID of the sound modality corresponds to the number of the coding model. Therefore, according to the ID of the sound modality corresponding to the separated audio to be encoded, the number of the corresponding coding model can be determined, and then the corresponding coding model can be called according to the number of the coding model.

[0088] The coding model may be a coding model trained based on an existing training set, and may be capable of encoding sounds of different sound modes to improve the targeted encoding of sounds and realize feature extraction.

[0089] Step S203: Call the encoding model corresponding to each separated audio to encode each separated audio, determine the encoding result of each separated audio, and connect all the encoding results to obtain a connection result.

[0090] In this application, the encoding model is called to encode the corresponding separated audio, thereby obtaining the encoding result of each separated audio. The encoding result at this time is obtained based on different encoding models, and therefore has accurate sound characteristics of the corresponding separated audio, which is helpful for subsequent emotion recognition work.

[0091] Connecting all encoding results is to fuse all encoding results. The features represented by the fused results cover all features of the entire audio data, ensuring the integrity of the audio data.

[0092] Because different sound modalities have different weights in the overall audio data, their influence and contribution to emotion may vary. When fusing all encoding results, the importance of each encoding result in the fusion can be determined based on the contribution of each sound modality to emotion. If the corresponding sound models of all encoding results have the same contribution to emotion, then the importance of all encoding results in the fusion will also be the same.

[0093] The importance of the coding result in the fusion result can be reflected by the position of the coding result in the fusion result, the proportion of the coding result in the fusion result, etc., that is, the more important the coding result is, the more important its position in the fusion result or the higher its proportion.

[0094] Optionally, calling a coding model corresponding to each separated audio to encode each separated audio, and determining the encoding result of each separated audio includes:

[0095] Use Mel spectrum to transform each separated audio to obtain the transformation result of the corresponding separated audio;

[0096] For any separated audio, the transformation result of the separated audio is input into the encoding model, the encoding result is output, and all separated audios are traversed to obtain the encoding result of each separated audio.

[0097] Before encoding the separated audio, the separated audio may be transformed, and the transformation adopts Mel spectrum transformation to facilitate the processing of the encoding model and ensure the accuracy of the encoding of the encoding model.

[0098] Optionally, after performing sound source separation on the acquired audio data using a sound modality extraction model to obtain separated audio of at least one sound modality, the method further includes:

[0099] comparing the composition information of at least one sound modality with the combination information in a second mapping table to determine a weight distribution result corresponding to the combination information that matches the composition information, wherein the second mapping table stores a mapping relationship between the combination information and the weight distribution result;

[0100] Connect all the encoding results to get the connection results including:

[0101] According to the weight distribution result, the encoding result of each separated audio is weighted and summed, and the weighted sum result is determined as the connection result.

[0102] In this embodiment, a second mapping table is preset artificially, and a mapping relationship between combination information corresponding to combinations of different sound modes and weight distribution results is stored in the first mapping table.

[0103] Combination information may refer to a combination of different sound modes. For example, human voice and piano may constitute a combination, human voice and bass may constitute a combination, and human voice, piano and bass together constitute a combination. The weights corresponding to different sound modes in each combination are different. Therefore, a combination form must correspond to a weight distribution result.

[0104] The weight distribution result includes the weight value of each sound mode in the corresponding combination. Therefore, when connecting, a weighted summation method can be used to multiply the encoding result corresponding to the sound mode by the weight value and then add them to obtain the combined result.

[0105] Optionally, all encoding results are concatenated to obtain a concatenated result including:

[0106] According to the importance of the sound mode corresponding to each separated audio, the encoding results corresponding to all the separated audios are connected end to end, and the result of the connection is determined as the connection result. Among them, the higher the importance of the sound mode, the higher the position of the encoding result of the corresponding separated audio in the connection result.

[0107] In this embodiment, the encoding results are concatenated end-to-end, based on the importance of the sound modality. The higher the importance of the sound modality, the higher the position of the corresponding separated audio encoding result in the concatenation result. This method can express the degree of emotional impact of different sound modalities during decoding.

[0108] For example, for audio data of three sound modes, namely vocals, piano and bass, if the importance is vocals, piano and bass respectively, then the encoding result corresponding to vocals in the connection result is at the front, the encoding result corresponding to piano is in the middle, and the encoding result corresponding to bass is at the end.

[0109] Optionally, all encoding results are concatenated to obtain a concatenated result including:

[0110] Input all encoding results into the trained full connector and output the connection results.

[0111] Among them, in this embodiment, the encoded results are connected using a trained full connector, without the need to give a preset weight distribution result or the importance of the sound mode. The use of a trained full connector using a training set can avoid the influence of the above-mentioned human factors and make the connection more accurate.

[0112] Step S204: input the connection result into a decoding model for decoding to obtain a decoding result, input the decoding result into a classification model, and output the emotion category of the audio data.

[0113] In this application, for decoding the connection result, the decoding model used can be an independently trained decoding model, and does not need to be trained together with the above-mentioned encoding model, full connector, etc.

[0114] The classification model is a trained model. The classification model gives the classification results obtained by training using the training set, that is, each emotion category corresponds to at least one element. The decoding result and the element are used to calculate the similarity, and the emotion category to which the element with the highest similarity belongs is determined to be the emotion category of the decoding result, that is, the emotion category of the audio data, thereby realizing emotion recognition of the audio data.

[0115] For example, using the pre-trained sound source separation system Demucs, the sound source separation part does not require further training in actual application, and there are no requirements for training data. In terms of generalization ability, using the sound source separation task as an aid can enable each feature extraction module to process the corresponding part separately. Since the information distribution of the same part between different music is relatively similar, this method can reduce the performance degradation when facing a new type of test set, and show better generalization and adaptability between monophonic and polyphonic music. The improvements proposed in this application can also be applied to other existing unimodal music emotion recognition methods. Performing sound source separation preprocessing before the network can improve the performance and generalization ability of most existing methods. In addition, after using the sound source separation module, each extracted part uses its own encoding module. Since the sound characteristics of each part are quite different, using a separately customized model for feature extraction can better target the characteristics of the part. For example, the bass part is relatively simple as a whole and does not require a deep network for learning. This operation reduces the depth of the network and the requirements for computing power and training data, making the model more lightweight and capable of being deployed in a real-time music emotion recognition system; at the same time, it can also remove background noise and extract the core elements of music (such as melody, rhythm, etc.) in relatively complex arrangements. This capability enables the model to perform well in both monophonic and polyphonic music; when a part is missing (such as monophonic music or music with less than four parts), the method of the present application can still operate normally and perform well. It can be seen that the method of the present application can effectively process polyphonic music.

[0116] The embodiment of the present application uses a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality. For any separated audio, the sound modality of the separated audio is compared with the sound modality in the first mapping table to determine the encoding model corresponding to the separated audio. The first mapping table stores the mapping relationship between the sound modality and the encoding model. The encoding model corresponding to the separated audio is called for encoding each separated audio to determine the encoding result of each separated audio. All the encoding results are concatenated to obtain a concatenation result. The concatenation result is input into the decoding model for decoding to obtain a decoding result. The decoding result is input into the classification model to output the emotion category of the audio data. Through audio separation, different models are used to perform different encoding processing on the separated audio, and then the separated audio is integrated to output the emotion classification result. It can realize the processing of complex audio, improve the generalization ability and robustness, and effectively retain the various features of the audio data, which helps to improve the recognition accuracy.

[0117] See also Figure 3 , is a flow chart of a method for emotion recognition of audio data provided in Example 3 of the present application, such as Figure 3 As shown, the emotion recognition method includes:

[0118] Step S301 : Using a sound modality extraction model, the acquired audio data is subjected to sound source separation to obtain separated audio of at least one sound modality.

[0119] Among them, the content of step S301 is the same as that of the above-mentioned step S201. For details, please refer to the description of step S201, which will not be repeated here.

[0120] Step S302: performing noise reduction processing on each separated audio to obtain separated audio after noise reduction processing.

[0121] In this application, if the sound source cannot be subjected to noise reduction, decoupling, and feature dimensionality reduction processing during the above-mentioned sound source separation, the above-mentioned processing is still required to improve the encoding accuracy of the encoding model, which can further reduce the training difficulty of the encoding model, thereby achieving the purpose of reducing the demand for training data.

[0122] Step S303: Perform feature dimensionality reduction on the separated audio after the noise reduction process to obtain separated audio with a dimension less than N.

[0123] In this application, N is an integer greater than 1. Feature dimensionality reduction of the separated audio after noise reduction processing can effectively reduce the training difficulty of the encoding model. Among them, the lowest dimension of the separated audio can reach one dimension. At this time, the training difficulty is greatly reduced. Of course, it may be accompanied by a decrease in accuracy. Therefore, the value of N can be freely selected according to the requirements of accuracy and training difficulty during use.

[0124] Step S304 : For any separated audio with a dimension smaller than N, the sound mode of the separated audio is compared with the sound mode in the first mapping table to determine a coding model corresponding to the separated audio.

[0125] Step S305 , calling the encoding model corresponding to each separated audio to encode each separated audio, determining the encoding result of each separated audio, and connecting all the encoding results to obtain a connection result.

[0126] Step S306: input the connection result into the decoding model for decoding to obtain a decoding result, input the decoding result into the classification model, and output the emotion category of the audio data.

[0127] Among them, the contents of steps S304 to S306 are the same as those of the above-mentioned steps S202 to S204. For details, please refer to the description of steps S202 to S204, which will not be repeated here.

[0128] The embodiment of the present application uses a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality, performs noise reduction processing on each separated audio to obtain separated audio after noise reduction processing, performs feature dimensionality reduction on the separated audio after noise reduction processing to obtain separated audio with a dimension less than N, and for any separated audio with a dimension less than N, compares the sound modality of the separated audio with the sound modality in the first mapping table to determine the encoding model corresponding to the separated audio, the first mapping table stores the mapping relationship between the sound modality and the encoding model, calls the encoding model corresponding to the separated audio for encoding each separated audio, determines the encoding result of each separated audio, concatenates all the encoding results to obtain a concatenation result, inputs the concatenation result into a decoding model for decoding to obtain a decoding result, inputs the decoding result into a classification model, and outputs the emotion category of the audio data. Through audio separation, different models are used to perform different encoding processing on the separated audio, and then the separated audio is integrated to output the emotion classification result. This can realize the processing of complex audio, improve generalization ability and robustness, and effectively retain various features of the audio data, which helps to improve recognition accuracy.

[0129] Corresponding to the emotion recognition method of audio data in the above embodiment, Figure 4 The structure block diagram of the emotion recognition device of the emotion prediction model provided in the fourth embodiment of the present application is shown. The emotion recognition device is applied to Figure 1 The server in the embodiment of the present invention is connected to a corresponding database, and the computer device corresponding to the server is connected to the corresponding database to obtain audio data from the database. In addition, if any model in the server needs to be trained, the training set in the database can also be obtained. The above-mentioned computer device can also be connected to a corresponding client, which is operated by the user and can send audio data to the server through the client. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0130] See also Figure 4 , the emotion recognition device comprises:

[0131] The sound source separation module 41 is used to use the sound mode extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound mode;

[0132] a coding model determination module 42 for comparing the sound mode of any separated audio with the sound modes in a first mapping table to determine a coding model corresponding to the separated audio, wherein the first mapping table stores a mapping relationship between sound modes and coding models;

[0133] The encoding connection module 43 is used to call the encoding model corresponding to each separated audio to encode, determine the encoding result of each separated audio, and connect all the encoding results to obtain a connection result;

[0134] The emotion classification module 44 is used to input the connection result into the decoding model for decoding to obtain a decoding result, input the decoding result into the classification model, and output the emotion category of the audio data.

[0135] Optionally, the encoding connection module 43 includes:

[0136] A data conversion unit, configured to convert each separated audio using the Mel spectrum to obtain a conversion result corresponding to the separated audio;

[0137] The audio encoding unit is used to input the transformation result of any separated audio into the encoding model, output the encoding result, traverse all separated audios, and obtain the encoding result of each separated audio.

[0138] Optionally, the sound source separation module 41 includes:

[0139] A first separation unit is configured to extract a human voice modality from the acquired audio data using a sound modality extraction model, and use the extraction result as separated audio corresponding to the human voice modality;

[0140] an elimination unit, configured to cancel the audio data using the separated audio corresponding to the human voice mode to obtain audio data without the human voice;

[0141] The second separation unit is used to extract audio representing other sound modes from the audio data without human voice using a sound mode extraction model to obtain separated audio of other sound modes.

[0142] Optionally, the emotion recognition device further includes:

[0143] a weight distribution determination module, configured to use a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality, compare the composition information of the at least one sound modality with the combination information in a second mapping table, and determine a weight distribution result corresponding to the combination information that matches the composition information, wherein the second mapping table stores a mapping relationship between the combination information and the weight distribution result;

[0144] The encoding connection module 43 includes:

[0145] The first connection unit is configured to perform weighted summation on the encoding results of each separated audio according to the weight distribution result, and determine the weighted summation result as the connection result.

[0146] Optionally, the encoding connection module 43 includes:

[0147] The second connection unit is used to connect the encoding results corresponding to all separated audios end to end according to the importance of the sound mode corresponding to each separated audio, and determine the result of the connection end to end as the connection result, wherein the higher the importance of the sound mode, the higher the position of the encoding result of the corresponding separated audio in the connection result.

[0148] Optionally, the encoding connection module 43 includes:

[0149] The third connection unit is used to input all encoding results into the trained full connector and output the connection result.

[0150] Optionally, the emotion recognition device further includes:

[0151] A noise reduction processing module is used to perform sound source separation on the acquired audio data using a sound modality extraction model to obtain separated audio of at least one sound modality, and then perform noise reduction processing on each separated audio to obtain separated audio after noise reduction processing;

[0152] A dimensionality reduction processing module is used to perform feature dimensionality reduction on the separated audio after noise reduction processing to obtain separated audio with a dimension less than N;

[0153] The coding model determination module 42 includes:

[0154] The coding model determining unit is used to compare the sound mode of the separated audio with the sound mode in the first mapping table for any separated audio with a dimension less than N, and determine the coding model corresponding to the separated audio.

[0155] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0156] Figure 5 This is a schematic diagram of the structure of a computer device provided in Example 5 of this application. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned embodiments of the method for emotion recognition of audio data are implemented.

[0157] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 5The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.

[0158] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.

[0159] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.

[0160] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0161] The present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing it.

[0162] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0163] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0164] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which can be electrical, mechanical or other forms.

[0165] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0166] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for emotion recognition of audio data, characterized in that: The emotion recognition method comprises: Using a sound modality extraction model, performing sound source separation on the acquired audio data to obtain separated audio of at least one sound modality; For any separated audio, compare the sound mode of the separated audio with the sound modes in a first mapping table to determine a coding model corresponding to the separated audio, wherein the first mapping table stores a mapping relationship between sound modes and coding models; Calling the encoding model corresponding to each separated audio to encode each separated audio, determining the encoding result of each separated audio, and concatenating all the encoding results to obtain a concatenated result; The connection result is input into a decoding model for decoding to obtain a decoding result, and the decoding result is input into a classification model to output the emotion category of the audio data.

2. The emotion recognition method according to claim 1, characterized in that The encoding model corresponding to each separated audio is called for encoding, and the encoding result of each separated audio is determined to include: Use Mel spectrum to transform each separated audio to obtain the transformation result of the corresponding separated audio; For any separated audio, the transformation result of the separated audio is input into the encoding model, the encoding result is output, and all separated audios are traversed to obtain the encoding result of each separated audio.

3. The emotion recognition method according to claim 1, characterized in that Using a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality includes: Using the sound modality extraction model, extract the human voice modality from the acquired audio data, and use the extraction result as the separated audio corresponding to the human voice modality; canceling the audio data using the separated audio corresponding to the vocal modality to obtain audio data from which the vocals have been removed; A sound modality extraction model is used to extract audio representing other sound modalities from the audio data from which human voices have been removed, thereby obtaining separated audio of other sound modalities.

4. The emotion recognition method according to claim 1, characterized in that After performing sound source separation on the acquired audio data using a sound modality extraction model to obtain separated audio of at least one sound modality, the method further includes: comparing the composition information of the at least one sound modality with combination information in a second mapping table to determine a weight distribution result corresponding to the combination information matching the composition information, wherein the second mapping table stores a mapping relationship between the combination information and the weight distribution results; Connect all the encoding results to get the connection results including: According to the weight distribution result, a weighted sum is performed on the encoding result of each separated audio, and the weighted sum result is determined as the connection result.

5. The emotion recognition method according to claim 1, characterized in that Connect all the encoding results to get the connection results including: According to the importance of the sound mode corresponding to each separated audio, the encoding results corresponding to all the separated audios are connected end to end, and the result of the connection is determined as the connection result. Among them, the higher the importance of the sound mode, the higher the position of the encoding result of the corresponding separated audio in the connection result.

6. The emotion recognition method according to claim 1, characterized in that Connect all the encoding results to get the connection results including: Input all encoding results into the trained full connector and output the connection results.

7. The emotion recognition method according to any one of claims 1 to 6, characterized in that: After performing sound source separation on the acquired audio data using the sound modality extraction model to obtain separated audio of at least one sound modality, the method further includes: Performing noise reduction processing on each separated audio to obtain separated audio after noise reduction processing; Performing feature dimensionality reduction on the separated audio after the noise reduction processing to obtain separated audio with a dimension less than N; For any separated audio, comparing the sound mode of the separated audio with the sound mode in the first mapping table to determine the coding model corresponding to the separated audio includes: For any separated audio with a dimension less than N, the sound mode of the separated audio is compared with the sound mode in the first mapping table to determine the encoding model corresponding to the separated audio.

8. An emotion recognition device for audio data, characterized in that: The emotion recognition device comprises: A sound source separation module is used to use a sound modality extraction model to perform sound source separation on the acquired audio data to obtain separated audio of at least one sound modality; a coding model determination module configured to compare, for any separated audio, the sound mode of the separated audio with the sound modes in a first mapping table to determine a coding model corresponding to the separated audio, wherein the first mapping table stores a mapping relationship between sound modes and coding models; The encoding connection module is used to call the encoding model corresponding to each separated audio to encode, determine the encoding result of each separated audio, and connect all the encoding results to obtain a connection result; The emotion classification module is used to input the connection result into a decoding model for decoding to obtain a decoding result, input the decoding result into a classification model, and output the emotion category of the audio data.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the emotion recognition method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the emotion recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Music emotion recognition method and device, storage medium and electronic equipment

    CN111858943A

  • Emotion recognition method and device, electronic equipment and storage medium

    CN111968679A