Voice emotion recognition model construction method and device
By encoding audio features and emotion tags on speech data, generating target encoding and training of speech emotion recognition models, the problem of information loss after speech data is converted into text is solved, and higher accuracy and accuracy of emotion recognition are achieved.
Patent Information
- Application Number
- CN202510606182.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the existing voice emotion recognition technology, key information is lost after voice data is converted into text, resulting in insufficient accuracy and high error rate of emotion recognition.
By encoding the audio characteristics and emotion labels of the speech data, audio encoding and emotion encoding are generated, and fused into target encoding, it is used to train the speech emotion recognition model and make full use of the context information in the speech data.
It improves the accuracy and accuracy of emotional recognition results, effectively solving the problem of insufficient accuracy in the existing technology.
Smart Images

Figure CN120452478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech emotion recognition, and in particular to a method and device for constructing a speech emotion recognition model. Background Art
[0002] Most existing speech emotion recognition methods first convert speech data into text data, then input the text data into a pre-trained emotion recognition model. The model then uses classification to determine and output the emotion label corresponding to the input text data, thereby obtaining the emotion corresponding to the speech data. However, this method has significant drawbacks. During the process of converting speech to text and then performing emotion recognition, a large amount of key information contained in the speech data, such as pitch, speaking speed, and tone, is lost. This leads to insufficient accuracy in emotion recognition and a high recognition error rate. Summary of the Invention
[0003] The present invention provides a method and device for constructing a speech emotion recognition model, which constructs a target speech emotion recognition model based on the audio coding and emotion coding of speech data, can fully utilize the contextual information in the speech data, and thus effectively improve the accuracy and precision of the emotion recognition results during the emotion recognition process.
[0004] According to one aspect of the present invention, a method for constructing a speech emotion recognition model is provided, the method comprising:
[0005] Get voice data;
[0006] Encoding the audio features of the speech data to obtain an audio code; and encoding the emotion tag corresponding to the speech data to obtain an emotion code;
[0007] Determining a target code based on the audio code and the emotion code;
[0008] The predetermined speech emotion recognition model to be trained is trained based on the target code to obtain a trained and updated target speech emotion recognition model.
[0009] According to another aspect of the present invention, a device for constructing a speech emotion recognition model is provided, the device comprising:
[0010] A voice data acquisition module, used to acquire voice data;
[0011] An encoding module, configured to encode the audio features of the speech data to obtain an audio code; and to encode the emotion tag corresponding to the speech data to obtain an emotion code;
[0012] A target coding determination module, configured to determine a target coding based on the audio coding and the emotion coding;
[0013] The target speech emotion recognition model obtaining module is used to train the predetermined speech emotion recognition model to be trained based on the target code to obtain the trained and updated target speech emotion recognition model.
[0014] The technical solution of the embodiment of the present invention obtains voice data and then encodes the audio features and corresponding emotion tags of the voice data to obtain audio codes and emotion codes. The audio codes and emotion codes are fused to generate target codes, and a predetermined voice emotion recognition model to be trained is trained based on the target codes to obtain a trained and updated target voice emotion recognition model. This technical solution constructs a target voice emotion recognition model based on the audio codes and emotion codes of the voice data, which can make full use of the contextual information in the voice data, thereby effectively improving the accuracy and precision of the emotion recognition results during the emotion recognition process. It effectively solves the problems of insufficient accuracy and high recognition error rate in existing voice emotion recognition technologies.
[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 This is a flowchart of a method for constructing a speech emotion recognition model according to the first embodiment of the present invention;
[0018] Figure 2 A schematic diagram of a speech emotion recognition model construction process provided in the second embodiment of the present invention;
[0019] Figure 3 A schematic diagram of another speech emotion recognition model construction process provided in Example 3 of the present invention;
[0020] Figure 4 A schematic diagram of another speech emotion recognition model construction process provided in the fourth embodiment of the present invention;
[0021] Figure 5 A flowchart of another method for constructing a speech emotion recognition model provided in Example 5 of the present invention;
[0022] Figure 6A flowchart of the application process of the target speech emotion recognition model provided in Example 6 of the present invention;
[0023] Figure 7 This is a structural diagram of a speech emotion recognition model construction device provided in Example 7 of the present invention. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] Example 1
[0027] Figure 1 This is a flow chart of a method for constructing a speech emotion recognition model according to the first embodiment of the present invention. This embodiment is applicable to situations where speech emotions are recognized. The method can be executed by a speech emotion recognition model construction device. The speech emotion recognition model construction device can be implemented in the form of hardware and / or software. The speech emotion recognition model construction device can be configured in a device. For example, the device can be a background server or other device with communication and computing capabilities. Figure 1 As shown, the method includes:
[0028] S110: Acquire voice data.
[0029] In this solution, voice data refers to the digital representation of various information in the form of voice.
[0030] In this embodiment, voice data can be obtained from a voice database; voice data can also be collected using a recording tool; and voice data can also be obtained from a cloud platform.
[0031] S120 , encoding the audio features of the speech data to obtain an audio code; and encoding the emotion tag corresponding to the speech data to obtain an emotion code.
[0032] Audio features refer to various parameters and attributes that describe the characteristics of an audio signal. For example, audio features can include amplitude, frequency, spectrum, formants, harmonics, and so on. Audio encoding is the process of converting audio features into a format more suitable for storage and transmission through specific algorithms.
[0033] In this embodiment, emotion tags are used to describe and define various emotional states. For example, emotion tags could be happiness, anger, sadness, fear, etc. Emotion encoding is the process of converting emotion tags through a specific algorithm into a format more suitable for storage and transmission. The number of emotion codes can be one or more.
[0034] In this solution, audio features are extracted from speech data, and then encoded to produce audio codes. For example, for 44KHz speech data, every 440 audio samples are extracted as an audio code, resulting in a 100Hz audio code, meaning 100 audio codes are obtained for every second of speech.
[0035] Audio coding can consist of a single or multi-layer structure. Multi-layer audio coding can capture and preserve richer and more detailed audio features by stacking different levels of coding.
[0036] Specifically, feature extraction algorithms can be used to extract features from speech data to obtain audio features. Alternatively, pre-trained feature extraction models can be used to extract features from speech data to obtain audio features. Feature extraction algorithms can include Mel-frequency cepstral coefficients and linear prediction cepstral coefficients. Feature extraction models can include machine learning models and deep learning models.
[0037] Furthermore, after obtaining the audio features, an encoder can be used to encode the audio features to obtain an audio code. Alternatively, a pre-trained encoding model can be used to encode the audio features to obtain an audio code. The encoder can be a MERT encoder or a Mel encoder. The encoding model can be a machine learning model, a deep learning model, or the like.
[0038] In this embodiment, the speech data can be processed based on a machine learning algorithm to determine the emotion label corresponding to the speech data; the speech data can also be processed based on a manual labeling method to determine the emotion label corresponding to the speech data.
[0039] Furthermore, after determining the emotion label corresponding to the speech data, an encoder can be used to encode the emotion label to obtain the emotion code. Alternatively, a pre-trained encoding model can be used to encode the emotion label to obtain the emotion code. The encoder can be a MERT encoder or a Mel encoder. The encoding model can be a machine learning model, a deep learning model, or the like.
[0040] S130: Determine a target code based on the audio code and the emotion code.
[0041] In this embodiment, the audio code and the emotion code may be concatenated to obtain the target code.
[0042] In this solution, the target code is a new code obtained by concatenating the audio code and the emotion code. This new code integrates audio and emotion information, and can more comprehensively describe the content and emotional characteristics of the speech data. The target code can be represented by a row vector or a matrix.
[0043] Optionally, the speech data includes speech data of at least one sentence of context information and / or context information and speech data of predicted emotions, and determining the target coding based on the audio coding and the emotion coding includes:
[0044] The audio codes corresponding to the speech data of the above and / or the below information, the audio codes corresponding to the speech data of the predicted emotion, and the emotion code are all spliced as fitting data to determine the target code; or,
[0045] The audio code corresponding to the speech data of the above and / or below information is used as prompt data, and the audio code corresponding to the speech data of the predicted emotion and the emotion code are spliced as fitting data to determine the target code.
[0046] In this embodiment, the voice data of the contextual information refers to the content before and after the conversation surrounding the target voice segment, which may cover multiple rounds of interaction information and is used to provide contextual support; the voice data for predicting emotions refers to the target voice segment to be analyzed.
[0047] The speech data of the context information and / or context information and the speech data of the predicted emotion may belong to the same character or may be generated by different characters.
[0048] In this solution, sentiment analysis of speech data during natural language processing often relies on contextual information. The same sentence, when placed in different contexts, can convey vastly different emotions. For example, if the speech data being predicted is "You're amazing," and the preceding context is "With your extraordinary wisdom and tireless efforts, you've successfully overcome numerous challenges and achieved groundbreaking results, which is truly admirable," then "You're amazing" would convey a strong sense of praise. However, if the preceding context were changed to "You frequently make mistakes at critical stages, causing huge losses to the team, yet you still resort to various excuses," "You're amazing" would create a sarcastic effect through the semantic contrast.
[0049] Specifically, the audio code corresponding to the speech data of the context information and / or context information, the audio code corresponding to the speech data of the predicted emotion, and the emotion code may be concatenated to generate the target code.
[0050] In this approach, the audio codes corresponding to the speech data of the preceding and / or following information can also be used as prompt data. This prompt data can provide critical context and guidance for the subsequent emotion recognition process, helping the model to more accurately capture the emotional characteristics of the speech. Based on this, the audio codes corresponding to the speech data for the predicted emotion are concatenated with the emotion code to generate the target code.
[0051] By integrating contextual information into speech data for emotion recognition, not only can the accuracy of emotion recognition be effectively improved, but also complex emotional expressions can be deeply analyzed, providing a more comprehensive and accurate judgment basis for sentiment analysis.
[0052] Optionally, the voice data containing at least one sentence of context information includes voice data of at least two characters, wherein the voice data of different characters are segmented by a third code.
[0053] In this embodiment, the third code can be represented by characters, numbers or complex data structures.
[0054] In this solution, the third encoding is a code specifically used to represent character-related information. When the voice data contains the voice data of at least two characters, the third encoding can be used to distinguish the voice data of different characters. For example, in a dialogue scene, multiple characters interact with each other. Character A says, "I encountered something really interesting today," and Character B responds, "Really? Tell me about it!" In this case, the third encoding can be used to deeply analyze the emotions expressed by Character B.
[0055] By segmenting the voice data of different characters through the third coding, the emotions of the characters can be presented in a quantifiable and analyzable way, providing strong support for subsequent in-depth research, analysis and decision-making.
[0056] S140 , training a predetermined speech emotion recognition model to be trained based on the target code to obtain a trained and updated target speech emotion recognition model.
[0057] Among them, the target speech emotion recognition model is an autoregressive language model.
[0058] In this solution, the target code is used as input to iteratively train the speech emotion recognition model to be trained. In this way, the model can effectively capture more contextual information and achieve in-depth mining of speech emotion features, thereby significantly improving the accuracy of emotion recognition and building a target speech emotion recognition model with better performance.
[0059] The technical solution of the embodiment of the present invention obtains speech data and then encodes the audio features and corresponding emotion tags of the speech data to obtain audio codes and emotion codes. The audio codes and emotion codes are fused to determine the target code, and a predetermined speech emotion recognition model to be trained is trained based on the target code to obtain a trained and updated target speech emotion recognition model. By implementing this technical solution, a target speech emotion recognition model is constructed based on the audio code and emotion code of the speech data, which can fully utilize the contextual information in the speech data, thereby effectively improving the accuracy and precision of the emotion recognition results during the emotion recognition process.
[0060] Example 2
[0061] Figure 2 This is a schematic diagram of a speech emotion recognition model construction process provided by the second embodiment of the present invention. The relationship between this embodiment and the above embodiment is a detailed description of the target coding determination process. Figure 2 As shown, the method includes:
[0062] S210: Acquire voice data.
[0063] S220 , encoding the audio features of the speech data to obtain an audio code; and encoding the emotion tag corresponding to the speech data to obtain an emotion code.
[0064] S230: splicing the audio code and the emotion code according to a preset splicing order to obtain a target code; wherein the preset splicing order includes splicing the emotion code before and / or after the audio code.
[0065] In this solution, the emotion code can be concatenated before the audio code to generate the target code. For example, if the emotion code is ABC and the audio code is DEF, the target code after concatenation can be ABCDEF or expressed in the form of [ABC, DEF].
[0066] In this embodiment, the emotion code can also be spliced after the audio code to generate a target code. For example, if the emotion code is ABC and the audio code is DEF, the target code after splicing can be DEFABC or expressed in the form of [DEF, ABC].
[0067] Splicing data according to a pre-set splicing order helps standardize and normalize data. This allows audio and emotion codes from different sources and formats to be combined in a unified manner, facilitating subsequent data storage, management, and transmission. It also helps ensure data consistency and stability during processing, reducing errors and deviations caused by inconsistent data formats.
[0068] S240 , training a predetermined speech emotion recognition model to be trained based on the target code to obtain a trained and updated target speech emotion recognition model.
[0069] The technical solution of the embodiment of the present invention obtains speech data and then encodes the audio features and corresponding emotion tags of the speech data to obtain audio codes and emotion codes. The audio codes and emotion codes are spliced together to determine the target code, and a predetermined speech emotion recognition model to be trained is trained based on the target code to obtain a trained and updated target speech emotion recognition model. By implementing this technical solution, a target speech emotion recognition model is constructed based on the audio code and emotion code of the speech data, which can fully utilize the contextual information in the speech data, thereby effectively improving the accuracy and precision of the emotion recognition results during the emotion recognition process.
[0070] Example 3
[0071] Figure 3 This is a schematic diagram of another speech emotion recognition model construction process provided by the third embodiment of the present invention. The relationship between this embodiment and the above embodiment is a detailed description of the target coding determination process. Figure 3 As shown, the method includes:
[0072] S310: Acquire voice data.
[0073] S320: Encode the audio features of the speech data to obtain an audio code; and encode the emotion tag corresponding to the speech data to obtain an emotion code.
[0074] S330. Perform layered processing on the audio code to obtain a first-layer audio code and a second-layer audio code; wherein the granularity of the first-layer audio code is greater than that of the second-layer audio code, or the number of the first-layer audio code is less than that of the second-layer audio code.
[0075] In this solution, the audio coding can be layered to obtain first-layer audio coding and second-layer audio coding of different granularities, wherein the granularity of the first-layer audio coding is greater than that of the second-layer audio coding.
[0076] Specifically, the audio code is mapped to the closest discretized first-layer audio code, residual information between the actual audio code and the first-layer audio code is calculated, and the residual information is compared with the discretized second-layer audio code to select the closest second-layer audio code.
[0077] The first-layer discretized audio code that is closest to the original audio code can be found based on a preset discretization rule. For example, a quantization interval can be set to divide the audio code into different discrete values according to its numerical range, and the closest discrete value can be selected as the corresponding first-layer discretized audio code.
[0078] In this embodiment, the residual information between the actual audio code and the first-layer audio code can be calculated using a time domain residual calculation method, or based on a machine learning model.
[0079] Furthermore, the similarity between the residual information and the discretized second-layer audio code can be analyzed to find the closest second-layer audio code. For example, the distance between the residual information and each discretized second-layer audio code can be calculated and used as a metric to quantify the degree of similarity. The smaller the distance value, the higher the similarity between the two. Ultimately, the second-layer audio code closest to the residual information is selected.
[0080] This layered coding strategy, by stacking the first and second layers of audio coding, can approximate the original audio features with greater accuracy. This significantly reduces encoding storage overhead and model training complexity. While traditional single-layer coding requires an N×M codebook to achieve the desired restoration, the layered coding framework only requires a two-layer codebook of N+M. Furthermore, as the number of coding layers increases, the improvements in compression efficiency and signal fidelity become increasingly significant.
[0081] In this solution, different numbers of first-layer audio codes and second-layer audio codes can be obtained by performing layered processing on the audio codes, wherein the number of the first-layer audio codes is smaller than that of the second-layer audio codes.
[0082] Specifically, the audio code is sampled at a preset sampling rate, a preset number of first-layer audio codes are extracted, and a preset number of second-layer audio codes are extracted from the preset number of first-layer audio codes. When extracting the audio code, the operation is performed in a bottom-up order, and the last extracted code is used as the first-layer audio code, which is a highly compressed audio code. Based on the first-layer audio code, the refinement operation is performed layer by layer. During the input process, each layer of audio code is input in sequence according to the hierarchical order, and the encoding output is also performed layer by layer during output. For example, the audio code is sampled by setting 100 sampling points, 1 first-layer audio code is extracted, and 1 second-layer audio code is extracted from the extracted 10 first-layer audio codes.
[0083] Through a bottom-up sampling approach, the last extracted code is used as the first layer of audio encoding, resulting in highly compressed results, reducing data volume and facilitating storage and transmission. Following the hierarchical order of operations can ensure the accuracy of the audio encoding to a certain extent. In audio analysis, precise audio encoding can provide more accurate audio feature information, which facilitates more accurate audio emotion recognition.
[0084] S340. Splice the first-layer audio code and the second-layer audio code with the emotion code respectively in a preset splicing order to obtain a target code; wherein the preset splicing order includes placing the emotion code before the first-layer audio code and the second-layer audio code for splicing; or placing the emotion code after the first-layer audio code and the second-layer audio code for splicing.
[0085] Specifically, the emotion coding may be placed before the first layer audio coding and the second layer audio coding to generate the target coding; or the emotion coding may be placed after the first layer audio coding and the second layer audio coding to generate the target coding.
[0086] In this solution, the audio coding is layered to obtain first-layer audio coding and second-layer audio coding of different granularities. During the splicing process, the first-layer audio coding and the second-layer audio coding can be superimposed, delayed, and spliced together.
[0087] In this embodiment, different numbers of first-layer audio codes and second-layer audio codes are obtained by layering the audio codes. During the splicing process, the first-layer audio codes and the second-layer audio codes can be superimposed or spliced layer by layer, that is, the 10 second-layer audio codes corresponding to one first-layer audio code can be superimposed or spliced; they can also be spliced by layer and combined in sequence into the target code, that is, the 10 first-layer audio codes are in front, and the corresponding 100 second-layer audio codes are in the back and combined in sequence into the target code.
[0088] S350: Training a predetermined speech emotion recognition model to be trained based on the target code to obtain a trained and updated target speech emotion recognition model.
[0089] The technical solution of the embodiment of the present invention obtains voice data and then encodes the audio features and corresponding emotion tags of the voice data to obtain audio codes and emotion codes. The audio codes are layered to obtain first-layer audio codes and second-layer audio codes, and the first-layer audio codes, the second-layer audio codes and the emotion codes are combined and spliced to determine the target code, and a predetermined speech emotion recognition model to be trained is trained based on the target code to obtain a trained and updated target speech emotion recognition model. By executing this technical solution, a target speech emotion recognition model is constructed based on the audio code and emotion code of the speech data, which can fully utilize the contextual information in the speech data, thereby effectively improving the accuracy and precision of the emotion recognition results during the emotion recognition process.
[0090] Example 4
[0091] Figure 4 This is a schematic diagram of another speech emotion recognition model construction process provided by the fourth embodiment of the present invention. The relationship between this embodiment and the above embodiment is a detailed description of the target coding determination process. Figure 4 As shown, the method includes:
[0092] S410: Acquire voice data.
[0093] S420: Encode the audio features of the speech data to obtain an audio code; and encode the emotion tag corresponding to the speech data to obtain an emotion code.
[0094] S430, adding a preset first code at the beginning position of the audio code and adding a preset second code at the end position of the audio code; wherein, the first code is used to identify the start of the prediction of the speech emotion recognition model to be trained; the second code is used to identify the end of the audio prediction of the speech emotion recognition model to be trained.
[0095] In this embodiment, the first code and the second code can be represented by characters, numbers, or complex data structures. The first code and the second code can be set based on the model training requirements. For example, the first code can be 0 and the second code can be 1.
[0096] In this embodiment, during the training process of the speech emotion recognition model, the first code is used to mark the time node when the speech emotion recognition model to be trained starts to make predictions, which means that the model will start the audio feature analysis and prediction process for the input audio code from that moment on. The second code is used to mark the end of the audio prediction of the speech emotion recognition model to be trained. It accurately indicates that the speech emotion recognition model to be trained has completed all the prediction tasks for the audio code, and can start the analysis and prediction of the emotional features of the input audio code. After outputting the corresponding predicted emotional code, it marks the official end of the model's emotional analysis process for the specific audio. Through the orderly coordination of these two codes, the prediction process of the model can be effectively standardized and managed to ensure the accuracy and efficiency of the training process of the speech emotion recognition model to be trained.
[0097] Specifically, the integrity marking of the audio code is achieved by embedding a preset first code at the start position of the audio code and adding a preset second code at the end position.
[0098] In this solution, the audio code can be modified directly to embed the first code and the second code into the starting position and the ending position of the audio code respectively; the digital watermark technology can also be used to embed the first code and the second code into the starting position and the ending position of the audio code respectively.
[0099] The first and second codes serve as clear start and end markers, allowing the speech emotion recognition model to be trained to accurately determine the start and end positions of the audio code. This prevents the model from mistakenly incorporating other irrelevant audio codes into the recognition range or incorrectly truncating the audio code, thereby improving the accuracy and stability of emotion recognition. Furthermore, the first code prompts the model to start prediction, allowing the model to quickly enter the working state and avoiding unnecessary calculations and waiting time. The second code allows the model to promptly analyze and process the audio code and predict the output emotion code, preventing the model from continuing to run ineffectively after the audio ends, thereby saving computing resources and improving overall recognition efficiency.
[0100] S440: Set the emotion code after the second code to construct a target code.
[0101] Specifically, the first code is added at the beginning of the audio code, the second code is added at the end, and the emotion code is set after the second code, so as to construct the target code.
[0102] S450: Training a predetermined speech emotion recognition model to be trained based on the target code to obtain a trained and updated target speech emotion recognition model.
[0103] The technical solution of the embodiment of the present invention obtains voice data and then encodes the audio features and corresponding emotion tags of the voice data to obtain audio code and emotion code. A preset first code is added at the beginning of the audio code and a preset second code is added at the end of the audio code, and the emotion code is set after the second code to construct a target code. The predetermined voice emotion recognition model to be trained is trained based on the target code to obtain a trained and updated target voice emotion recognition model. By executing this technical solution, a target voice emotion recognition model is constructed based on the audio code and emotion code of the voice data, which can fully utilize the contextual information in the voice data, thereby effectively improving the accuracy and precision of the emotion recognition results during the emotion recognition process.
[0104] Example 5
[0105] Figure 5 This is a flowchart of another method for constructing a speech emotion recognition model provided by the fifth embodiment of the present invention. The relationship between this embodiment and the above embodiments is a detailed description of the training process of the speech emotion recognition model to be trained. The speech emotion recognition model to be trained is an autoregressive language model, such as Figure 5 As shown, the method includes:
[0106] S510: Acquire voice data.
[0107] S520: Encode the audio features of the speech data to obtain an audio code; and encode the emotion tag corresponding to the speech data to obtain an emotion code.
[0108] S530: Determine a target code based on the audio code and the emotion code.
[0109] S540: Determine initial parameters of the speech emotion recognition model to be trained.
[0110] In this embodiment, the speech emotion recognition model to be trained is an autoregressive language model.
[0111] Specifically, the initial parameters of the speech emotion recognition model to be trained include network structure parameters and training hyperparameters. Network structure parameters include the number of model layers, the number of attention heads, and the embedding dimension; these parameters determine the model's feature extraction capabilities and information exchange efficiency. Training hyperparameters include key indicators such as the learning rate and the number of training rounds, which are used to control the model's training process and convergence speed. By rationally configuring these parameters, the model's performance in speech emotion recognition tasks can be effectively optimized.
[0112] S550: input the target code into the speech emotion recognition model to be trained, and output a predicted code, wherein the predicted code includes a predicted speech code and a predicted emotion code;
[0113] In this solution, the target code is input into the speech emotion recognition model to be trained, and the speech emotion recognition model to be trained performs in-depth analysis and processing on the target code, performs prediction operations, and finally outputs the predicted speech code and predicted emotion code.
[0114] S560: Determine a loss function value based on the predicted code and the target code;
[0115] The loss function is a function used to measure the difference between the model's predictions and the actual results. Loss functions can include mean squared error, mean absolute error, mean absolute percentage error, logarithmic loss function, etc.
[0116] Specifically, during the model training process, the loss function value is calculated using prediction coding and target coding. By calculating the loss function value, the difference between the model prediction coding and the target coding can be scientifically quantified, and then the prediction performance and generalization ability of the model in the current training state can be comprehensively and accurately evaluated.
[0117] Optionally, determining a loss function value based on the predicted encoding and the target encoding includes:
[0118] Determining a weight of the audio encoding and a weight of the emotion encoding;
[0119] According to the weight of the audio coding and the weight of the emotion coding, the loss function value of the audio coding and the loss function value of the emotion coding are weightedly combined to determine the loss function value.
[0120] The weight of the audio coding and the weight of the emotion coding may be set based on a machine learning algorithm; or the weight of the audio coding and the weight of the emotion coding may be set based on experience or domain knowledge.
[0121] In this embodiment, the sum of the weight of the audio coding and the weight of the emotion coding is equal to one.
[0122] In this solution, the loss function values for audio coding and emotion coding are weighted according to their respective weights to obtain the final loss function value. Specifically, the loss function value for audio coding is multiplied by the audio coding weight, the loss function value for emotion coding is multiplied by the emotion coding weight, and the two products are then added together.
[0123] This weighted combination approach comprehensively considers the impact of audio encoding and emotion encoding on the final result, thereby determining a comprehensive loss function value. This comprehensive loss function value can be used for subsequent model training and optimization.
[0124] In this embodiment, different weights can be assigned to audio coding and emotion coding when calculating the loss function value. Optionally, the weight of emotion coding can be set to be greater than the weight of audio coding, that is, a higher weight is assigned to emotion coding and a lower weight is assigned to audio coding; the weight of emotion coding can also be set to one and the weight of audio coding to zero. At this time, audio coding only serves as auxiliary information to help the model understand the contextual semantics, so that the model can output the predicted coding more accurately. However, audio coding will not participate in the parameter optimization link of the model. This strategy can effectively guide the training direction of the speech emotion recognition model to be trained, so that it can converge quickly and significantly improve the accuracy and efficiency of emotion prediction.
[0125] S570. Adjust the initial parameters of the speech emotion recognition model to be trained according to the loss function value to obtain a target speech emotion recognition model after training and update.
[0126] In this approach, the gradient of each initial parameter in the model is calculated using the backpropagation algorithm using the loss function value. The gradient represents the rate of change of the loss function under the current parameters and can indicate the direction and magnitude of parameter adjustment. Based on the calculated gradient, an appropriate optimization algorithm, such as stochastic gradient descent (SGD) or adaptive moment estimation (Adam), is then selected to update the initial parameters.
[0127] Furthermore, during the parameter update process, the entire target encoding needs to be iteratively trained multiple times. The above process is repeated for each iteration. Through continuous iterative training, the model parameters are gradually adjusted, so that the model's predicted encoding becomes increasingly close to the target encoding, and the loss function value gradually decreases. When the loss function value reaches a predetermined minimum value, or when the change in the loss function value after multiple iterations is no longer significant, the model is considered to have converged. The resulting model is the target speech emotion recognition model after training and update.
[0128] In this solution, the target encoding, including both audio and emotion encodings, is used as training data during the training of the speech emotion recognition model. The model predicts and fits the audio and emotion encodings in the training data. Unlike existing emotion classification models, the speech emotion recognition model not only learns the emotion classification information corresponding to the textual semantic output, but also learns richer information, such as audio features such as context, intonation, voice tone, stress, and speech rate. After comprehensively considering all audio features, the model outputs the corresponding emotion encoding, significantly improving the accuracy of emotion judgment.
[0129] The technical solution of the embodiment of the present invention is to encode the audio features of the speech data and the corresponding emotion labels respectively after the speech data is acquired, thereby obtaining audio coding and emotion coding. The audio coding and emotion coding are spliced to determine the target coding. Next, the initial parameters of the speech emotion recognition model to be trained are determined, the target coding is input into the speech emotion recognition model to be trained, and the model outputs the predicted coding. Based on the predicted coding and the target coding, the loss function value is calculated and determined, and the initial parameters of the speech emotion recognition model to be trained are adjusted according to the loss function value, and finally the target speech emotion recognition model after training and update is obtained. By executing this technical solution, the target speech emotion recognition model is constructed based on the audio coding and emotion coding of the speech data, which can make full use of the contextual information in the speech data, thereby effectively improving the accuracy and precision of the emotion recognition results in the emotion recognition process.
[0130] Example 6
[0131] Figure 6 The flowchart of the application process of the target speech emotion recognition model provided by the sixth embodiment of the present invention is a detailed description of the application process of the trained target speech emotion recognition model. Figure 6 As shown, the method includes:
[0132] S610: Acquire target voice data.
[0133] In this solution, target speech data refers to the digital representation of various information in the form of speech.
[0134] In this embodiment, the target voice data may be obtained from a voice database; the target voice data may also be collected using a recording tool; or the target voice data may be obtained from a cloud platform.
[0135] S620: Encode the target speech data and input it into the target speech emotion recognition model, and output the target emotion code.
[0136] In this solution, target audio features are obtained by extracting features from target speech data, and the target audio features are encoded to obtain target audio codes.
[0137] The target audio codec can be composed of a single-layer or multi-layer structure. Multi-layer audio codecs can capture and preserve richer and more detailed audio features by stacking different levels of codecs.
[0138] Specifically, a feature extraction algorithm can be used to extract features from the target speech data to obtain target audio features. Alternatively, a pre-trained feature extraction model can be used to extract features from the target speech data to obtain target audio features. Feature extraction algorithms can include Mel-frequency cepstral coefficients and linear prediction cepstral coefficients. Feature extraction models can include machine learning models and deep learning models.
[0139] Furthermore, after obtaining the target audio features, an encoder can be used to encode the target audio features to obtain the target audio code. Alternatively, a pre-trained encoding model can be used to encode the target audio features to obtain the target audio code. The encoder can include a MERT encoder and a Mel encoder. The encoding model can include a machine learning model, a deep learning model, and the like.
[0140] In this embodiment, after obtaining the target audio code, the target audio code is input into the target speech emotion recognition model, and the target speech emotion recognition model predicts the target audio code and outputs the target emotion code.
[0141] Specifically, when processing the input target speech data, the target speech emotion recognition model predicts and generates a complete predicted audio code and a predicted emotion code. However, when finally sampling and outputting the data, only the predicted emotion code is output. This approach allows the model to fully consider and utilize all contextual and audio information when predicting the emotion code.
[0142] S630: Determine a target emotion label corresponding to the target speech data based on the target emotion code.
[0143] In this solution, decoding technology can be used to parse the target emotion code and accurately output the emotion label corresponding to the target speech data.
[0144] The technical solution of an embodiment of the present invention obtains target speech data, encodes the target speech data, and then inputs it into the target speech emotion recognition model. The target emotion code is then output, and based on the target emotion code, a target emotion label corresponding to the target speech data is determined. By implementing this technical solution, emotion recognition is performed on the encoded speech data by constructing a target speech emotion recognition model, effectively improving the accuracy and precision of the recognition results.
[0145] Example 7
[0146] Figure 7 This is a structural diagram of a device for constructing a speech emotion recognition model provided by the seventh embodiment of the present invention. Figure 7 As shown, the device includes:
[0147] A voice data acquisition module 710 is used to acquire voice data;
[0148] The encoding module 720 is configured to encode the audio features of the speech data to obtain an audio code; and to encode the emotion tag corresponding to the speech data to obtain an emotion code;
[0149] A target coding determination module 730 is configured to determine a target coding based on the audio coding and the emotion coding;
[0150] The target speech emotion recognition model obtaining module 740 is used to train the predetermined speech emotion recognition model to be trained based on the target code to obtain the trained and updated target speech emotion recognition model.
[0151] Optionally, the target encoding determination module 730 is specifically configured to:
[0152] The audio code and the emotion code are spliced together in a preset splicing order to obtain a target code; wherein the preset splicing order includes splicing the emotion code before and / or after the audio code.
[0153] Optionally, the target encoding determination module 730 is further configured to:
[0154] performing layered processing on the audio codes to obtain first-layer audio codes and second-layer audio codes; wherein the granularity of the first-layer audio codes is greater than that of the second-layer audio codes, or the number of the first-layer audio codes is less than that of the second-layer audio codes;
[0155] The first-layer audio code and the second-layer audio code are respectively spliced with the emotion code in a preset splicing order to obtain a target code; wherein the preset splicing order includes placing the emotion code before the first-layer audio code and the second-layer audio code for splicing; or placing the emotion code after the first-layer audio code and the second-layer audio code for splicing.
[0156] Optionally, the target encoding determination module 730 is further configured to:
[0157] Adding a preset first code at the beginning of the audio code and adding a preset second code at the end of the audio code; wherein the first code is used to mark the start of the prediction of the speech emotion recognition model to be trained; and the second code is used to mark the end of the audio prediction of the speech emotion recognition model to be trained;
[0158] The emotion code is set after the second code to construct a target code.
[0159] Optionally, the speech emotion recognition model to be trained is an autoregressive language model, and the target speech emotion recognition model obtaining module 740 includes:
[0160] An initial parameter determination unit, used to determine the initial parameters of the speech emotion recognition model to be trained;
[0161] A prediction code output unit, configured to input the target code into the speech emotion recognition model to be trained and output a prediction code, wherein the prediction code includes a predicted speech code and a predicted emotion code;
[0162] a loss function value determining unit, configured to determine a loss function value based on the predicted encoding and the target encoding;
[0163] The target speech emotion recognition model obtaining unit is used to adjust the initial parameters of the speech emotion recognition model to be trained according to the loss function value to obtain the target speech emotion recognition model after training and updating.
[0164] Optional loss function value determination unit, specifically used for:
[0165] Determining a weight of the audio encoding and a weight of the emotion encoding;
[0166] According to the weight of the audio coding and the weight of the emotion coding, the loss function value of the audio coding and the loss function value of the emotion coding are weightedly combined to determine the loss function value.
[0167] Optionally, the speech data includes speech data of at least one sentence of context information and / or context information and speech data of predicted emotions. The target coding determination module 730 is further configured to:
[0168] The audio codes corresponding to the speech data of the above and / or the below information, the audio codes corresponding to the speech data of the predicted emotion, and the emotion code are all spliced as fitting data to determine the target code; or,
[0169] The audio code corresponding to the speech data of the above and / or below information is used as prompt data, and the audio code corresponding to the speech data of the predicted emotion and the emotion code are spliced as fitting data to determine the target code.
[0170] A speech emotion recognition model construction device provided by an embodiment of the present invention can execute a speech emotion recognition model construction method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.
[0171] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0172] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for constructing a speech emotion recognition model, characterized in that: include: Get voice data; Encoding the audio features of the speech data to obtain audio code; and encoding the emotion tag corresponding to the speech data to obtain an emotion code; Determining a target code based on the audio code and the emotion code; The predetermined speech emotion recognition model to be trained is trained based on the target code to obtain a trained and updated target speech emotion recognition model.
2. The method according to claim 1, characterized in that Determining a target code based on the audio code and the emotion code includes: The audio code and the emotion code are spliced together in a preset splicing order to obtain a target code; wherein the preset splicing order includes splicing the emotion code before and / or after the audio code.
3. The method according to claim 1, characterized in that Determining a target code based on the audio code and the emotion code includes: performing layered processing on the audio codes to obtain first-layer audio codes and second-layer audio codes; wherein the granularity of the first-layer audio codes is greater than that of the second-layer audio codes, or the number of the first-layer audio codes is less than that of the second-layer audio codes; The first-layer audio code and the second-layer audio code are respectively spliced with the emotion code in a preset splicing order to obtain a target code; wherein the preset splicing order includes placing the emotion code before the first-layer audio code and the second-layer audio code for splicing; or placing the emotion code after the first-layer audio code and the second-layer audio code for splicing.
4. The method according to claim 1, wherein Determining a target code based on the audio code and the emotion code includes: Adding a preset first code at the beginning of the audio code and adding a preset second code at the end of the audio code; wherein the first code is used to mark the start of the prediction of the speech emotion recognition model to be trained; and the second code is used to mark the end of the audio prediction of the speech emotion recognition model to be trained; The emotion code is set after the second code to construct a target code.
5. The method according to claim 1, wherein The speech emotion recognition model to be trained is an autoregressive language model, and the predetermined speech emotion recognition model to be trained is trained based on the target code to obtain a target speech emotion recognition model after training and updating, including: Determine the initial parameters of the speech emotion recognition model to be trained; Inputting the target code into the speech emotion recognition model to be trained, and outputting a predicted code, wherein the predicted code includes a predicted speech code and a predicted emotion code; Determining a loss function value based on the predicted code and the target code; The initial parameters of the speech emotion recognition model to be trained are adjusted according to the loss function value to obtain a target speech emotion recognition model after training and updating.
6. The method according to claim 5, characterized in that The determining of a loss function value based on the predicted encoding and the target encoding includes: Determining a weight of the audio encoding and a weight of the emotion encoding; According to the weight of the audio coding and the weight of the emotion coding, the loss function value of the audio coding and the loss function value of the emotion coding are weightedly combined to determine the loss function value.
7. The method according to claim 1, characterized in that The speech data includes speech data of at least one sentence of context information and / or context information and speech data of predicted emotions, and determining the target coding according to the audio coding and the emotion coding includes: The audio codes corresponding to the speech data of the above and / or the below information, the audio codes corresponding to the speech data of the predicted emotion, and the emotion code are all spliced as fitting data to determine the target code; or, The audio code corresponding to the speech data of the above and / or below information is used as prompt data, and the audio code corresponding to the speech data of the predicted emotion and the emotion code are spliced as fitting data to determine the target code.
8. The method according to claim 7, characterized in that The voice data containing at least one sentence of context information includes voice data of at least two characters, wherein the voice data of different characters are segmented by a third code.
9. The method according to any one of claims 1 to 8, characterized in that The application process of the target speech emotion recognition model includes: Obtain target voice data; Encoding the target speech data and inputting it into the target speech emotion recognition model to output the target emotion code; Based on the target emotion code, a target emotion label corresponding to the target speech data is determined.
10. A device for constructing a speech emotion recognition model, characterized in that: include: A voice data acquisition module, used to acquire voice data; An encoding module, configured to encode the audio features of the speech data to obtain an audio code; and encoding the emotion tag corresponding to the speech data to obtain an emotion code; A target coding determination module, configured to determine a target coding based on the audio coding and the emotion coding; The target speech emotion recognition model obtaining module is used to train the predetermined speech emotion recognition model to be trained based on the target code to obtain the trained and updated target speech emotion recognition model.
Citation Information
Patent Citations
Voice emotion recognition method and device, medium and electronic equipment
CN111429946A
Intent recognition method, device and equipment based on artificial intelligence and storage medium
CN113935333A
Speech synthesis method and device based on emotion migration, equipment and medium
CN116758894A
Service execution method and device based on voice emotion recognition
CN117316189A
Speech emotion recognition method and apparatus
WO2024008215A2