Methods, devices, electronic equipment, and storage media for recognizing voice emotions

By employing a neural network algorithm optimized by dynamic population evolution, a generative adversarial network based on dynamic symmetry breaking theory, and a random forest classifier with quantum entropy encoding, the problems of difficult data acquisition, vanishing gradient, and insufficient feature dimensionality reduction in speech emotion recognition are solved, achieving high-precision emotion recognition.

CN119832939BActive Publication Date: 2025-10-28CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411834783.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-10-28
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Existing speech emotion recognition technologies suffer from difficulties in acquiring emotional speech data, limited sample sizes, gradient vanishing and gradient exploding problems during feature extraction, difficulty in retaining important information by feature dimensionality reduction techniques, and difficulty in capturing the inherent diversity and uncertainty of data by traditional classifiers, resulting in insufficient recognition accuracy.

Method used

Feature extraction is performed using a neural network algorithm based on dynamic population evolution optimization, data augmentation is performed using a generative adversarial network based on dynamic symmetry breaking theory, feature dimensionality reduction is performed using an autoencoder neural network based on feature refinement, and classification is performed using a random forest classifier based on quantum entropy encoding to form an emotion recognition model.

Benefits of technology

It improves the accuracy and stability of speech emotion recognition, enhances the model's generalization ability and its ability to handle complex environments, and significantly improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832939B_ABST
    Figure CN119832939B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, electronic device, and storage medium for recognizing speech emotions, relating to the field of data processing technology. The method includes: acquiring target speech data and an emotion recognition model for speech emotion recognition, the emotion recognition model including at least a feature extraction model, a feature dimensionality reduction model, and a classifier; inputting the target speech data into the feature extraction model for feature extraction to obtain speech features corresponding to the target speech data; inputting the speech features into the feature dimensionality reduction model for feature dimensionality reduction to obtain low-dimensional features corresponding to the speech features; and inputting the low-dimensional features into the classifier for prediction to obtain the emotion category of the speech data. Thus, by performing feature extraction, feature dimensionality reduction, and classification on the speech data, the accuracy of speech emotion recognition is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for recognizing voice emotions, a device for recognizing voice emotions, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With the development of human-computer interaction technology, emotion recognition has become increasingly important in various intelligent systems. Voice emotion recognition, as a crucial means of recognizing and understanding human emotions, is widely used in customer service, emotion monitoring, intelligent assistants, and interactive entertainment. However, existing voice emotion recognition technologies face numerous challenges in practical applications. First, acquiring emotional voice data is difficult, and the limited number of samples makes it difficult for models to possess sufficient generalization ability. Second, during feature extraction, conventional neural network algorithms are prone to problems such as vanishing and exploding gradients, affecting the effective extraction of features and the stability of the model. Third, feature dimensionality reduction techniques struggle to retain important information while reducing data dimensionality, leading to a decrease in subsequent classification accuracy. Finally, traditional classifier algorithms struggle to effectively capture and utilize the inherent diversity and uncertainty of complex emotional data, resulting in insufficient accuracy in emotion recognition. Summary of the Invention

[0003] The present invention provides a method, apparatus, electronic device, and computer-readable storage medium for recognizing voice emotions, in order to solve or partially solve the problem of inaccurate voice emotion recognition.

[0004] This invention discloses a method for recognizing speech emotions, including:

[0005] The system acquires target speech data, an emotion recognition model for speech emotion recognition, and a data augmentation model for data augmentation. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The data augmentation model is a model trained based on dynamic symmetry breaking theory. The feature extraction model is a model trained based on a neural network algorithm with dynamic population evolution optimization. The feature dimensionality reduction model is a model trained based on a feature refinement model. The classifier is a classifier trained based on a random forest classifier with quantum entropy encoding.

[0006] The target speech data is input into the feature extraction model for feature extraction to obtain the speech features corresponding to the target speech data;

[0007] The speech features are input into the feature dimensionality reduction model to perform feature dimensionality reduction, thereby obtaining the low-dimensional features corresponding to the speech features.

[0008] The low-dimensional features are input into the classifier for prediction to obtain the emotion category of the speech data.

[0009] This invention also discloses a voice emotion recognition device, comprising:

[0010] The data acquisition module is used to acquire target speech data, as well as an emotion recognition model for speech emotion recognition and a data augmentation model for data augmentation. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The data augmentation model is a model trained based on dynamic symmetry breaking theory. The feature extraction model is a model trained based on a neural network algorithm with dynamic population evolution optimization. The feature dimensionality reduction model is a model trained based on a feature refinement-based feature dimensionality reduction model. The classifier is a classifier trained based on a random forest classifier with quantum entropy encoding.

[0011] The feature extraction module is used to input the target speech data into the feature extraction model to extract features and obtain the speech features corresponding to the target speech data.

[0012] The feature dimensionality reduction module is used to input the speech features into the feature dimensionality reduction model to perform feature dimensionality reduction and obtain the low-dimensional features corresponding to the speech features.

[0013] The prediction module is used to input the low-dimensional features into the classifier for prediction to obtain the emotion category of the speech data.

[0014] This invention also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0015] The memory is used to store computer programs;

[0016] When the processor executes a program stored in the memory, it implements the method described in the embodiments of the present invention.

[0017] This invention also discloses a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the processors to perform the methods described in this invention.

[0018] The embodiments of the present invention have the following advantages:

[0019] In this embodiment of the invention, during the process of speech emotion recognition, an emotion recognition model for speech emotion recognition is obtained. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The feature extraction model is a model trained based on a neural network algorithm with dynamic population evolution optimization. The feature dimensionality reduction model is a model trained based on a feature refinement model. The classifier is a classifier trained based on a random forest classifier with quantum entropy encoding. First, the target speech data is input into the feature extraction model for feature extraction to obtain the speech features corresponding to the target speech data. Then, the speech features are input into the feature dimensionality reduction model for feature reduction to obtain the low-dimensional features corresponding to the speech features. Finally, the low-dimensional features are input into the classifier for prediction to obtain the emotion category of the speech data. Thus, by performing feature extraction, feature dimensionality reduction, and classification on the speech data, the accuracy of speech emotion recognition is effectively improved. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the steps of a voice emotion recognition method provided in an embodiment of the present invention;

[0021] Figure 2 This is a schematic diagram of the training process of the neural network algorithm provided in this embodiment of the invention;

[0022] Figure 3 This is a schematic diagram of the training process of the autoencoder neural network algorithm provided in this embodiment of the invention;

[0023] Figure 4 This is a structural block diagram of a voice emotion recognition device provided in an embodiment of the present invention. Detailed Implementation

[0024] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] As an example, in speech emotion recognition, the following challenges arise: First, acquiring emotional speech data is difficult, and the limited number of samples hinders the model's generalization ability. Second, during feature extraction, conventional neural network algorithms are prone to vanishing and exploding gradients, affecting the effective extraction of features and the model's stability. Third, feature dimensionality reduction techniques struggle to retain important information while reducing data dimensionality, leading to decreased classification accuracy. Finally, traditional classifier algorithms struggle to effectively capture and utilize the inherent diversity and uncertainty of complex emotional data, resulting in insufficient accuracy in emotion recognition.

[0026] In this invention, the emotion recognition model is optimized. By optimizing the training process, the trained emotion recognition model can accurately identify the emotion category corresponding to the speech data. Specifically, during the recognition process, an emotion recognition model for speech emotion recognition is obtained. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The feature extraction model is trained using a neural network algorithm based on dynamic population evolution optimization. The feature dimensionality reduction model is trained using a feature refinement-based feature dimensionality reduction model. The classifier is trained using a random forest classifier based on quantum entropy encoding. First, the target speech data is input into the feature extraction model for feature extraction to obtain the speech features corresponding to the target speech data. Then, the speech features are input into the feature dimensionality reduction model for feature reduction to obtain the low-dimensional features corresponding to the speech features. Finally, the low-dimensional features are input into the classifier for prediction to obtain the emotion category of the speech data. Thus, by performing feature extraction, feature dimensionality reduction, and classification on the speech data, the accuracy of speech emotion recognition is effectively improved.

[0027] Reference Figure 1 The diagram illustrates a flowchart of a speech emotion recognition method provided in an embodiment of the present invention, which may specifically include the following steps:

[0028] Step 101: Obtain target speech data and an emotion recognition model for speech emotion recognition. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The feature extraction model is a model trained using a neural network algorithm based on dynamic population evolution optimization. The feature dimensionality reduction model is a model trained using a feature dimensionality reduction model based on feature refinement. The classifier is a classifier trained using a random forest classifier based on quantum entropy encoding.

[0029] Step 102: Input the target speech data into the feature extraction model to extract features and obtain the speech features corresponding to the target speech data;

[0030] Step 103: Input the speech features into the feature dimensionality reduction model to perform feature dimensionality reduction and obtain the low-dimensional features corresponding to the speech features;

[0031] Step 104: Input the low-dimensional features into the classifier for prediction to obtain the emotion category of the speech data.

[0032] In this embodiment of the invention, the emotion recognition model may include at least a feature extraction model, a feature dimensionality reduction model, and a classifier, wherein the feature extraction model is a model trained by a neural network algorithm based on dynamic population evolution optimization, the feature dimensionality reduction model is a model trained by a feature dimensionality reduction model based on feature refinement, and the classifier is a classifier trained by a random forest classifier based on quantum entropy encoding.

[0033] After training the corresponding emotion recognition model, emotion data is input into the model. A feature extraction model extracts features to obtain the speech features corresponding to the target speech data. Then, a feature dimensionality reduction model reduces the dimensionality of the speech features to obtain low-dimensional features. Finally, a classifier predicts the emotion category of the speech data. Thus, by performing feature extraction, dimensionality reduction, and classification on the speech data, the accuracy of speech emotion recognition is effectively improved. The emotion category can include happiness, sadness, anger, surprise, and neutrality, etc., and this invention does not limit this.

[0034] It should be noted that the training process of the emotion recognition model can include data augmentation and model training. The data augmentation process can include expanding the training data through a data augmentation model, while the model training process can include the training of corresponding models such as the data augmentation model, feature extraction model, feature dimensionality reduction model, and classifier.

[0035] Before training the model, data can be collected and labeled. The training data can include various publicly available emotional speech databases and on-site speech data collected through legal channels, covering a variety of languages, accents and emotion types. All collected data is stored in a lossless compression format and is accompanied by metadata tags, including but not limited to the recording environment, speaker's age, gender and accent information. The collected data is stored in a structured database, and each file is indexed by a unique identifier to facilitate subsequent data processing and access.

[0036] For emotion recognition in speech, emotion categories can be labeled as: happiness, sadness, anger, surprise, and neutral. Each emotion category is carefully reviewed and labeled by emotion recognition experts to ensure consistency and accuracy. In one example, the labeled speech data can be shown in Table 1 below:

[0037]

[0038] Table 1

[0039] Once the training data is determined, the processes of data collection, annotation, and preprocessing are extremely time-consuming and labor-intensive. Furthermore, insufficient training samples can easily lead to poor model generalization ability, affecting the model's prediction accuracy. To address this, in this embodiment of the invention, a generative adversarial network algorithm based on dynamic symmetry breaking theory can be used to generate samples, thereby expanding the data and enhancing its richness and the model's robustness.

[0040] Furthermore, considering that speech data is often affected by environmental noise in practical applications, in this embodiment of the invention, a time-frequency mask is also incorporated into the generative adversarial network framework to combat various noises during the data propagation stage, thereby enhancing the model's ability to process speech data in complex environments.

[0041] In some feasible implementations, a data augmentation model to be trained is obtained. This model includes at least a generator and a discriminator, and is configured with a symmetry breaking stimulus term. Then, a first target speech data is determined for training the data augmentation model. Next, first network parameters corresponding to the generator and second network parameters corresponding to the discriminator are randomly generated. The generator and discriminator are then trained alternately using the first target speech data, the first network parameters, and the second network parameters to obtain the generator loss function and the discriminator loss function. The generator parameters are then iterated using the generator loss function and the symmetry breaking stimulus term until a preset stopping condition is met, resulting in a target generator. The discriminator parameters are then iterated using the discriminator loss function until a preset stopping condition is met, resulting in a target discriminator. Finally, the target generator and the target discriminator are combined to obtain the data augmentation model.

[0042] In the specific implementation, for the generator G c and discriminator D c Network parameters and The generator is randomly initialized, and its goal is to learn a mapping function. The random noise vector z c The distribution of converted speech data is used, and the discriminator aims to distinguish between generated speech data and real speech data. Therefore, the initialization method is as follows:

[0043]

[0044] in, These are the parameters of the generator; σ represents the parameters of the discriminator; σ represents the standard deviation of the initialization parameters; the init() function is responsible for initializing the network parameters according to a normal distribution; ← represents the parameter update operation.

[0045] In adversarial training, the generator and discriminator are trained alternately. The discriminator updates its parameters to better distinguish between generated and real data, while the generator updates its parameters to generate speech samples that better match the distribution of real data. In this way, the generator and discriminator promote each other in continuous adversarial training, thereby improving the performance of the entire model.

[0046] During adversarial training, it is necessary to minimize the generator's loss function. and the loss function of maximizing the discriminator Specifically, the generator loss function corresponding to the generator can be calculated using formula (1):

[0047]

[0048] The discriminator loss function corresponding to the discriminator is calculated using formula (2):

[0049]

[0050] in, Let the generator loss function be... z is the discriminator loss function; c To start from the prior distribution Random noise obtained from sampling; x c This is real data; p data For the true data distribution; G c The ( ) function is a generator function; D c The ( ) function is the discriminator function; ☐ represents the expectation operation; ~ indicates that it follows a specific distribution.

[0051] After calculating the corresponding loss function through the above process, dynamic symmetry breaking adjustment can be performed. In each training cycle, the symmetry breaking stimulus term is dynamically adjusted according to the difference between the generated data and the real data, so that the generator can maintain data diversity while avoiding the problem of pattern collapse. Specifically, the symmetry breaking stimulus term can be calculated using formula (3):

[0052] Δ c =λ c ·(||G c (z c )-x c ||2-γ c )Formula (3)

[0053] Then, the parameters of the generator are updated using formula (4):

[0054]

[0055] Where, Δc Ω represents the symmetry breaking excitation term; c λ is the gradient adjustment function; c It is the adjustment coefficient for symmetry breaking excitation; γ c It is a preset symmetry breaking threshold; ||G c (z c )-x c ||2 is the Euclidean distance between the generated data and the real data; || ||2 is the L2 norm, calculated in the same way as the Euclidean distance. It is the learning rate of the generative adversarial network; Indicates to The gradient of λ. Optionally, λ c Set to 0.1, Set to 0.02, γ c Set it to 0.5.

[0056] In some feasible implementations, during the training process of the data augmentation model, an adaptive gradient modulation mechanism can be used to dynamically adjust the gradient update rate of the generator during training, so as to capture and simulate complex speech emotion features more precisely. Specifically, during the training of the generator, the adjustment gradient of the generator is adjusted based on the gradient adjustment function using formula (5):

[0057] Ω c =exp(-κ) c ·(1-Acc(G c (z c )))) Formula (5)

[0058] Among them, Acc(G c (z c )) represents the classification accuracy obtained by the discriminator from the generated samples, κ c It is a parameter for adjusting sensitivity; Ω c Ensure that adjustment stops when the accuracy of the generated samples' sentiment expression reaches a preset threshold; when the accuracy falls below the preset threshold, enhance gradient adjustment to quickly improve generator performance. Optionally, κ c It can be set to 2.

[0059] As in the above embodiment, in order to enhance the model's ability to process speech data in complex environments, a time-frequency mask is also combined in the generative adversarial network framework to combat various noises during the data propagation stage. Specifically, the first speech data used to train the model can be obtained, followed by obtaining a time-frequency mask and a random noise vector for the first speech data, and then optimizing the first speech data based on the time-frequency mask and the random noise vector to obtain the first target speech data.

[0060] In the specific implementation, the speech data is optimized using a time-frequency mask in each batch of generated data to maintain high quality even in noisy environments. The first speech data can be optimized using formulas (6) and (7) to obtain the first target speech data:

[0061] M c =f c (G c (Z c ))Formula (6)

[0062] G′ c (z c ) = M c ⊙G c (Z c )Formula (7)

[0063] Among them, Z c M is the random noise vector of the entire batch. c For time-frequency mask; f c The parentheses function is the mask generation function; ⊙ represents element-wise multiplication; G′ c (Z c ) is the generated data after applying a mask.

[0064] Furthermore, the mask generation function depends not only on a single generated sample but also on the statistical characteristics of the entire data batch to improve the dynamic adaptability of the mask. Specifically, the time-frequency mask for the first speech data can be calculated using formula (8):

[0065]

[0066] Among them, Z c It is the random noise vector of the entire batch. The Var() and Mean() functions calculate the variance and mean of the random noise vector, respectively, and the Max() function is used to calculate the maximum value.

[0067] It should be noted that during the model augmentation process using training data, the above process can be iterated repeatedly until a preset stopping iteration condition is met, indicating that the model training is complete. This allows for data augmentation based on the dynamic symmetry breaking theory of generative adversarial networks. By employing symmetry breaking incentive terms, the pattern collapse problem encountered in traditional generative adversarial networks during sample generation is solved, thereby enhancing the diversity and richness of generated samples. In one embodiment, the preset stopping iteration condition is reaching a preset maximum number of iterations. For example, the preset maximum number of iterations can be set to 1000.

[0068] For the feature extraction model, after the training data is augmented using a data augmentation model, the augmented training data can be input into the feature extraction model for training. Optionally, in this embodiment of the invention, a 6-layer fully connected neural network is used for feature extraction. However, in related technologies, some schemes use neural networks for feature extraction. In certain neural network structures, problems such as vanishing gradients, exploding gradients, or getting trapped in local optima may be encountered, affecting the stability of training and the performance of the model. To address this, in this embodiment of the invention, a neural network algorithm based on dynamic population evolution optimization is used as the feature extraction model. In traditional population evolution algorithms, the evolution of all individuals is based on fixed rules. This invention utilizes a self-correction mechanism to automatically adjust the evolution rules according to the characteristics of the current training data and adjusts the probabilities of crossover and mutation according to the changing trend of the loss function, thereby making the algorithm more flexible in adapting to different data distributions and improving the model's generalization ability and training efficiency.

[0069] In some feasible implementations, a second target speech data for training the feature extraction model can be determined first. This second target speech data can be speech data augmented based on a data augmentation model. Then, according to a biomimetic algorithm initialization method, an initial population is generated during the initialization phase. The population includes several individuals, each representing a configuration of network weights in the feature extraction model. The weights and biases corresponding to each individual are obtained, and the composite loss function of the individual in the second target speech data is calculated based on the weights and biases. Then, the fitness of the individual is obtained, and a target individual is selected from the individuals based on the fitness. Crossover and mutation operations are performed on the target individuals to generate new individuals. Then, the feature extraction model is iterated based on the composite loss function and the new individuals until a preset stopping iteration condition is met, thus obtaining the trained feature extraction model.

[0070] In the specific implementation, based on the biomimetic algorithm initialization method, an initial population is generated during the initialization phase. Each individual represents a network weight configuration, and the population size is represented as N. p Let the weights and biases of the i-th individual be initialized as follows:

[0071]

[0072] Among them, W pi Let b be the weight of the neural network corresponding to the i-th individual. pi Let be the bias of the neural network corresponding to the i-th individual. This represents the weight matrix of the i-th individual in the initial state; σ represents the bias of the i-th individual in the initial state; 2 Represents the initial variance; This indicates that the mean is 0 and the variance is σ.2 The normal distribution; It follows a normal distribution. Optionally, σ 2 It can be set to 0.01.

[0073] For each individual in the population, the corresponding neural network configuration can be used to process the input training data, calculate the model output, and evaluate its performance according to a predetermined loss function. Specifically, for the i-th individual, the extraction model loss function on the training set can be calculated based on the individual's corresponding weights and biases using formula (9):

[0074]

[0075] Among them, L pi is the loss of the neural network corresponding to the i-th individual; mps represents the number of samples input in the current batch; l() function represents the composite loss function, and fsig() function represents the neural network model function; Represent the features of the j-th sample; This represents the label of the j-th sample.

[0076] In an alternative example, the composite loss function may include a regularization term, which can increase the generalization ability of the model. Specifically, the regularization term is calculated using formula (10):

[0077]

[0078] Among them, W pi,k Let k be the weight of the neural network corresponding to the i-th individual;

[0079] The composite loss function is calculated using formula (11):

[0080]

[0081] Wherein, the MSE() function is the mean squared error function, Reg(W pi ) is the regularization term, λ ps This is the regularization parameter.

[0082] Once the composite loss function is determined, the best-performing individual in the current population can be selected based on its fitness and retained as a candidate solution for the next generation. Specifically, selection is based on the individual's fitness, with superior individuals having a higher probability of being selected. The probability of an individual being selected can be calculated using formula (12), and individuals with a probability greater than or equal to a preset threshold are selected as target individuals.

[0083]

[0084] Among them, Pselect (i) represents the probability that the i-th individual is selected; γ pse It is a parameter that controls the selected pressure; L pk Let be the loss of the neural network corresponding to the k-th individual.

[0085] Furthermore, new individuals can be generated through crossover and mutation operations. Crossover allows two superior individuals to exchange some genes, generating new offspring; mutation randomly alters some genes within an individual to increase population diversity. Specifically, crossover randomly selects two individuals for gene exchange, as shown below:

[0086] W′ pi =α pcs W p1 +(1-α pcs W p2

[0087] b′ pi =α pcs b p1 +(1-α pcs )b p2

[0088] Where, α pcs It's the crossover rate, W p1 b represents the weights of the neural network corresponding to the first selected individual. p1 W is the bias of the neural network corresponding to the first selected individual. p2 b represents the weights of the neural network corresponding to the selected second individual. p2 W′ is the bias of the neural network corresponding to the selected second individual. pi b′ represents the weights of the neural network corresponding to the individuals after the crossover operation. pi Let be the bias of the neural network corresponding to the individual after the crossover operation; and let be the mutation operation, which applies a small-amplitude random perturbation to the weights of the newly generated individuals, as follows:

[0089]

[0090] Where, τ 2 The variance, W″ represents the variation. pi b″ represents the weights of the neural network corresponding to the individual after the mutation operation. pi This represents the bias of the neural network corresponding to the individual after the mutation operation.

[0091] It should be noted that the above steps are repeated iteratively until a preset stopping iteration condition is met, indicating that the model training is complete. In one embodiment, the preset stopping iteration condition is reaching a preset maximum number of iterations. Thus, during the training process of the feature extraction model, feature extraction is performed based on a dynamic population evolution optimization model. Through a self-correction mechanism, the evolutionary rules are automatically adjusted according to the characteristics of the training data, improving the model's adaptability to different data distributions and significantly enhancing the stability and efficiency of feature extraction. Optionally, the preset maximum number of iterations is set to 1000.

[0092] After the feature extraction model obtained through the above training extracts the corresponding feature data, the extracted feature data can be input into the feature dimensionality reduction model for training. Specifically, this invention uses an autoencoder neural network based on feature refinement as the dimensionality reduction model. The autoencoder neural network based on feature refinement consists of three parts: an encoder, a decoder, and a feature adjustment unit. The encoder is responsible for mapping high-dimensional input data to a low-dimensional feature space, and the decoder is used to reconstruct the dimensionality-reduced features back to the original space to ensure the reversibility of the dimensionality reduction process. The feature adjustment unit dynamically adjusts the dimensionality-reduced feature space through recursive feature adaptive optimization, so that important features are strengthened and secondary features are gradually weakened, thereby making the dimensionality-reduced feature representation have good simplicity while retaining data information.

[0093] In some feasible implementations, the first feature data used to train the feature dimensionality reduction model can be determined first. Then, the first feature data is input into the encoder for mapping to obtain the corresponding initial low-dimensional features. The initial low-dimensional features are then input into the feature adjustment unit for feature adjustment. First, the feature weights corresponding to the low-dimensional features are determined, and the initial low-dimensional features are adjusted to the target low-dimensional features corresponding to the feature weights. Then, the target low-dimensional features are recursively optimized. In each iteration, the feature weights of each low-dimensional feature are adjusted according to the performance of each low-dimensional feature in the previous round to obtain the target feature weights corresponding to each low-dimensional feature. Then, the target low-dimensional features and their corresponding target feature weights are input into the decoder for reconstruction to obtain the second feature data corresponding to the target low-dimensional features. Finally, the first feature data and the second feature data are matched, and the feature dimensionality reduction model is iterated based on the matching results until the preset stopping iteration condition is met, thus obtaining the trained feature dimensionality reduction model.

[0094] In a practical implementation, we can assume that the data input to the autoencoder neural network is X. r The encoder employs a multi-layer nonlinear mapping structure to map high-dimensional data to an initial low-dimensional feature space. Therefore, the first feature data can be mapped to the initial low-dimensional feature space using formula (13) to obtain the corresponding low-dimensional features.

[0095] Zr =Sig enc (W r X r +b r )Formula (13)

[0096] Among them, Z r W represents the initial low-dimensional features. r Let b be the weight matrix of the encoder. r Sig is the bias vector of the encoder. enc The () function is the multi-level Sigmoid activation function for the encoder.

[0097] After the low-dimensional features are generated, the feature adjustment module automatically generates feature weights based on the importance of features in the current feature space. During initialization, this module assigns the same initial weights to all features so that they can be gradually adjusted based on feature contributions in subsequent steps. Specifically, this can be achieved by obtaining an initial weight matrix for the low-dimensional features; then, the low-dimensional features are input into the feature adjustment unit, and the initial low-dimensional features are adjusted using the initial weight matrix to obtain the corresponding target low-dimensional features.

[0098] The initial weight matrix is ​​as follows:

[0099] A r =diag(α) r )

[0100] Among them, A r The initial weight matrix; α r Let α be the feature weight vector. r Each element α in r,i Initializing to the same value indicates that all features have the same importance in the initial stage; diag() is a function to extract the diagonal elements of the matrix;

[0101] The target low-dimensional features are represented as follows:

[0102] Z′ r =A r Z r

[0103] Among them, Z′ r The target low-dimensional features are adjusted by feature weights.

[0104] After the target low-dimensional features are input into the feature adjustment unit, the feature adjustment module recursively optimizes the initially generated low-dimensional features. In each iteration, the module adjusts the weights of each feature based on the performance of the features in the previous round, gradually strengthening those features with significant influence and gradually weakening redundant or noisy features, thus recursively optimizing the target low-dimensional features. In the t-th iteration, the feature weights are updated using formula (14):

[0105]

[0106] in, This represents the feature weights in the (t+1)th iteration. η represents the feature weights in the t-th iteration. r L is the learning rate for the feature dimensionality reduction model. r The () function is the loss function of the feature dimensionality reduction model, Y r For label data, The function represents the gradient of the loss function with respect to the feature weights;

[0107] In this model, the loss function can be the reconstruction error loss function.

[0108] To ensure the effectiveness of the dimensionality reduction process, the decoder remaps the low-dimensional features back to the high-dimensional space to ensure that no important information is lost during the dimensionality reduction process. The target low-dimensional features can be input into the decoder for reconstruction. Specifically, the weight matrix and bias vector corresponding to the decoder can be obtained, and the target low-dimensional features, the weight matrix, and the bias vector can be processed according to the multi-layer activation function configured in the decoder to construct the second feature data corresponding to the target low-dimensional features.

[0109] For example, the reconstruction process of the decoder can be represented as:

[0110] X′ r =Sig dec (W′ r Z′ r +b′ r )

[0111] Where, X′ r For the reconstructed high-dimensional data, W′ r Let b′ be the weight matrix of the decoder. r Sig is the bias vector of the decoder. dec The () function is the multi-layer Sigmoid activation function for the decoder.

[0112] It should be noted that during the training of the feature reduction model, the above steps can be iterated repeatedly until a preset stopping iteration condition is met, indicating that the model training is complete. Thus, during the training of the feature reduction model, an autoencoder neural network based on feature refinement is used as the feature reduction model. Through recursive feature adaptive optimization, the importance of the reduced features is dynamically adjusted, achieving effective feature compression while preserving important information and reducing redundancy. In one embodiment, the preset stopping iteration condition is reaching a preset maximum number of iterations. Optionally, the preset maximum number of iterations can be set to 1000.

[0113] After training the feature dimensionality reduction model through the above process, the dimensionality-reduced data can be further input into the classifier for model training. Specifically, this invention uses a random forest based on quantum entropy encoding as the classification algorithm, employing a quantum entropy encoding mechanism to optimize the tree node selection process, thereby significantly improving classification accuracy and efficiency. Quantum entropy encoding is a technique that uses the probability amplitude of qubits to encode information. Through this mechanism, this invention can more effectively capture and utilize the inherent uncertainty and diversity of data, thus optimizing the decision tree construction process.

[0114] In some feasible implementations, the training process of the random forest algorithm based on quantum entropy encoding may include: initializing a random forest classifier composed of multiple decision trees, each decision tree containing several nodes; during the construction of each decision tree, obtaining the quantum entropy corresponding to the node; then using quantum entropy to calculate the information gain for the decision tree; and splitting the nodes according to the information gain to determine the result growth tree structure corresponding to the decision tree; then comparing the result growth tree structure with the preset real structure; iterating the decision tree based on the comparison result until a preset stopping iteration condition is met to obtain the trained decision tree; and finally integrating all trained decision trees to obtain the final random forest classifier.

[0115] In practical implementation, a random forest composed of multiple decision trees can be initialized first. The parameter initialization method is expressed as follows:

[0116]

[0117] in, N represents the parameters of the i-th tree; T It is the total number of trees; σ u This is the initial standard deviation; the `init()` function is used to initialize the parameters according to a Gaussian distribution. Optionally, N... T Set it to 100.

[0118] In the construction of each decision tree, the selection of nodes is achieved through a quantum entropy encoding mechanism. By evaluating the change in quantum entropy after each feature split, the split that maximizes the quantum entropy is selected, thereby enhancing the model's sensitivity to data and classification accuracy. Specifically, the quantum entropy corresponding to the node can be calculated using formula (14):

[0119]

[0120] Where S represents the dataset of the current node; f is the segmentation feature considered; θ is the segmentation threshold; p k S represents the probability of the k-th label in S; K is the total number of categories. It is the partitioned subset; λ u It is a balance coefficient used to adjust the weights of the segmentation uniformity and information gain.

[0121] In one embodiment, p k The calculation method considers the frequency of occurrence of each category and the statistical weight of that category at a specific node, and is expressed as follows:

[0122]

[0123] In the formula, w i It is the weight of the i-th data point, y i It is the category of the i-th data point, [y i =k] is an indicator function, when y i The value is 1 when it equals k, and 0 otherwise.

[0124] Furthermore, To account for the impact of imbalanced data, a correction factor μ is used. j To adjust the influence of each subset, the correction coefficient can be calculated based on the diversity and distribution density of the subsets, and is expressed as:

[0125]

[0126] In the formula, Is category k in subset The relative frequency in Representing a subset Volume or coverage in feature space; Div() function is the diversity function; Den() function is the density function.

[0127] After calculating the corresponding quantum entropy through the above process, a tree structure can be grown based on the result of the quantum entropy encoding. The growth process of each tree involves recursively splitting the dataset until the stopping condition is met. The node splitting decision is based on maximizing information gain. Specifically, the information gain for the decision tree can be calculated using formula (15):

[0128]

[0129] Among them, IG u Information gain reflects the degree of entropy reduction after segmenting nodes using feature f.

[0130] Furthermore, through iterative training, the parameters and structure of each tree are continuously optimized, allowing the entire model to gradually adapt to the characteristics of the training data. In each iteration, the quantum entropy encoding strategy and game theory strategy are adjusted and optimized by evaluating the classification performance of the entire forest to achieve the optimal classification effect. Among them, game theory strategies can be used to adjust the growth strategies between decision trees to optimize the overall forest performance. Specifically, in each iteration, the growth strategies between the decision trees are adjusted using formula (16):

[0131]

[0132] in, The utility function of the i-th decision tree depends on its own parameters and the parameters θ of the other decision trees. -i ;Γ u This represents the optimization function used to maximize the overall synergistic utility of the random forest.

[0133] Finally, through an ensemble learning strategy, all optimized tree models are integrated to form the final random forest classifier. The decisions of all trees are ensembled through a voting mechanism to determine the final classification result, represented as:

[0134]

[0135] Where x is the input feature vector, specifically the feature vector after dimensionality reduction; This is the prediction result of the i-th tree; the mode() function selects the most frequently occurring class label as the final prediction, thereby optimizing the tree node selection process during the training of the classifier model using a random forest classifier based on quantum entropy encoding. With the goal of maximizing quantum entropy, it improves classification accuracy and efficiency, and achieves efficient classification of complex data.

[0136] After training the corresponding emotion recognition model through the above process, the trained model can be used to process new samples to achieve speech emotion recognition. In one embodiment, the collected raw data is input into the trained feature extraction and feature dimensionality reduction model for feature processing. Further, the processed features are input into a classifier model for classifier training, thereby obtaining the classification result. In this embodiment, the classification categories include: happiness, sadness, anger, surprise, and neutrality. The processed emotion recognition result can be output to the application system or end user, and can be used in various application scenarios, such as customer service, emotion monitoring, and interactive entertainment.

[0137] It should be noted that the embodiments of the present invention include, but are not limited to, the examples described above. It is understood that those skilled in the art can make further settings according to actual needs under the guidance of the ideas in the embodiments of the present invention, and the present invention does not limit such settings.

[0138] In this embodiment of the invention, during the process of speech emotion recognition, an emotion recognition model for speech emotion recognition is obtained. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The feature extraction model is a model trained based on a neural network algorithm with dynamic population evolution optimization. The feature dimensionality reduction model is a model trained based on a feature refinement model. The classifier is a classifier trained based on a random forest classifier with quantum entropy encoding. First, the target speech data is input into the feature extraction model for feature extraction to obtain the speech features corresponding to the target speech data. Then, the speech features are input into the feature dimensionality reduction model for feature reduction to obtain the low-dimensional features corresponding to the speech features. Finally, the low-dimensional features are input into the classifier for prediction to obtain the emotion category of the speech data. Thus, by performing feature extraction, feature dimensionality reduction, and classification on the speech data, the accuracy of speech emotion recognition is effectively improved.

[0139] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following examples are provided for illustrative purposes:

[0140] Reference Figure 2This diagram illustrates the training process of the neural network algorithm provided in this embodiment of the invention. The feature extraction model can be a model built based on the neural network algorithm. The corresponding training process can include: initializing the population - calculating the loss for each individual - evaluating performance based on the loss - selecting the best-performing individual - checking operations to generate new individuals - mutation operations to increase diversity - checking the stopping iteration condition - if satisfied, the model training is completed; if not satisfied, the loss is calculated for each individual to complete the model training. Thus, feature extraction is performed based on a neural network algorithm optimized by dynamic population evolution. Through a self-correction mechanism, the evolution rules are automatically adjusted according to the characteristics of the training data, which improves the model's adaptability to different data distributions and significantly enhances the stability and efficiency of feature extraction.

[0141] Reference Figure 3 This diagram illustrates the training process of the autoencoder neural network algorithm provided in this embodiment of the invention. The feature dimensionality reduction model can be obtained by training an autoencoder neural network algorithm. The corresponding training process may include: input data - encoder dimensionality reduction - generation of low-dimensional features - feature adjustment module - initial feature weights - weight optimization - generation of optimized features - decoder reconstruction of high-dimensional features - calculation of loss function - determination of whether the iteration condition is met - if met, stop iteration; if not met, continue weight optimization - end training. Thus, based on the feature refinement autoencoder neural network as the feature dimensionality reduction model, through recursive feature adaptive optimization, the importance of the dimensionality-reduced features is dynamically adjusted, achieving effective feature compression, retaining important information while reducing redundancy.

[0142] Furthermore, for other model training processes, please refer to the descriptions in the aforementioned embodiments, which will not be repeated here. Through the implementation of data augmentation techniques, the number of samples is significantly increased, improving the model's generalization ability and robustness, and enhancing the emotion recognition performance in diverse environments. Moreover, during feature extraction, dynamically adjusting feature weights improves the representational power of important features, reduces instability during training, and thus improves the overall recognition accuracy. Additionally, the feature dimensionality reduction stage effectively reduces the data dimensionality, making subsequent classification processes more efficient while ensuring the integrity of important information and optimizing feature representation. Finally, with the support of quantum entropy encoding, the random forest classifier can better capture the inherent uncertainty of the data, improving the accuracy of classification decisions and enhancing the model's adaptability in practical applications.

[0143] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0144] Reference Figure 4 The diagram shows a structural block diagram of a voice emotion recognition device provided in an embodiment of the present invention, which may specifically include the following modules:

[0145] The data acquisition module 401 is used to acquire target speech data, an emotion recognition model for speech emotion recognition, and a data augmentation model for data augmentation. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The data augmentation model is a model trained based on dynamic symmetry breaking theory. The feature extraction model is a model trained based on a neural network algorithm with dynamic population evolution optimization. The feature dimensionality reduction model is a model trained based on a feature refinement-based feature dimensionality reduction model. The classifier is a classifier trained based on a random forest classifier with quantum entropy encoding.

[0146] Data augmentation module 402 is used to input the voice data into the data augmentation model for data augmentation to obtain target voice data;

[0147] Feature extraction module 403 is used to input the target speech data into the feature extraction model for feature extraction to obtain the speech features corresponding to the target speech data;

[0148] The feature dimensionality reduction module 404 is used to input the speech features into the feature dimensionality reduction model to perform feature dimensionality reduction and obtain the low-dimensional features corresponding to the speech features.

[0149] The prediction module 405 is used to input the low-dimensional features into the classifier for prediction to obtain the emotion category of the speech data.

[0150] In some feasible embodiments, the apparatus further includes:

[0151] The model acquisition module is used to acquire the data augmentation model to be trained. The data augmentation model includes at least a generator and a discriminator, and the data augmentation model is configured with a symmetry breaking stimulus term.

[0152] The first data determination module is used to determine the first target speech data for training the data augmentation model;

[0153] The first parameter generation module is used to randomly generate the first network parameters corresponding to the generator;

[0154] The second parameter generation module is used to randomly generate the second network parameters corresponding to the discriminator;

[0155] The loss function calculation module is used to alternately train the generator and the discriminator using the first target speech data, the first network parameters and the second network parameters to obtain the generator loss function corresponding to the generator and the discriminator loss function corresponding to the discriminator.

[0156] The first iteration module is used to iterate the parameters of the generator using the generator loss function and the symmetry breaking incentive term until a preset stopping iteration condition is met, thereby obtaining the target generator.

[0157] The first processing module is used to iterate the parameters of the discriminator using the discriminator loss function until a preset stopping iteration condition is met, thereby obtaining the target discriminator.

[0158] The model determination module is used to combine the target generator and the target discriminator to obtain a data augmentation model.

[0159] In some feasible embodiments, the loss function calculation module is specifically used for:

[0160] The generator loss function corresponding to the generator is calculated using formula (1):

[0161]

[0162] The discriminator loss function corresponding to the discriminator is calculated using formula (2):

[0163]

[0164] in, Let the generator loss function be... z is the discriminator loss function; c To start from the prior distribution Random noise obtained from sampling; x c This is real data; p data For the true data distribution; G c ( ) represents the generator function; D c ( ) represents the discriminator function; ☐ represents the expectation operation; ~ indicates that it follows a specific distribution.

[0165] In some feasible embodiments, the first iteration module is specifically used for:

[0166] The symmetry breaking excitation term is calculated using formula (3):

[0167] Δ c =λ c ·(||G c (z c )-x c ||2-γ c )Formula (3)

[0168] The parameters of the generator are updated using formula (4):

[0169]

[0170] Where, Δ c Ω represents the symmetry breaking excitation term; c λ is the gradient adjustment function; c It is the adjustment coefficient for symmetry breaking excitation; γ c It is a preset symmetry breaking threshold; ||G c (z c )-x c ||2 is the Euclidean distance between the generated data and the real data; || ||2 is the L2 norm, calculated in the same way as the Euclidean distance. It is the learning rate of the generative adversarial network; Indicates to The gradient.

[0171] In some feasible embodiments, λ c Set to 0.1, Set to 0.02, γ c Set it to 0.5.

[0172] In some feasible embodiments, it also includes:

[0173] The gradient adjustment module is used to adjust the adjustment gradient of the generator based on the gradient adjustment function according to formula (5) during the training of the generator:

[0174] Ω c =exp(-κ) c ·(1-Acc(G c (z c )))) Formula (5)

[0175] Among them, Acc(G c (z c )) represents the classification accuracy obtained by the discriminator from the generated samples, κ c It is a parameter for adjusting sensitivity; Ω cEnsure that adjustments are stopped when the accuracy of the generated sample's emotional expression reaches a preset threshold; when the accuracy is less than the preset threshold, enhance gradient adjustments to quickly improve generator performance.

[0176] In some feasible embodiments, the first data determining module is specifically used for:

[0177] Acquire first speech data for training the data augmentation model;

[0178] Obtain the time-frequency mask and random noise vector for the first speech data;

[0179] The first speech data is optimized based on the time-frequency mask and the random noise vector to obtain the first target speech data.

[0180] In some feasible embodiments, the first data determining module is specifically used for:

[0181] The first speech data is optimized using formulas (6) and (7) to obtain the first target speech data:

[0182] M c =f c (G c (Z c ))Formula (6)

[0183] G′ c (z c ) = M c ⊙G c (Z c )Formula (7)

[0184] Among them, Z c M is the random noise vector of the entire batch. c For time-frequency mask; f c () represents the mask generating function; ⊙ represents element-wise multiplication; G′ c (Z c ) is the generated data after applying a mask.

[0185] In some feasible embodiments, the first data determining module is specifically used for:

[0186] The time-frequency mask for the first speech data is calculated using formula (8):

[0187]

[0188] Among them, Z c It is the random noise vector of the entire batch. Var() and Mean() calculate the variance and mean of the random noise vector, respectively, and Max() calculates the maximum value.

[0189] In some feasible embodiments, it also includes:

[0190] The second data determination module is used to determine the second target speech data for training the feature extraction model;

[0191] An initialization module is used to generate an initial population during the initialization phase according to the biomimetic algorithm initialization method. The population includes several individuals, and each individual represents a configuration of network weights in the feature extraction model.

[0192] The calculation module is used to obtain the weights and biases corresponding to the individual, and calculate the composite loss function of the individual in the second target speech data based on the weights and biases;

[0193] An individual selection module is used to obtain the fitness of the individuals and select a target individual from the individuals based on the fitness.

[0194] Individual processing module, used to perform crossover and mutation operations on the target individual to generate new individuals;

[0195] The second iteration module is used to iterate the feature extraction model according to the composite loss function and the new individual until a preset stopping iteration condition is met, thereby obtaining a trained feature extraction model.

[0196] In some feasible embodiments, the initialization module is specifically used for:

[0197] Population size is represented as N. p Let the weights and biases of the i-th individual be initialized as follows:

[0198]

[0199] Among them, W pi Let b be the weight of the neural network corresponding to the i-th individual. pi Let be the bias of the neural network corresponding to the i-th individual. This represents the weight matrix of the i-th individual in the initial state; σ represents the bias of the i-th individual in the initial state; 2 Represents the initial variance; This indicates that the mean is 0 and the variance is σ. 2 The normal distribution; It follows a normal distribution.

[0200] In some feasible embodiments, the computing module is specifically used for:

[0201] Using formula (9), for the i-th individual, the extraction model loss function on the training set is calculated based on the individual's corresponding weights and biases:

[0202]

[0203] Among them, L pi is the loss of the neural network corresponding to the i-th individual; mps represents the number of samples input in the current batch; l() represents the composite loss function, and fsig() represents the neural network model function; Represent the features of the j-th sample; This represents the label of the j-th sample.

[0204] In some feasible embodiments, the composite loss function includes a regularization term, which is calculated using formula (10):

[0205]

[0206] Among them, W pi,k Let k be the weight of the neural network corresponding to the i-th individual;

[0207] The composite loss function is calculated using formula (11):

[0208]

[0209] Where MSE() is the mean squared error function, Reg(W pi ) is the regularization term, λ ps This is the regularization parameter.

[0210] In some feasible embodiments, the individual selection module is specifically used for:

[0211] The probability of the individual being selected is calculated using formula (12), and individuals with a probability greater than or equal to a preset threshold are selected as target individuals:

[0212]

[0213] Among them, P select (i) represents the probability that the i-th individual is selected; γ pse It is a parameter that controls the selected pressure; L pk Let be the loss of the neural network corresponding to the k-th individual.

[0214] In some feasible embodiments, the individual processing module is specifically used for:

[0215] Crossover operations randomly select two individuals for gene exchange, represented as:

[0216] W′ pi=α pcs W p1 +(1-α pcs W p2

[0217] b′ pi =α pcs b p1 +(1-α pcs )b p2

[0218] Where, α pcs It's the crossover rate, W p1 b represents the weights of the neural network corresponding to the first selected individual. p1 W is the bias of the neural network corresponding to the first selected individual. p2 b represents the weights of the neural network corresponding to the selected second individual. p2 W′ is the bias of the neural network corresponding to the selected second individual. pi b′ represents the weights of the neural network corresponding to the individuals after the crossover operation. pi Let be the bias of the neural network corresponding to the individual after the crossover operation; and let be the mutation operation, which applies a small-amplitude random perturbation to the weights of the newly generated individuals, as follows:

[0219]

[0220] Where, τ 2 The variance, W″ represents the variation. pi b″ represents the weights of the neural network corresponding to the individual after the mutation operation. pi This represents the bias of the neural network corresponding to the individual after the mutation operation.

[0221] In some feasible embodiments, the feature dimensionality reduction model includes at least an encoder, a decoder, and a feature adjustment unit, and the device further includes:

[0222] The third data determination module is used to determine the first feature data for training the feature dimensionality reduction model.

[0223] The first mapping module is used to input the first feature data into the encoder for mapping to obtain the corresponding initial low-dimensional features;

[0224] The feature adjustment module is used to input the initial low-dimensional features into the feature adjustment unit for feature adjustment. First, the feature weights corresponding to the low-dimensional features are determined, and then the initial low-dimensional features are adjusted to the target low-dimensional features corresponding to the feature weights.

[0225] An iterative module is used to recursively optimize the target low-dimensional features. In each iteration, the feature weights of each low-dimensional feature are adjusted according to the performance of each low-dimensional feature in the previous round to obtain the target feature weights corresponding to each low-dimensional feature.

[0226] The reconstruction module is used to input the target low-dimensional feature and the corresponding target feature weight into the decoder for reconstruction, so as to obtain the second feature data corresponding to the target low-dimensional feature.

[0227] The model building module is used to match the first feature data and the second feature data, and iterate the feature dimensionality reduction model based on the matching results until a preset stopping iteration condition is met, so as to obtain the trained feature dimensionality reduction model.

[0228] In some feasible embodiments, the first mapping module is specifically used for:

[0229] The first feature data is mapped to the initial low-dimensional feature space using formula (13) to obtain the corresponding low-dimensional features:

[0230] Z r =Sig enc (W r X r +b r )Formula (13)

[0231] Among them, Z r W represents the initial low-dimensional features. r Let b be the weight matrix of the encoder. r Sig is the bias vector of the encoder. enc () is the multi-layer Sigmoid activation function of the encoder.

[0232] In some feasible embodiments, the feature adjustment module is specifically used for:

[0233] Obtain the initial weight matrix for the low-dimensional features;

[0234] The low-dimensional features are input into the feature adjustment unit, and the initial low-dimensional features are adjusted by the initial weight matrix to obtain the corresponding target low-dimensional features;

[0235] The initial weight matrix is ​​as follows:

[0236] A r =diag(α) r )

[0237] Among them, A r The initial weight matrix; α r Let α be the feature weight vector. rEach element α in r,i Initializing to the same value indicates that all features have the same importance in the initial stage; diag() is a function to extract the diagonal elements of the matrix;

[0238] The target low-dimensional features are represented as follows:

[0239] Z′ r =A r Z r

[0240] Among them, Z′ r The target low-dimensional features are adjusted by feature weights.

[0241] In some feasible embodiments, the iterative module is specifically used for:

[0242] The target low-dimensional features are recursively optimized, and in the t-th iteration, the feature weights are updated using formula (14):

[0243]

[0244] in, This represents the feature weights in the (t+1)th iteration. η represents the feature weights in the t-th iteration. r L is the learning rate for the feature dimensionality reduction model. r () represents the loss function of the feature reduction model, Y r For label data, This represents the gradient of the loss function with respect to the feature weights;

[0245] The loss function of the feature dimensionality reduction model is the reconstruction error loss function.

[0246] In some feasible embodiments, the reconstruction module is specifically used for:

[0247] Obtain the weight matrix and bias vector corresponding to the decoder;

[0248] The target low-dimensional feature, the weight matrix, and the bias vector are processed according to the multi-layer activation function configured in the decoder to construct the second feature data corresponding to the target low-dimensional feature.

[0249] In some feasible embodiments, it also includes:

[0250] The classifier initialization module is used to initialize a random forest classifier composed of multiple decision trees, wherein each decision tree includes several nodes.

[0251] A quantum entropy determination module is used to obtain the quantum entropy corresponding to each node during the construction of each decision tree;

[0252] The segmentation module is used to calculate the information gain of the decision tree using the quantum entropy, and to segment the nodes according to the information gain to determine the result growth tree structure corresponding to the decision tree.

[0253] The decision tree iteration module is used to compare the resulting growth tree structure with the preset real structure, and iterate the decision tree based on the comparison result until the preset stopping iteration condition is met to obtain the trained decision tree.

[0254] The ensemble module is used to integrate all trained decision trees to obtain the final random forest classifier.

[0255] In some feasible embodiments, the quantum entropy determination module is specifically used for:

[0256] The quantum entropy corresponding to the node is calculated using formula (14):

[0257]

[0258] Where S represents the dataset of the current node; f is the segmentation feature considered; θ is the segmentation threshold; p k S represents the probability of the k-th label in S; K is the total number of categories. It is the partitioned subset; λ u It is a balance coefficient used to adjust the weights of the segmentation uniformity and information gain.

[0259] In some feasible embodiments, the segmentation module is specifically used for:

[0260] The information gain for the decision tree is calculated using formula (15):

[0261]

[0262] Among them, IG u Information gain reflects the degree of entropy reduction after segmenting nodes using feature f.

[0263] In some feasible embodiments, the decision tree iteration module is specifically used for:

[0264] In each iteration, the growth strategy between the decision trees is adjusted using formula (16):

[0265]

[0266] in, The utility function of the i-th decision tree depends on its own parameters and the parameters θ of the other decision trees. -i ;Γ uThis represents the optimization function used to maximize the overall synergistic utility of the random forest.

[0267] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0268] In addition, this invention also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described speech emotion recognition method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0269] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described speech emotion recognition method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0270] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0271] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, EEPROM, Flash, and eMMC, etc.) containing computer-usable program code.

[0272] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0273] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0274] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0275] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0276] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0277] The present invention has provided a detailed description of a method and device for recognizing voice emotions. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for recognizing voice emotion, characterized in that, include: Acquire target speech data and an emotion recognition model for speech emotion recognition. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The feature extraction model is a model trained using a neural network algorithm based on dynamic population evolution optimization. The feature dimensionality reduction model is a model trained using a feature dimensionality reduction model based on feature refinement. The classifier is a classifier trained using a random forest classifier based on quantum entropy encoding. The target speech data is input into the feature extraction model for feature extraction to obtain the speech features corresponding to the target speech data; The speech features are input into the feature dimensionality reduction model to perform feature dimensionality reduction, thereby obtaining the low-dimensional features corresponding to the speech features. The low-dimensional features are input into the classifier for prediction to obtain the emotion category of the speech data; The method further includes: Obtain a data augmentation model to be trained, the data augmentation model including at least a generator and a discriminator, and the data augmentation model is configured with a symmetry breaking stimulus term; Determine the first target speech data for training the data augmentation model; Randomly generate the first network parameters corresponding to the generator; Randomly generate the second network parameters corresponding to the discriminator; The generator and the discriminator are trained alternately using the first target speech data, the first network parameters, and the second network parameters to obtain the generator loss function and the discriminator loss function. The generator parameters are iterated using the generator loss function and the symmetry breaking incentive term until a preset stopping iteration condition is met, thereby obtaining the target generator; The discriminator parameters are iterated using the discriminator loss function until a preset stopping iteration condition is met, thereby obtaining the target discriminator; The target generator and the target discriminator are combined to obtain a data augmentation model; The step of iterating the generator parameters using the generator loss function and the symmetry breaking stimulus term until a preset stopping condition is met to obtain the target generator includes: The symmetry breaking excitation term is calculated using formula (3): Δ c = λ c · (||G c (z c ) - x c ||² - γ c ) Equation (3) The parameters of the generator are updated using formula (4): Where, Δ c G represents the symmetry breaking excitation term; c () represents the generator function; z c To start from the prior distribution Random noise obtained from sampling; x c This is real data; Ω c λ is the gradient adjustment function; c It is the adjustment coefficient for symmetry breaking excitation; γ c It is a preset symmetry breaking threshold; ||G c (z c )-x c ||2 is the Euclidean distance between the generated data and the real data; ||||2 is the L2 norm, calculated in the same way as the Euclidean distance. Let the generator loss function be... It is the learning rate of the generative adversarial network; Indicates to The gradient.

2. The method according to claim 1, characterized in that, The step of alternately training the generator and the discriminator using the first target speech data, the first network parameters, and the second network parameters to obtain the generator loss function and the discriminator loss function includes: The generator loss function corresponding to the generator is calculated using formula (1): The discriminator loss function corresponding to the discriminator is calculated using formula (2): in, p is the discriminator loss function; data For the true data distribution; D c () represents the discriminator function; ☐ represents the expectation operation; ~ indicates that it follows a specific distribution.

3. The method according to claim 1, characterized in that, λ c Set to 0.1, Set to 0.02, γ c Set it to 0.

5.

4. The method according to claim 2, characterized in that, Also includes: During the training of the generator, the adjustment gradient of the generator is adjusted based on the gradient adjustment function using formula (5): Ω c =exp(-κ) c ·(1-Acc(G c (z c )))) Formula (5) Among them, Acc(G c (z c )) represents the classification accuracy obtained by the discriminator from the generated samples, κ c It is a parameter for adjusting sensitivity; Ω c Ensure that adjustments are stopped when the accuracy of the generated sample's emotional expression reaches a preset threshold; when the accuracy is less than the preset threshold, enhance gradient adjustments to quickly improve generator performance.

5. The method according to claim 2, characterized in that, The determination of the first target speech data for training the data augmentation model includes: Acquire first speech data for training the data augmentation model; Obtain the time-frequency mask and random noise vector for the first speech data; The first speech data is optimized based on the time-frequency mask and the random noise vector to obtain the first target speech data.

6. The method according to claim 5, characterized in that, The step of optimizing the first speech data based on the time-frequency mask and the random noise vector to obtain the first target speech data includes: The first speech data is optimized using formulas (6) and (7) to obtain the first target speech data: M c =f c (G c (Z c )) Formula (6) G′ c (z c ) = M c ⊙G c (Z c ) Formula (7) Among them, Z c M is the random noise vector of the entire batch. c For time-frequency mask; f c () represents the mask generating function; ⊙ represents element-wise multiplication; G′ c (Z c ) is the generated data after applying a mask.

7. The method according to claim 6, characterized in that, The step of obtaining the time-frequency mask for the first speech data includes: The time-frequency mask for the first speech data is calculated using formula (8): Among them, Z c It is the random noise vector of the entire batch. Var() and Mean() calculate the variance and mean of the random noise vector, respectively, and Max() calculates the maximum value.

8. The method according to claim 1, characterized in that, Also includes: Determine the second target speech data for training the feature extraction model; According to the biomimetic algorithm initialization method, in the initialization stage, an initial population is generated, which includes several individuals, each of which represents a configuration of network weights in the feature extraction model. Obtain the weights and biases corresponding to the individual, and calculate the composite loss function of the individual in the second target speech data based on the weights and biases; Obtain the fitness of the individuals, and select a target individual from the individuals based on the fitness; Crossover and mutation operations are performed on the target individual to generate a new individual; The feature extraction model is iterated based on the composite loss function and the new individual until a preset stopping condition is met, thus obtaining a trained feature extraction model.

9. The method according to claim 8, characterized in that, The initialization method based on the biomimetic algorithm generates an initial population during the initialization phase, including: Population size is represented as N. p Let the weights and biases of the i-th individual be initialized as follows: Among them, W pi Let b be the weight of the neural network corresponding to the i-th individual. pi Let be the bias of the neural network corresponding to the i-th individual. This represents the weight matrix of the i-th individual in the initial state; σ represents the bias of the i-th individual in the initial state; 2 Represents the initial variance; This indicates that the mean is 0 and the variance is σ. 2 The normal distribution; It follows a normal distribution.

10. The method according to claim 8 or 9, characterized in that, The step of calculating the composite loss function of the individual in the second target speech data based on the weights and the bias includes: Using formula (9), for the i-th individual, the extraction model loss function on the training set is calculated based on the individual's corresponding weights and biases: Among them, L pi is the loss of the neural network corresponding to the i-th individual; mps represents the number of samples input in the current batch; The function represents the composite loss function, and fsig() represents the neural network model function. Represent the features of the j-th sample; This represents the label of the j-th sample.

11. The method according to claim 8, characterized in that, The composite loss function includes a regularization term, which is calculated using formula (10): Among them, W pi,k Let k be the weight of the neural network corresponding to the i-th individual; The composite loss function is calculated using formula (11): Where MSE() is the mean squared error function, Reg(W pi ) is the regularization term, λ ps This is the regularization parameter.

12. The method according to claim 8, characterized in that, The step of selecting a target individual from the individuals based on the fitness includes: The probability of the individual being selected is calculated using formula (12), and individuals with a probability greater than or equal to a preset threshold are selected as target individuals: Among them, P select (i) represents the probability that the i-th individual is selected; γ pse It is a parameter that controls the selected pressure; L pk Let be the loss of the neural network corresponding to the k-th individual.

13. The method according to claim 8, characterized in that, The process of performing crossover and mutation operations on the target individual to generate a new individual includes: Crossover operations randomly select two individuals for gene exchange, represented as: W′ pi =a pcs W p1 +(1-a pcs )W p2 b′ pi =a pcs b p1 +(1-a pcs )b p2 Where, α pcs It's the crossover rate, W p1 b represents the weights of the neural network corresponding to the first selected individual. p1 W is the bias of the neural network corresponding to the first selected individual. p2 b represents the weights of the neural network corresponding to the selected second individual. p2 W′ is the bias of the neural network corresponding to the selected second individual. pi b′ represents the weights of the neural network corresponding to the individuals after the crossover operation. pi Let be the bias of the neural network corresponding to the individual after the crossover operation; and let be the mutation operation, which applies a small-amplitude random perturbation to the weights of the newly generated individuals, as follows: Where, τ 2 W″ represents the variance of the variation. pi b″ represents the weights of the neural network corresponding to the individual after the mutation operation. pi This represents the bias of the neural network corresponding to the individual after the mutation operation.

14. The method according to claim 1, characterized in that, The feature dimensionality reduction model includes at least an encoder, a decoder, and a feature adjustment unit, and the method further includes: Determine the first feature data to be used for training the feature dimensionality reduction model; The first feature data is input into the encoder for mapping to obtain the corresponding initial low-dimensional features; The initial low-dimensional features are input into the feature adjustment unit for feature adjustment. First, the feature weights corresponding to the low-dimensional features are determined, and then the initial low-dimensional features are adjusted to the target low-dimensional features corresponding to the feature weights. The target low-dimensional features are recursively optimized. In each iteration, the feature weights of each low-dimensional feature are adjusted based on the performance of each low-dimensional feature in the previous round to obtain the target feature weights corresponding to each low-dimensional feature. The target low-dimensional feature and the corresponding target feature weight are input into the decoder for reconstruction to obtain the second feature data corresponding to the target low-dimensional feature; The first feature data and the second feature data are matched, and the feature dimensionality reduction model is iterated based on the matching results until a preset stopping iteration condition is met, thereby obtaining the trained feature dimensionality reduction model.

15. The method according to claim 14, characterized in that, The step of inputting the first feature data into the encoder for mapping to obtain the corresponding initial low-dimensional features includes: The first feature data is mapped to the initial low-dimensional feature space using formula (13) to obtain the corresponding low-dimensional features: Z r =Sig enc (W r X r +b r )Formula (13) Among them, Z r W represents the initial low-dimensional features. r Let b be the weight matrix of the encoder. r Sig is the bias vector of the encoder. enc () is the multi-layer Sigmoid activation function of the encoder.

16. The method according to claim 14 or 15, characterized in that, The step of inputting the initial low-dimensional features into the feature adjustment unit for feature adjustment includes first determining the feature weights corresponding to the low-dimensional features, and then adjusting the initial low-dimensional features to the target low-dimensional features corresponding to the feature weights, including: Obtain the initial weight matrix for the low-dimensional features; The low-dimensional features are input into the feature adjustment unit, and the initial low-dimensional features are adjusted by the initial weight matrix to obtain the corresponding target low-dimensional features; The initial weight matrix is ​​as follows: A r =diag(a r ) Among them, A r The initial weight matrix; α r Let α be the feature weight vector. r Each element α in r,i Initializing to the same value indicates that all features have the same importance in the initial stage; diag() is a function to extract the diagonal elements of the matrix; The target low-dimensional features are represented as follows: Z′ r =A r From r Among them, Z′ r The target low-dimensional features are adjusted by feature weights.

17. The method according to claim 14 or 15, characterized in that, The recursive optimization of the target low-dimensional features, in each iteration, adjusts the feature weights of each low-dimensional feature based on the performance of each low-dimensional feature in the previous iteration to obtain the target feature weights corresponding to each low-dimensional feature, includes: The target low-dimensional features are recursively optimized, and in the t-th iteration, the feature weights are updated using formula (14): in, This represents the feature weights in the (t+1)th iteration. η represents the feature weights in the t-th iteration. r L is the learning rate for the feature dimensionality reduction model. r () represents the loss function of the feature reduction model, Y r For label data, This represents the gradient of the loss function with respect to the feature weights; The loss function of the feature dimensionality reduction model is the reconstruction error loss function.

18. The method according to claim 14 or 15, characterized in that, The step of inputting the target low-dimensional features and the corresponding target feature weights into the decoder for reconstruction to obtain second feature data corresponding to the target low-dimensional features includes: Obtain the weight matrix and bias vector corresponding to the decoder; The target low-dimensional feature, the weight matrix, and the bias vector are processed according to the multi-layer activation function configured in the decoder to construct the second feature data corresponding to the target low-dimensional feature.

19. The method according to claim 1, characterized in that, Also includes: Initialize a random forest classifier consisting of multiple decision trees, each decision tree including several nodes; During the construction of each decision tree, the quantum entropy corresponding to the node is obtained; The quantum entropy is used to calculate the information gain for the decision tree, and the nodes are split according to the information gain to determine the result growth tree structure corresponding to the decision tree; The resulting growth tree structure is compared with the preset real structure, and the decision tree is iterated based on the comparison results until the preset stopping iteration condition is met, so as to obtain the trained decision tree. Integrate all trained decision trees to obtain the final random forest classifier.

20. The method according to claim 19, characterized in that, The process of obtaining the quantum entropy corresponding to the node includes: The quantum entropy corresponding to the node is calculated using formula (14): Where S represents the dataset of the current node; f is the segmentation feature considered; θ is the segmentation threshold; p k S represents the probability of the k-th label in S; K is the total number of categories. It is the partitioned subset; λ u It is a balance coefficient used to adjust the weights of the segmentation uniformity and information gain.

21. The method according to claim 19, characterized in that, The calculation of information gain for the decision tree using the quantum entropy includes: The information gain for the decision tree is calculated using formula (15): Among them, IG u Information gain reflects the degree of entropy reduction after segmenting nodes using feature f.

22. The method according to claim 19, characterized in that, The iteration of the decision tree based on the comparison results includes: In each iteration, the growth strategy between the decision trees is adjusted using formula (16): in, N represents the parameters of the i-th tree; T It is the total number of trees; The utility function of the i-th decision tree depends on its own parameters and the parameters θ of the other decision trees. -i ;Γ u This represents the optimization function used to maximize the overall synergistic utility of the random forest.

23. A device for recognizing voice emotion, characterized in that, include: The data acquisition module is used to acquire target speech data, as well as an emotion recognition model for speech emotion recognition and a data augmentation model for data augmentation. The emotion recognition model includes at least a feature extraction model, a feature dimensionality reduction model, and a classifier. The data augmentation model is a model trained based on dynamic symmetry breaking theory. The feature extraction model is a model trained based on a neural network algorithm with dynamic population evolution optimization. The feature dimensionality reduction model is a model trained based on a feature refinement-based feature dimensionality reduction model. The classifier is a classifier trained based on a random forest classifier with quantum entropy encoding. The feature extraction module is used to input the target speech data into the feature extraction model to extract features and obtain the speech features corresponding to the target speech data. The feature dimensionality reduction module is used to input the speech features into the feature dimensionality reduction model to perform feature dimensionality reduction and obtain the low-dimensional features corresponding to the speech features. The prediction module is used to input the low-dimensional features into the classifier for prediction to obtain the sentiment category of the speech data; The device further includes: The model acquisition module is used to acquire the data augmentation model to be trained. The data augmentation model includes at least a generator and a discriminator, and the data augmentation model is configured with a symmetry breaking stimulus term. The first data determination module is used to determine the first target speech data for training the data augmentation model; The first parameter generation module is used to randomly generate the first network parameters corresponding to the generator; The second parameter generation module is used to randomly generate the second network parameters corresponding to the discriminator; The loss function calculation module is used to alternately train the generator and the discriminator using the first target speech data, the first network parameters and the second network parameters to obtain the generator loss function corresponding to the generator and the discriminator loss function corresponding to the discriminator. The first iteration module is used to iterate the parameters of the generator using the generator loss function and the symmetry breaking incentive term until a preset stopping iteration condition is met, thereby obtaining the target generator. The first processing module is used to iterate the parameters of the discriminator using the discriminator loss function until a preset stopping iteration condition is met, thereby obtaining the target discriminator. The model determination module is used to combine the target generator and the target discriminator to obtain a data augmentation model; Specifically, the first iteration module is used for: The symmetry breaking excitation term is calculated using formula (3): Δ c = λ c · (||G c (z c ) - x c ||² - γ c ) Equation (3) The parameters of the generator are updated using formula (4): Where, Δ c Ω represents the symmetry breaking excitation term; c λ is the gradient adjustment function; c It is the adjustment coefficient for symmetry breaking excitation; γ c It is a preset symmetry breaking threshold; ||G c (z c )-x c ||2 is the Euclidean distance between the generated data and the real data; ||||2 is the L2 norm, calculated in the same way as the Euclidean distance. It is the learning rate of the generative adversarial network; Indicates to The gradient.

24. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes a program stored in the memory, it implements the method as described in any one of claims 1-22.

25. A computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the method as described in any one of claims 1-22.

Citation Information

Patent Citations

  • Adversarial sample detection method and device, electronic equipment and medium

    CN112329837A

  • Wheat powdery mildew spore segmentation method for small sample image data set

    CN112862792A