Voice emotion recognition model training method and device, equipment and medium

By preprocessing, feature extraction, initial mapping, feature adjustment and enhancement, as well as fine-grained reward classification and strategy optimization of the speech dataset, the shortcomings of existing speech emotion recognition technology in accuracy and stability are solved, and more efficient and accurate speech emotion recognition is achieved.

CN120656491APending Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511052029.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing speech emotion recognition technology lacks accuracy and anti-interference capabilities in emotion recognition, and is unable to cope with complex emotional features and overlapping and fuzzy features, resulting in limited effectiveness in practical applications.

Method used

By obtaining a speech dataset for preprocessing and feature extraction, using a pre-built training model for initial mapping, and through feature adjustment and enhancement, combined with fine-grained reward classification and strategy optimization, the accuracy and stability of the model are improved.

Benefits of technology

It improves the accuracy and stability of the speech emotion recognition model in complex scenarios, enhances the model's ability to capture subtle differences in speech, and improves the accuracy and generalization of sentiment analysis and speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656491A_ABST
    Figure CN120656491A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision making, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a voice emotion recognition model training method, device and equipment and a medium, and the method comprises the steps: obtaining a voice data set, carrying out the preprocessing and feature extraction of the voice data set, and obtaining key features; performing initial mapping on the key features by using a pre-constructed training model to obtain an initial mapping result; performing feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; performing fine-grained reward classification on the feature enhancement result to obtain a fine-grained reward classification result; and performing strategy optimization on the training model according to a fine-grained reward classification result to obtain an optimization model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent decision-making technology, and in particular to a speech emotion recognition model training method, device, equipment and medium. Background Art

[0002] Speech emotion recognition refers to the process of collecting and analyzing human speech signals to extract the emotional characteristic information contained therein, such as the ups and downs of tone, changes in speaking speed, volume, and subtle differences in timbre, and then judging and classifying the emotional state expressed in the speech (such as joy, anger, sadness, fear, surprise, calmness, etc.) based on these characteristics. Its core is to mine emotional information from speech signals and achieve accurate identification of emotion categories.

[0003] In the field of financial technology, voice emotion recognition technology can be used in customer service and fraud detection. By analyzing the emotional state of customers during telephone conversations, financial institutions can more effectively identify potential fraudulent behavior, while improving customer service experience and ensuring transaction security and customer satisfaction.

[0004] In the field of healthcare, speech emotion recognition technology can be applied to mental health assessment and monitoring. By analyzing the patient's voice characteristics, such as tone, speaking speed and volume, it can identify the patient's possible mental health problems, such as depression or anxiety, and provide doctors with non-invasive, real-time mental health assessment methods.

[0005] In summary, existing speech emotion recognition technology still has significant shortcomings in real-world scenarios, with insufficient accuracy and robustness, significantly limiting its effectiveness in practical applications. On the one hand, its ability to capture subtle emotional shifts in speech is limited, making it difficult to cope with the complexity and abstract nature of emotion discrimination. On the other hand, existing technology is unable to accurately identify and clearly distinguish complex emotional characteristics, as well as the potential overlap and ambiguity between different emotions.

[0006] Therefore, there are problems with accuracy and stability in speech emotion recognition in the existing technology that need to be solved urgently. Summary of the Invention

[0007] The present invention provides a speech emotion recognition model training method, device, equipment and medium to solve the accuracy and stability problems of speech emotion recognition in complex scenarios, thereby improving model performance.

[0008] In a first aspect, a method for training a speech emotion recognition model is provided, comprising: Acquire a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features; Performing initial mapping on the key features using a pre-built training model to obtain an initial mapping result; Performing feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; Performing fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; The training model is subjected to strategy optimization according to the fine-grained reward classification result to obtain an optimized model.

[0009] In a second aspect, a speech emotion recognition model training device is provided, comprising: An extraction module is used to obtain a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features; A mapping module, configured to perform initial mapping on the key features using a pre-built training model to obtain an initial mapping result; An enhancement module, configured to perform feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; a classification module, configured to perform fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; An optimization module performs strategy optimization on the training model according to the fine-grained reward classification result to obtain an optimized model.

[0010] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned speech emotion recognition model training method are implemented.

[0011] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned speech emotion recognition model training method are implemented.

[0012] In the solution implemented by the above-mentioned speech emotion recognition model training method, device, equipment and medium, by obtaining a speech data set and performing preprocessing and feature extraction stages, it is possible to filter out irrelevant noise from massive amounts of raw speech data, unify the data format, and accurately locate key features, laying a solid foundation for subsequent processing, ensuring that the information input to the model is pure and valuable, and improving the efficiency and accuracy of model learning. Next, the training model is used to perform an initial mapping of key features, preliminarily establish a connection between speech features and target outputs, and form an initial mapping result. This step is the starting point of model learning and provides an iterative benchmark for subsequent optimization. The initial mapping result is then subjected to feature adjustment and enhancement to further explore potential useful information and strengthen the performance of key features, so that the model can more keenly capture subtle differences and important patterns in speech, and enhance the representativeness and discriminability of the enhanced feature results. Fine-grained reward classification of enhanced feature results can accurately distinguish the categories or effects corresponding to different speech features, assign reasonable reward weight scores to different features, optimize the training model strategy based on the fine-grained reward classification results, and adjust the model parameters and learning strategies in a targeted manner, so that the model can more accurately process speech tasks in subsequent applications, such as speech recognition and sentiment analysis, thereby improving the overall model performance and generalization ability, better adapting to complex and changing speech application scenarios, and achieving efficient and accurate speech processing and analysis goals. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0014] Figure 1 This is a schematic diagram of an application environment of a speech emotion recognition model training method according to an embodiment of the present invention; Figure 2 This is a flow chart of a method for training a speech emotion recognition model according to an embodiment of the present invention; Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S2; Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S5; Figure 5 This is a structural diagram of a speech emotion recognition model training device according to an embodiment of the present invention; Figure 6 is a structural diagram of a computer device in one embodiment of the present invention; Figure 7 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0016] The embodiment of the present invention provides a speech emotion recognition model training method, which can be applied to Figure 1 In an application environment, the client communicates with the server through a network. The server can obtain a speech data set, preprocess and extract features on the speech data set to obtain key features; use a pre-built training model to perform initial mapping on the key features to obtain initial mapping results; perform feature adjustment and enhancement on the initial mapping results to obtain enhanced feature results; perform fine-grained reward classification on the enhanced feature results to obtain fine-grained reward classification results; and optimize the strategy of the training model based on the fine-grained reward classification results to obtain an optimized model. The present invention provides a speech emotion recognition model training device, which can screen out key features from the original speech data through preprocessing and feature extraction for target result business, remove irrelevant information, lay the foundation for subsequent steps, and improve data quality and processing efficiency. The initial mapping allows the model to initially learn the relationship between features and targets, while feature adjustment and enhancement further strengthen the key features, so that the model can grasp the characteristics of speech data more accurately. Fine-grained reward classification can accurately distinguish different categories, providing a clear direction for model optimization. Optimizing the model based on this result can significantly improve model performance, enabling it to perform better in tasks such as speech recognition and sentiment analysis, significantly enhancing accuracy and generalization capabilities, and better adapting to various speech scenarios and task requirements. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented as a standalone server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0017] See also Figure 2 As shown, Figure 2 A flow chart of a method for training a speech emotion recognition model provided by an embodiment of the present invention includes the following steps: S1. Acquire a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features.

[0018] In this context, a speech dataset refers to a collection of collected, organized, and annotated speech signals, typically originating from diverse speakers, environments, and language backgrounds. Each data point in a speech dataset typically consists of an audio signal and its corresponding label information, such as a text transcription (for speech recognition), an emotion category (for emotion recognition), or other relevant annotations.

[0019] In an embodiment of the present invention, the step of obtaining a speech data set, preprocessing and extracting features from the speech data set, and obtaining key features includes: performing noise processing on the speech signal in the speech data set to obtain a preliminary speech signal; performing normalization processing on the preliminary speech signal to obtain a preprocessed speech signal; The key features of the preprocessed speech signal are extracted using a time-frequency joint analysis technique.

[0020] In an embodiment of the present invention, noise processing of the speech signal in the speech dataset to obtain a preliminary speech signal involves identifying the noise frequency components and distribution characteristics in the speech signal through spectral analysis. Based on the spectral analysis, the power spectral density of the noise is estimated, and a filtering algorithm (such as Wiener filtering) is then employed to attenuate the noise frequency band, thereby separating the pure speech component from the noisy speech, thereby obtaining a preliminary speech signal with reduced noise interference and more amenable to subsequent processing. The Wiener filtering calculates the frequency response of the Wiener filter based on the estimated noise power spectral density and the power spectral density of the speech signal. The Wiener filter attenuates the noise frequency band by minimizing the mean square error of the error signal, thereby enhancing the useful components of the speech signal.

[0021] In an example of the present invention, the normalization processing of the preliminary speech signal to obtain a preprocessed speech signal refers to performing statistical analysis on the amplitude of the preliminary speech signal, calculating its characteristic parameters such as average energy and peak amplitude, and then using an amplitude normalization method (such as linearly scaling the signal amplitude to the interval [-1,1]) to obtain a preprocessed speech signal with uniform amplitude distribution and uniform energy level.

[0022] For example, in the fintech sector, financial institutions' intelligent customer service systems often need to process a large number of customer voice inquiries. These voice signals may originate from different devices and environments, and may vary in volume, pitch, and background noise. By normalizing the initial voice signals, voice signals from different sources can be aligned to a uniform amplitude range and frequency characteristics, thereby improving speech recognition accuracy and system response speed.

[0023] In an embodiment of the present invention, the method of extracting key features of the preprocessed speech signal using a time-frequency joint analysis technique includes: Extracting Mel-frequency cepstral coefficients and frequency domain features of the preprocessed speech signal using feature parsing and frequency domain analysis techniques; Performing frame division and computation processing on the preprocessed speech signal to obtain time domain features; The Mel-frequency cepstral coefficients, the frequency domain features, and the time domain features are concatenated to obtain key features.

[0024] In an embodiment of the present invention, the Mel-frequency cepstral coefficients and frequency domain features of the preprocessed speech signal are extracted using feature parsing and frequency domain analysis techniques. The preprocessed speech signal is first framed and windowed to convert it into a short-term stationary signal. Each frame of the signal is converted from the time domain to the frequency domain by fast Fourier transform to obtain a spectrum. The spectrum is filtered using a Mel filter bank to simulate the human ear's perception of sound frequency. The linear frequency is converted into Mel frequency. The energy after filtering is logarithmized and discrete cosine transformed to obtain the Mel-frequency cepstral coefficients. The Mel-frequency cepstral coefficients are a feature vector extracted from the speech signal and used to characterize the spectral characteristics of the speech. Statistics such as the spectral centroid, spectral bandwidth, and spectral flux, i.e., frequency domain features, are calculated in the frequency domain.

[0025] In an example of the present invention, the framing and calculation processing of the pre-processed speech signal to obtain time domain features refers to framing the pre-processed speech signal, dividing it into multiple short-time frames, calculating the short-time energy of each frame, that is, the mean of the sum of the squares of the signal amplitudes, and simultaneously counting the number of times the signal waveform of each frame crosses the zero level to obtain the short-time zero-crossing rate, which is used to distinguish between unvoiced and voiced sounds. The short-time energy is a measure of the strength of the frame signal by calculating the sum of the squares of all sampling points in each frame, which can reflect the activity level of the speech signal. The short-time zero-crossing rate is a count of the number of times the signal waveform in each frame crosses the zero point, which is used to distinguish between unvoiced sounds (high zero-crossing rate) and voiced sounds (low zero-crossing rate).

[0026] In an embodiment of the present invention, the key features are obtained by concatenating the Mel-frequency cepstral coefficients, the frequency-domain features, and the time-domain features. These feature vectors are concatenated end-to-end in a predetermined order (e.g., Mel-frequency cepstral coefficients first, then frequency-domain features, and finally time-domain features) to form a composite feature vector of higher dimension. For example, if the Mel-frequency cepstral coefficients are 13-dimensional, the frequency-domain features are 10-dimensional, and the time-domain features are 5-dimensional, then the dimension of the concatenated key features is 28-dimensional.

[0027] In the examples of the present invention, by collecting a diverse set of speech data, the universality of subsequent processing is ensured, and preprocessing can eliminate interference such as environmental noise and channel distortion, making the speech signal purer and laying a good foundation for feature extraction. Secondly, the feature extraction process transforms the original speech signal into more representative key features through the fusion of Mel-frequency cepstral coefficients, frequency domain features, and time domain features. These features can not only retain the essential information of the speech (such as spectral envelope, pitch, energy changes), but also remove redundancy, while enhancing the model's robustness to scenarios such as noise and changes in speech speed, ultimately achieving efficient conversion from raw speech data to effective information.

[0028] S2. Use a pre-built training model to perform initial mapping on the key features to obtain an initial mapping result.

[0029] In the examples of the present invention, the training model is a deep learning-based model used to perform initial mapping of key features and generate initial mapping results. The model is usually composed of multiple layers or modules, including an input layer, a hidden layer (such as a fully connected layer, a convolutional layer or a recurrent layer) and an output layer. By learning the inherent patterns and regularities in a large amount of labeled data, the original high-dimensional feature vector can be mapped to a new feature space. In this new space, the representation of features is more compact and more discriminative, which can better retain key information while removing redundant or noise components. The initial mapping results are feature vectors generated after preliminary processing of the original data (such as speech signals, images or other multidimensional data). These feature vectors have undergone steps such as format conversion, feature encoding and dimensionality reduction, thereby reducing the complexity and dimensionality of the data while retaining key information.

[0030] Examples of the present invention, see Figure 3 As shown, the key features are initially mapped using the pre-built training model to obtain the initial mapping results, including: S21, converting the key features into a format according to a preset format rule to obtain a converted feature; S22, performing feature encoding processing on the converted features to obtain abstract feature representation; S23. Perform dimensionality reduction processing on the abstract feature representation to obtain an initial mapping result.

[0031] In the examples of the present invention, the format conversion of the key features according to the preset format rules to obtain the conversion features refers to matching the key feature dimensions with the target dimension values ​​in the preset rules. If the dimension of the key feature is smaller than the target dimension, it can be filled, for example, by adding zero values ​​or other preset values ​​at the end of the feature vector until its dimension meets the target requirement; if the dimension of the key feature is larger than the target dimension, it is necessary to trim it, that is, remove the part of the feature vector that exceeds the target dimension, usually choosing to trim redundant or relatively unimportant feature components. Finally, the conversion features that meet the preset format rules are obtained.

[0032] In the example of the present invention, the feature encoding processing of the conversion feature to obtain an abstract feature representation is to input the conversion feature into a multi-layer coding network. In the convolution layer, according to the set convolution kernel size and number, the convolution kernel is allowed to slide on the two-dimensional or multi-dimensional space of the input feature, and the local feature pattern (such as a specific frequency combination in the speech signal, etc.) is extracted by calculating the dot product operation between the convolution kernel parameters and the corresponding input feature area. Then, the pooling layer is entered, and the maximum pooling, average pooling, etc. are used. According to the preset pooling window (such as 2×2) and step size, the feature map output by the convolution layer is downsampled to reduce the spatial dimension of the feature while retaining the key local feature information. For sequential conversion features, in the recurrent layer, the memory mechanism of the recurrent unit is used to process the sequence data in sequence according to the time step. The transmission and update of information are controlled by the gating structure (input gate, forget gate, output gate, etc.) to capture the time dependency in the sequence data, such as the changing trend of speech emotions over time. After the combined encoding operation of multiple layers of convolution, pooling, and recurrent layers, the middle layer of the network will fuse the local and global features extracted by each layer, and finally generate a feature vector of fixed dimension, that is, the abstract feature representation. This vector integrates the complex structure and semantic information of the input features, providing a more representative feature basis for subsequent model processing.

[0033] In the examples of the present invention, the dimensionality reduction processing of the abstract feature representation to obtain the initial mapping result is performed by converting the high-dimensional abstract feature representation into a low-dimensional feature representation through the encoder in the autoencoder, and gradually compressing the input features into a low-dimensional space through a nonlinear activation function. The autoencoder consists of two parts: an encoder and a decoder. The encoder is responsible for mapping the input high-dimensional abstract feature representation to a low-dimensional feature space to generate a low-dimensional encoded representation; the decoder attempts to reconstruct this low-dimensional encoded representation back into the original high-dimensional feature representation.

[0034] For example, in the healthcare field, during mental health monitoring, a patient's voice signals can be used to identify emotional states such as anxiety, depression, or stress. First, key features are extracted from the patient's voice data. These features are then processed through feature encoding to generate high-dimensional abstract feature representations. However, high-dimensional features may lead to increased computational complexity and overfitting problems. Therefore, through dimensionality reduction processing, such as using an autoencoder, the abstract feature representation is mapped to a low-dimensional space to obtain an initial mapping result. The autoencoder compresses high-dimensional features into low-dimensional features through the encoder part while retaining key information. The initial mapping result after dimensionality reduction not only reduces the data dimension but also improves the training efficiency and generalization ability of the model.

[0035] In the examples of the present invention, by inputting key features into the training model for initial mapping, the original features can be converted into higher-level abstract feature representations. These abstract features are often more discriminative and robust, and can better adapt to different task requirements.

[0036] S3. Perform feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result.

[0037] In this example, the enhanced feature result is a data structure that combines model performance feedback and feature representation. It is obtained by concatenating the classification accuracy reward with a feature vector that has been formatted and constrained. This concatenation not only preserves the original feature information but also incorporates the model's classification accuracy evaluation on a specific task, allowing the enhanced feature result to directly reflect the model's performance.

[0038] In an embodiment of the present invention, the feature adjustment and enhancement of the initial mapping result to obtain an enhanced feature result includes: Performing format constraints on the initial mapping result to obtain a format constraint result; Performing classification accuracy reward on the format constraint result to obtain a classification accuracy feature result; Perform vector concatenation on the vector corresponding to the classification accuracy feature result and the format constraint result to obtain an enhanced feature result.

[0039] In this embodiment of the present invention, the format constraint applied to the initial mapping result to obtain the format constraint result is performed by checking the initial mapping result against a preset format specification (e.g., a specific output format) to determine whether it meets the preset format requirements. If the initial mapping result meets the format specification, it is marked as valid and proceeds to the next step of processing; if not, the result is discarded.

[0040] In an embodiment of the present invention, the classification accuracy reward for the format-constrained results is obtained by first performing a linear transformation and nonlinear mapping on the format-constrained results in a high-dimensional space to explore the complex correlations between features. The transformed features are then converted into probability distributions of various emotion and state labels. The training model analyzes the probability distribution based on the corresponding patterns and regularities between speech features and emotion and state labels learned during the training phase, selects the category with the highest probability, and obtains the predicted emotion and state labels. The predicted emotion and state labels are then compared with the actual emotion and state labels (e.g., pre-labeled real emotions and corresponding states such as "anger"). If the two are completely consistent (e.g., if the actual label is "anger," the model's output label after format constraints is also "anger"), the classification is determined to be accurate, and a 1-point classification accuracy reward is directly awarded to reinforce the recognition feedback of the emotion category and state. If they are inconsistent, the classification logic can be subsequently adjusted to optimize the output through calculation errors such as cross-entropy loss. Finally, a reward score, i.e., the classification accuracy feature result, is generated based on the comparison results.

[0041] The enhanced feature result is obtained by concatenating the vector corresponding to the classification accuracy feature result with the format constraint result. This converts the classification accuracy feature result into a vector of dimension M (e.g., M=1), which carries quantitative feedback on the model's classification performance (e.g., a 1-point reward). The vector corresponding to the classification accuracy feature result is then added to the format-constrained feature vector to form a new, expanded feature vector, the enhanced feature result. This new vector not only contains the original feature information but also incorporates direct feedback on model performance, allowing it to be used for further model training or optimization.

[0042] In this example, format constraints and classification accuracy rewards are applied to the initial mapping results to generate enhanced feature results. This significantly improves the standardization and reliability of model output, enabling the model to make accurate decisions in complex scenarios. Format constraints ensure that model output adheres to uniform standards, facilitating subsequent data processing and analysis. Classification accuracy rewards, through quantitative feedback, encourage the model to continuously optimize its feature recognition capabilities and enhance prediction accuracy.

[0043] For example, in the field of financial technology, when a customer communicates with an intelligent customer service representative via voice, the system first performs feature extraction and initial mapping on the collected voice signal to obtain the initial mapping result. At this time, the format constraint requires that the output must follow a specific format. If the model output conforms to this format, a format reward can be obtained, which ensures the standardization of the recognition results and facilitates the system to quickly parse and store them. The output emotion and intensity labels are compared with the real labels marked by humans. If the customer's "anxiety" emotion caused by account anomalies is accurately identified and the intensity is "high", a classification accuracy reward is given to strengthen the model's recognition ability of this type of emotional features. Finally, the classification accuracy reward result and the format constraint result are vectorized and spliced ​​to obtain the enhanced feature result.

[0044] S4. Perform fine-grained reward classification on the enhanced feature result to obtain a fine-grained reward classification result.

[0045] In this example, the fine-grained reward classification method breaks down the model's classification accuracy reward into multiple dimensions, such as the positive or negative emotion and the specific state, to provide more specific feedback. By calculating the positive or negative labeling results and the state classification results and concatenating them with the initial enhanced feature result vector, a more informative fine-grained reward classification result is generated.

[0046] In an embodiment of the present invention, the fine-grained reward classification is performed on the enhanced feature result to obtain a fine-grained reward classification result, including: Performing emotion positive and negative labeling on the enhanced feature results according to a predefined positive and negative emotion dictionary to obtain a positive and negative labeling result; Performing emotional state classification on the enhanced feature result according to a predefined emotional state evaluation system to obtain a state classification result; Vector concatenation is performed on the positive and negative labeling results, the state classification results, and the enhanced feature results to obtain a fine-grained reward classification result.

[0047] In an example of the present invention, the positive and negative emotion labeling of the enhanced feature results according to the predefined positive and negative emotion dictionary is performed to obtain the positive and negative labeling results, which is to extract the emotion label from the structured information of the enhanced feature results, and then match the label with the predefined positive and negative emotion dictionary (such as negative emotion words such as "anger" and "anxiety", and positive emotion words such as "satisfaction" and "pleasure"). If the emotion label belongs to the negative emotion vocabulary set, it is assigned a score in the range of 0-0.1 (mild), 0.2-0.3 (moderate), and 0.4 (strong) according to the preset negative emotion intensity level (such as mild, moderate, and strong); if it belongs to the positive emotion vocabulary set, it is assigned a score in the range of 0.6-0.7 (mild), 0.8-0.9 (moderate), and 1.0 (strong) according to the positive emotion positivity level; if it cannot match any dictionary word, or the emotion is neutral, it is assigned 0.5 points. The final value is the positive and negative labeling result. This score quantitatively reflects the positive and negative attributes and intensity of the emotions in the enhanced feature results, providing a quantitative basis for subsequent more detailed sentiment analysis and model optimization.

[0048] In this embodiment of the present invention, the emotional state classification of the enhanced feature results based on a predefined emotional state assessment system is performed to obtain a state classification result. This involves extracting an emotional state description from the structured information of the enhanced feature results. Then, based on the predefined emotional state assessment system (for example, in a financial customer service scenario, customer emotional states can be categorized into broad categories such as consultation, complaint, and suggestion, further subdivided into subcategories such as urgent complaint and general consultation), the current emotional state category is determined and assigned an emotional state category score. Scoring criteria are set for each category based on its importance to the business process, processing priority, or complexity. For example, "urgent complaint" might be assigned a score of 0.9-1, while "general consultation" might be assigned a score of 0.1-0.3. If the emotional state cannot be accurately classified into a predefined category, a fuzzy matching algorithm (such as calculating semantic similarity) is used based on its similarity to each standard category. Interpolation and other methods are used to determine an intermediate score. The final output value is the state classification result. This result effectively measures the value and impact of emotional states in real business scenarios, providing feedback that is more tailored to practical needs for model strategy optimization.

[0049] In an example of the present invention, the positive and negative labeling results and the state classification results are vector-concatenated with the enhanced feature results to obtain a fine-grained reward classification result. The positive and negative labeling results and the state classification results are first converted into single-dimensional vectors (e.g., the positive and negative labeling result 0.8 is converted to [0.8], and the state classification result 0.7 is converted to [0.7]), and then, according to a preset vector concatenation rule, these two single-dimensional vectors are combined with the enhanced feature result vector containing format constraints and classification accuracy information. Usually, the positive and negative labeling result vector and the state classification result vector are appended to the end of the enhanced feature result vector in sequence. For example, if the enhanced feature result vector is [1,2,3], the concatenated vector is [1,2,3,0.8,0.7], thereby forming a new feature vector. This new vector fully integrates the structural information of the original enhanced feature result, the quantitative evaluation of the positive and negative emotions, and the business value measurement of the emotional state, fully presents the multi-dimensional attributes of the speech emotion characteristics, and finally generates a fine-grained reward classification result.

[0050] In the examples of the present invention, the benefit of fine-grained reward classification of enhanced feature results is that it can more comprehensively reflect the complexity of emotions and states. By combining the positive and negative labeling of emotions and the state classification results, it is possible to identify the positive or negative tendency of emotions, and further refine the specific states of emotions, such as "anger", "sadness", "excitement", etc. This fine-grained classification method provides the model with richer feedback information. Through this fine-grained reward mechanism, the intelligent customer service system can more accurately identify customer emotions. For example, when a user applies for a loan, it can promptly perceive the customer's dissatisfaction with the interest rate policy, automatically transfer the customer to manual service and push soothing words to improve the customer experience. At the same time, it can prevent potential service complaint risks for financial institutions and optimize service strategies.

[0051] S5. Optimize the strategy of the training model according to the fine-grained reward classification result to obtain an optimized model.

[0052] In this example, the policy optimization employs a reinforcement learning approach, improving model performance through fine-grained reward classification. This process uses a group relative policy optimization algorithm to perform multi-output sampling on input samples and calculate relative rewards within the group to mitigate the impact of differences in reward scale. These rewards are then used to construct a policy gradient, pushing model parameters toward higher rewards. Furthermore, a KL divergence regularization term is used to control the magnitude of policy updates, ensuring a stable training process. Through continuous iteration until the KL divergence constraint is satisfied, the policy update is complete, and the optimized model is obtained.

[0053] Examples of the present invention, see Figure 4 As shown, the strategy optimization of the training model is performed based on the fine-grained reward classification result to obtain an optimized model, including: S51, calculating the relative reward value within the group of the key feature according to the fine-grained reward classification result; S52, performing strategy update and divergence constraint on the training model according to the relative reward value within the group; S53. If the divergence constraint is greater than a preset threshold, return to the step of calculating based on the fine-grained reward classification result; if the divergence constraint is not greater than the preset threshold, stop iterative updating to obtain the optimized model.

[0054] In an example of the present invention, the calculation of the relative reward value within the group of the key feature based on the fine-grained reward classification result is combined with a group relative strategy optimization algorithm. When calculating the relative reward value within the group of the key feature based on the fine-grained reward classification result, for the key feature, multiple emotion prediction outputs (such as 16) are first sampled from the training model. For each output, first, the positive and negative labeling results are determined based on the matching of the identified emotion category in the predefined emotion dictionary, which reflects the positive and negative attributes and intensity of the emotion; then, the state classification result is determined based on the specific state of the emotion, which reflects the recognition accuracy of the emotion state; finally, the two results are weighted and summed to obtain the total fine-grained reward score for each output.

[0055] Sampling refers to the process of randomly selecting a subset of all possible emotion prediction outputs generated by a trained model for a specific input (e.g., key features) for analysis. For example, for each input sample, the model may produce multiple different emotion predictions. Sampling involves selecting a certain number (e.g., 16) of these outputs for subsequent processing.

[0056] Furthermore, the average of the total fine-grained reward scores of these outputs is calculated, and the total fine-grained reward score of each individual output is subtracted from the average fine-grained reward score to calculate the advantage value of each output relative to the average value within the group. The advantage value represents the performance of a single output relative to other outputs in the group. If the advantage value is positive, it means that the total reward score of the output is higher than the average within the group, and the performance is good; if the advantage value is negative, it means that the total reward score of the output is lower than the average within the group, and the performance is poor. The advantage value reflects the relative performance of a single output within the group. In this way, the scale difference caused by fluctuations in the absolute value of the reward can be eliminated. Ultimately, these advantage values ​​serve as relative reward values ​​within the group, which can accurately reflect the relative value of key features within the group.

[0057] In an example of the present invention, the strategy update and divergence constraint of the training model based on the relative reward value within the group is performed. When the relative reward value within the group is substituted into the clipping objective function of the proximal policy optimization algorithm to calculate the policy gradient, this gradient reflects the direction and degree of difference between the model output behavior (such as emotion recognition results) and the expected behavior (the result corresponding to the high relative reward value within the group) under the current parameters. Based on this, the gradient descent algorithm is used to adjust the parameters such as the convolution kernel weights of the convolution layer, the gated unit weights of the recurrent layer, and the connection weights of the fully connected layer in the multi-layer network in the direction of maximizing the relative reward value within the group, so as to change the way the model extracts and maps key features of the input speech. KL divergence is used to measure the difference in the output distribution of the new and old strategies. After incorporating it into the objective function as a regularization term, the parameter adjustment range will be limited during each parameter update. If the distribution difference between the new strategy and the old strategy is too large (that is, the KL divergence exceeds a certain threshold), the parameter update will be suppressed to avoid performance fluctuations or overfitting due to excessive adjustment of the model, and ensure that the model is gradually optimized within a stable parameter space through the iterative process of repeatedly calculating gradients, updating parameters, calculating KL divergence and adjusting.

[0058] In an example of the present invention, if the divergence constraint is greater than a preset threshold, the step of returning to the calculation based on the fine-grained reward classification result is returned; if the divergence constraint is not greater than the preset threshold, the iterative update is stopped to obtain the optimized model. After each update, it is determined whether the divergence constraint is greater than the preset threshold. If the divergence constraint is greater than the preset threshold, it means that the model update is large and needs to be optimized further. Therefore, the step of calculating the relative reward value within the group is returned to re-update the strategy. If the divergence constraint is not greater than the preset threshold, it means that the model has reached a relatively stable state. At this time, the iterative update is stopped, and the optimized model is finally obtained. This method ensures that the model update can respond to the learning needs of key features while maintaining the smoothness and stability of the update process, thereby improving the stability and generalization ability of the model.

[0059] In the example of the present invention, the strategy optimization of the training model is performed based on the results of fine-grained reward classification. The benefit of the optimized model is that it can significantly improve the performance and adaptability of the model on specific tasks. Through fine-grained reward classification, the model can receive specific feedback on its performance in recognizing different emotion categories or states, which enables the model to more accurately identify and learn those features that are most critical to improving classification accuracy. The calculation of the relative reward value within the group and the divergence constraint mechanism in the strategy optimization process ensure the stability and smoothness of the model update and prevent training instability caused by drastic changes in parameters. Ultimately, the optimized model not only improves recognition accuracy, but can also be better generalized to new data, thereby providing more reliable and efficient services in practical applications.

[0060] For example, in the field of healthcare, in medical voice diagnosis scenarios, the training model needs to identify emotions (such as anxiety and pain) in the patient's voice to assist in the diagnosis of the disease. Through fine-grained reward classification, the model can distinguish between sub-states such as "anxiety - mild / severe" and "pain - physiological / psychological inducements". During strategy optimization, the relative reward value within the group is used to strengthen the model's learning of "postoperative pain - high-frequency groaning characteristics" and "chronic disease anxiety - low-pitched repetitive demands characteristics". Divergence constraints are used to control parameter updates to ensure that the model can stably recognize the emotional expressions of different patients. The optimized model can more accurately associate emotions with conditions (such as identifying "anxiety accompanied by rapid breathing" to indicate the risk of respiratory diseases), helping doctors to detect potential health problems in patients earlier.

[0061] It can be seen that in the above scheme, for the target result business, a speech data set is obtained, and the speech data set is preprocessed and feature extracted to obtain key features; the key features are initially mapped using a pre-built training model to obtain an initial mapping result; the initial mapping result is feature adjusted and enhanced to obtain an enhanced feature result; the enhanced feature result is fine-grained reward classification to obtain a fine-grained reward classification result; the training model is strategy optimized based on the fine-grained reward classification result to obtain an optimized model.

[0062] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0063] In one embodiment, a speech emotion recognition model training device is provided, which corresponds to a speech emotion recognition model training method in the above embodiment. Figure 5 As shown, the speech emotion recognition model training device includes an extraction module 101, a mapping module 102, an enhancement module 103, a classification module 104 and an optimization module 105. The functional modules are described in detail as follows: Extraction module 101, used to obtain a speech data set, preprocess and extract features from the speech data set to obtain key features; A mapping module 102 is configured to perform an initial mapping on the key features using a pre-built training model to obtain an initial mapping result; An enhancement module 103 is configured to adjust and enhance the features of the initial mapping result to obtain an enhanced feature result; A classification module 104 is configured to perform fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; The optimization module 105 performs strategy optimization on the training model according to the fine-grained reward classification result to obtain an optimized model.

[0064] In one embodiment, the extraction module 101, after obtaining a speech dataset, performs preprocessing and feature extraction on the speech dataset to obtain key features for: performing noise processing on the speech signal in the speech data set to obtain a preliminary speech signal; performing normalization processing on the preliminary speech signal to obtain a preprocessed speech signal; The key features of the preprocessed speech signal are extracted using a time-frequency joint analysis technique.

[0065] In one embodiment, the extraction module 101, after obtaining a speech data set, performs preprocessing and feature extraction on the speech data set to obtain key features, and is further configured to: Extracting Mel-frequency cepstral coefficients and frequency domain features of the preprocessed speech signal using feature parsing and frequency domain analysis techniques; Performing frame division and computation processing on the preprocessed speech signal to obtain time domain features; The Mel-frequency cepstral coefficients, the frequency domain features, and the time domain features are concatenated to obtain key features.

[0066] In one embodiment, the mapping module 102 performs an initial mapping on the key features using a pre-built training model to obtain an initial mapping result, which is used to: Performing format conversion on the key features according to preset format rules to obtain converted features; Performing feature encoding processing on the conversion features to obtain abstract feature representation; Performing dimensionality reduction processing on the abstract feature representation to obtain an initial mapping result.

[0067] In one embodiment, the enhancement module 103 performs feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result for: Performing format constraints on the initial mapping result to obtain a format constraint result; Performing classification accuracy reward on the format constraint result to obtain a classification accuracy feature result; Perform vector concatenation on the vector corresponding to the classification accuracy feature result and the format constraint result to obtain an enhanced feature result.

[0068] In one embodiment, the classification module 104 performs fine-grained reward classification on the enhanced feature result to obtain a fine-grained reward classification result for: Performing emotion positive and negative labeling on the enhanced feature results according to a predefined positive and negative emotion dictionary to obtain a positive and negative labeling result; Performing emotional state classification on the enhanced feature result according to a predefined emotional state evaluation system to obtain a state classification result; Vector concatenation is performed on the positive and negative labeling results, the state classification results, and the enhanced feature results to obtain a fine-grained reward classification result.

[0069] In one embodiment, the optimization module 105 performs strategy optimization on the training model based on the fine-grained reward classification result to obtain an optimized model for: Calculating the relative reward value within the group of the key feature according to the fine-grained reward classification result; Performing strategy updating and divergence constraint on the training model according to the relative reward value within the group; If the divergence constraint is greater than a preset threshold, return to the step of calculating based on the fine-grained reward classification result; if the divergence constraint is not greater than the preset threshold, stop iterative updating to obtain the optimized model.

[0070] The present invention provides a speech emotion recognition model training device. Targeting a target result, the device obtains a speech dataset, preprocesses and extracts features from the dataset to obtain key features. A pre-built training model is used to perform initial mapping on the key features, obtaining an initial mapping result. Feature adjustment and enhancement are performed on the initial mapping result to obtain an enhanced feature result. Fine-grained reward classification is performed on the enhanced feature result to obtain a fine-grained reward classification result. The training model is then strategically optimized based on the fine-grained reward classification result to obtain an optimized model. Preprocessing and feature extraction can filter key features from the raw speech data, remove irrelevant information, and lay the foundation for subsequent steps, improving data quality and processing efficiency. Initial mapping allows the model to initially learn the relationship between features and targets, while feature adjustment and enhancement further strengthen key features, enabling the model to more accurately grasp the characteristics of speech data. Fine-grained reward classification accurately distinguishes different categories, providing a clear direction for model optimization. Model optimization based on this result improves model performance, resulting in superior performance in tasks such as speech recognition and sentiment analysis, significantly enhancing accuracy and stability, and better adapting to various speech scenarios and task requirements.

[0071] For the specific definition of a speech emotion recognition model training device, please refer to the definition of a speech emotion recognition model training method above, which will not be repeated here. The various modules in the above-mentioned speech emotion recognition model training device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0072] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a speech emotion recognition model training method.

[0073] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of a speech emotion recognition model training method. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed: Acquire a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features; Performing initial mapping on the key features using a pre-built training model to obtain an initial mapping result; Performing feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; Performing fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; The training model is subjected to strategy optimization according to the fine-grained reward classification result to obtain an optimized model.

[0074] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Acquire a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features; Performing initial mapping on the key features using a pre-built training model to obtain an initial mapping result; Performing feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; Performing fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; The training model is subjected to strategy optimization according to the fine-grained reward classification result to obtain an optimized model.

[0075] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0076] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0077] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0078] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.

[0079] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A speech emotion recognition model training method, characterized in that: include: Acquire a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features; Performing initial mapping on the key features using a pre-built training model to obtain an initial mapping result; Performing feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; Performing fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; The training model is subjected to strategy optimization according to the fine-grained reward classification result to obtain an optimized model.

2. The speech emotion recognition model training method according to claim 1, wherein The acquiring of a speech data set, preprocessing and feature extraction of the speech data set, and obtaining key features include: performing noise processing on the speech signal in the speech data set to obtain a preliminary speech signal; performing normalization processing on the preliminary speech signal to obtain a preprocessed speech signal; The key features of the preprocessed speech signal are extracted using a time-frequency joint analysis technique.

3. The speech emotion recognition model training method according to claim 2, wherein: The method of extracting key features of the preprocessed speech signal by using a time-frequency joint analysis technique includes: Extracting Mel-frequency cepstral coefficients and frequency domain features of the preprocessed speech signal using feature parsing and frequency domain analysis techniques; Performing frame division and computation processing on the preprocessed speech signal to obtain time domain features; The Mel-frequency cepstral coefficients, the frequency domain features, and the time domain features are concatenated to obtain key features.

4. The speech emotion recognition model training method according to claim 1, wherein The initial mapping of the key features using the pre-built training model to obtain the initial mapping results includes: Performing format conversion on the key features according to preset format rules to obtain converted features; Performing feature encoding processing on the conversion features to obtain abstract feature representation; Performing dimensionality reduction processing on the abstract feature representation to obtain an initial mapping result.

5. The speech emotion recognition model training method according to claim 1, wherein The step of adjusting and enhancing the features of the initial mapping result to obtain an enhanced feature result includes: Performing format constraints on the initial mapping result to obtain a format constraint result; Performing classification accuracy reward on the format constraint result to obtain a classification accuracy feature result; Perform vector concatenation on the vector corresponding to the classification accuracy feature result and the format constraint result to obtain an enhanced feature result.

6. The speech emotion recognition model training method according to claim 1, wherein: The performing fine-grained reward classification on the enhanced feature result to obtain a fine-grained reward classification result includes: Performing emotion positive and negative labeling on the enhanced feature results according to a predefined positive and negative emotion dictionary to obtain a positive and negative labeling result; Performing emotional state classification on the enhanced feature result according to a predefined emotional state evaluation system to obtain a state classification result; Vector concatenation is performed on the positive and negative labeling results, the state classification results, and the enhanced feature results to obtain a fine-grained reward classification result.

7. The speech emotion recognition model training method according to claim 1, wherein: The performing strategy optimization on the training model according to the fine-grained reward classification result to obtain an optimized model includes: Calculating the relative reward value within the group of the key feature according to the fine-grained reward classification result; Performing strategy updating and divergence constraint on the training model according to the relative reward value within the group; If the divergence constraint is greater than a preset threshold, return to the step of calculating based on the fine-grained reward classification result; if the divergence constraint is not greater than the preset threshold, stop iterative updating to obtain the optimized model.

8. A speech emotion recognition model training device, characterized in that: include: An extraction module is used to obtain a speech data set, perform preprocessing and feature extraction on the speech data set, and obtain key features; A mapping module, configured to perform initial mapping on the key features using a pre-built training model to obtain an initial mapping result; An enhancement module, configured to perform feature adjustment and enhancement on the initial mapping result to obtain an enhanced feature result; a classification module, configured to perform fine-grained reward classification on the enhanced feature results to obtain a fine-grained reward classification result; An optimization module performs strategy optimization on the training model according to the fine-grained reward classification result to obtain an optimized model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the speech emotion recognition model training method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech emotion recognition model training method according to any one of claims 1 to 7 is implemented.