Speech Emotion Recognition Method and Device for Non-Performing Asset Disposal
By constructing the Gaussian hybrid model-hidden Markov chain joint model and the noise-emotion misjudgment association matrix, and combining multimodal information for speech emotion recognition, the problem of low speech emotion recognition accuracy under complex background noise is solved, and higher recognition accuracy and disposal efficiency are achieved.
Patent Information
- Application Number
- CN202510449612.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art is difficult to accurately recognize voice emotions under complex background noise, especially in non-performing asset disposal scenarios. The frequency, intensity and duration of noise are complex, resulting in a decrease in the accuracy of voice emotions recognition.
By collecting historical noise samples, Gaussian mixed model-hidden Markov chain joint model and noise-emotion misjudgment correlation matrix are constructed, and the probability of misjudgment of noise on emotions is accurately analyzed, and comprehensive emotion recognition is carried out in combination with multimodal information.
It effectively improves the accuracy of speech emotion recognition under complex background noise, enhances the accuracy of speech emotion recognition in non-performing asset disposal scenarios, helps staff understand customer emotions more accurately, and improves disposal efficiency.
Smart Images

Figure CN119993217B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and more specifically, to a speech emotion recognition method and device for non-performing asset disposal. Background Art
[0002] Speech emotion recognition is of crucial significance for understanding and analyzing human communication. Especially in specific scenarios such as non-performing asset disposal, accurately recognizing the emotions in speech helps staff better grasp the psychological state of the communication object and improve the efficiency and quality of disposal. However, in the actual phone communication scenario of non-performing asset disposal, complex background noises, such as the noisy mechanical sounds in the disposal site and the current interference sounds of the phone line, seriously interfere with the accuracy of speech emotion recognition, posing a huge challenge to existing speech emotion recognition technologies.
[0003] In the prior art, a Chinese patent application with the publication number CN109243492A discloses a speech emotion recognition system and recognition method. Although it increases the detection means for phone fraud systems, can perform multi-dimensional analysis on speech data, and improves the detection accuracy of the system by 5%, this method has limitations in dealing with complex noises in the non-performing asset disposal scenario. In the phone scenario of non-performing asset disposal, the types of noises are diverse and complex. This prior art only performs conventional processing on speech data through a speech preprocessing module, making it difficult to effectively remove the interference of various noises on speech emotion features. Its method of extracting acoustic parameters closely related to emotions in the speech signal cannot accurately capture weak emotion features in a high-noise environment, resulting in a significant impact on the emotion recognition accuracy. For example, in the case of continuous interference from mechanical sounds in the disposal site, the weak emotion features of the target speech may be masked, making it difficult for this method to accurately recognize the emotion type.
[0004] A Chinese patent application with the publication number CN117612569A proposes a speech emotion recognition method, device, equipment, and storage medium, which improves the emotion recognition accuracy by dividing speech signals, performing noise reduction processing on part of the speech, and combining context information. However, in the non-performing asset disposal scenario, this method still has deficiencies. Its preset signal-to-noise ratio prediction model and speech enhancement model are designed based on general scenarios and have poor adaptability to the unique noise characteristics in the non-performing asset disposal scenario. During the non-performing asset disposal process, the characteristics of noises such as frequency, intensity, and duration are complex and diverse. The noise reduction and feature extraction methods of this prior art are difficult to comprehensively and specifically handle these noises, resulting in inaccurate recognition of speech emotions under complex background noises and unable to meet the high-precision requirements of the non-performing asset disposal scenario for speech emotion recognition.
[0005] Currently, in the scenario of non-performing asset disposal, existing speech emotion recognition technologies generally have insufficient ability to extract weak emotion features when dealing with complex background noise, unable to effectively distinguish different emotion types, resulting in a significant decline in recognition accuracy and difficulty in meeting the actual application requirements. Summary of the Invention
[0006] This invention is mainly applied to the telephone communication scenario of non-performing asset disposal. During the process of non-performing asset disposal, staff communicate with customers by phone to understand information such as the customers' repayment willingness and financial status. However, telephone communication is often interfered by background noise, such as the noise in the disposal venue and the line current noise. This invention can analyze the call speech in real time, accurately identify the emotions of customers and staff, help staff adjust communication strategies in a timely manner, and improve the efficiency of non-performing asset disposal.
[0007] To overcome the above defects of the prior art, this invention provides a speech emotion recognition method and device for non-performing asset disposal. This invention constructs a model and an association matrix by collecting historical noise samples, accurately analyzes the misjudgment probability of noise on emotions, then enhances the emotion features of the speech signal, and finally uses multi-modal information and the model for comprehensive emotion recognition. This invention effectively solves the problem that the accuracy of speech emotion recognition is affected under complex background noise, improves the accuracy of speech emotion recognition in the non-performing asset disposal scenario, and helps to promote the non-performing asset disposal work more efficiently.
[0008] To achieve the above object, this invention provides the following technical solutions:
[0009] A speech emotion recognition method for non-performing asset disposal, including:
[0010] Collect non-speech segment noise samples of historical non-performing asset disposal calls, construct a Gaussian mixture model-hidden Markov chain joint model according to the noise samples, and generate a noise-emotion misjudgment association matrix M ne ;
[0011] Obtain the current speech signal, and according to the constructed Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment association matrix M ne , obtain the misjudgment probability of each emotion type caused by the noise in the current speech signal;
[0012] According to the misjudgment probability of each emotion type caused by the noise in the current speech signal, enhance the emotion features of the current speech signal to obtain the current speech signal with enhanced emotions;
[0013] Perform emotion recognition on the current speech signal with enhanced emotions to obtain the emotion type of the current speech signal with enhanced emotions.
[0014] Further, the construction of the Gaussian mixture model - hidden Markov chain joint model based on the noise samples includes:
[0015] Perform frame segmentation on each noise sample to form a noise feature matrix N of the noise samples;
[0016] According to the noise feature matrix N of the noise samples, use the spectral clustering - Mahalanobis distance metric algorithm to perform clustering analysis on the noise samples, and divide the noise samples into n2 types of steady - state noises;
[0017] For each type of steady - state noise, construct a Gaussian mixture model - hidden Markov chain joint model, and a total of n2 Gaussian mixture model - hidden Markov chain joint models are constructed.
[0018] Further, the formation of the noise feature matrix N of the noise samples includes:
[0019] Perform frame segmentation on each noise sample, set the frame length to t1 and the frame shift to t2, to obtain a plurality of segmented signals;
[0020] Perform a fast Fourier transform on each segmented signal to convert the time - domain signal to the frequency - domain, and generate a time - frequency matrix;
[0021] According to the time - frequency matrix of each segmented signal, extract the spectral centroid, the energy ratio of 20 Mel bands, and the zero - crossing rate of each segmented signal;
[0022] Combine the spectral centroid, the energy ratio of 20 Mel bands, and the zero - crossing rate of each segmented signal into a noise feature vector, and form the noise feature matrix N of the noise samples from the noise feature vectors of each segmented signal.
[0023] Further, the division of the noise samples into n2 types of steady - state noises includes:
[0024] Calculate the Mahalanobis distance matrix D of the noise feature matrix N M ;
[0025] Calculate the Mahalanobis distance matrix D M The median σ = median(D M ), and calculate the similarity matrix W according to the median σ and the Mahalanobis distance matrix D M ;
[0026] Calculate the degree matrix D based on the constructed similarity matrix W, and calculate the Laplacian matrix L according to the degree matrix D and the similarity matrix W;
[0027] Perform eigenvalue decomposition on the obtained Laplacian matrix L to obtain n1 eigenvectors, and take the first n2 largest eigenvectors to form a reduced - dimension space; where, n1 > n2;
[0028] Using the constructed dimensionality reduction space, the K-means++ algorithm is used to divide the noise samples into n2 types of steady-state noise, and the steady-state noise includes mechanical sounds in the disposal site, line current sounds, and keyboard tapping sounds.
[0029] Further, the generated noise-emotion misjudgment correlation matrix includes:
[0030] Statistically obtain n3 types of emotion types, and obtain the probability distribution of misjudgment of different emotion types caused by different types of steady-state noise in the noise samples, to obtain the emotion misjudgment probability matrix P; the element p in the emotion misjudgment probability matrix P i'j' represents the probability distribution of misjudgment of the j'-th emotion type caused by the i'-th type of steady-state noise; where, i' is the index of the steady-state noise type, i' = 1, 2,..., n2, and j' is the index of the emotion type, j' = 1, 2,..., n3;
[0031] Normalize the element p in the emotion misjudgment probability matrix P i'j' to obtain the influence coefficient m of the i'-th type of steady-state noise on the j'-th emotion type i'j' , and form the noise-emotion misjudgment correlation matrix M from m i'j' . ne .
[0032] Further, the method for obtaining the probability of misjudgment of each emotion type caused by the noise in the current voice signal includes:
[0033] Perform frame processing on the current voice signal to form the feature matrix V of the current voice signal;
[0034] Input the feature matrix V of the current voice signal into the constructed n2 Gaussian mixture model-hidden Markov chain joint models to obtain the probability q that the current voice signal belongs to each type of steady-state noise i' , where, i' = 1, 2,..., n2;
[0035] According to the probability q that the current voice signal belongs to each type of steady-state noise i' and the noise-emotion misjudgment correlation matrix M ne , calculate the probability u of misjudgment of each emotion type caused by the noise in the current voice signal j' .
[0036] Further, the method for enhancing the emotion features of the current voice signal includes:
[0037] Extract voice segment data from historical non-performing asset disposal calls as historical voice data, and count the emotional characteristics of the voice signals of each type of emotional type in the historical voice data to obtain the probability distribution of the emotional characteristics of each type of emotional type; according to the probability of misjudgment of each type of emotional type caused by the noise in the current voice signal and the historical voice data, determine the emotional types that need to be enhanced in the current voice signal, and the number of emotional types that need to be enhanced is n4;
[0038] Extract the emotional characteristics corresponding to the emotional type j that needs to be enhanced from the probability distribution of the emotional characteristics of each type of emotional type, and perform emotional characteristic enhancement on the current voice signal according to the emotional type j that needs to be enhanced in the current voice signal and the emotional characteristics corresponding to the emotional type j to obtain the current voice signal after emotional enhancement; where j = 1,..., n4.
[0039] Further, the emotional characteristics include the first acoustic characteristics and the first prosodic characteristics;
[0040] The determination of the emotional types that need to be enhanced in the current voice signal includes:
[0041] According to the probability u of misjudgment of each type of emotional type caused by the noise in the current voice signal j' and the probability distribution of the emotional characteristics of each type of emotional type, calculate the posterior probability that the current voice signal belongs to each type of emotional type;
[0042] Determine the n4 emotional types with the largest posterior probability that the current voice signal belongs to each type of emotional type as the emotional types that need to be enhanced in the current voice signal.
[0043] Further, the obtaining of the current voice signal after emotional enhancement includes:
[0044] According to the first acoustic characteristics corresponding to the emotional type j that needs to be enhanced, adjust the spectrum of the current voice signal to enhance the first acoustic characteristics of the emotional type j to obtain the voice signal after acoustic characteristic enhancement;
[0045] According to the first prosodic characteristics corresponding to the emotional type j that needs to be enhanced, adjust the fundamental frequency curve and rhythm of the current voice signal to enhance the first prosodic characteristics of the emotional type j to obtain the voice signal after prosodic characteristic enhancement;
[0046] Fuse the voice signals after acoustic characteristic enhancement and prosodic characteristic emotional enhancement to obtain the current voice signal after emotional enhancement.
[0047] A voice emotion recognition device for non-performing asset disposal, which is used to implement the above-mentioned voice emotion recognition method for non-performing asset disposal. The device includes:
[0048] Noise Modeling Module: Used to collect non-speech segment noise samples of historical non-performing asset disposal calls, construct a Gaussian mixture model-hidden Markov chain joint model based on the noise samples, and generate a noise-emotion misjudgment correlation matrix M ne ;
[0049] Misjudgment Probability Calculation Module: Used to obtain the current voice signal, and based on the constructed Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment correlation matrix M ne , obtain the probability of misjudgment of each emotion type caused by the noise in the current voice signal;
[0050] Emotion Enhancement Module: According to the probability of misjudgment of each emotion type caused by the noise in the current voice signal, perform emotion feature enhancement on the current voice signal to obtain the current voice signal after emotion enhancement;
[0051] Emotion Recognition Module: Used to perform emotion recognition on the current voice signal after emotion enhancement to obtain the emotion type of the current voice signal after emotion enhancement.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0053] By collecting historical noise samples to construct a model and a correlation matrix, the present invention can accurately analyze the misjudgment probability of various emotions caused by noise in the current voice signal. On this basis, emotion feature enhancement is performed on the voice signal, effectively compensating for the emotion features interfered by noise and highlighting the key emotion information. Finally, comprehensive emotion recognition is carried out using multi-modal information and the model, fully considering the acoustic, prosodic and semantic features of speech and the context information. The overall solution improves the accuracy of speech emotion recognition in complex background noise from noise analysis, emotion feature enhancement to multi-modal recognition, meets the requirement of accurate speech emotion recognition in the non-performing asset disposal scenario, helps to more accurately grasp the emotions of both parties in communication, and improves the disposal efficiency and quality. Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0055] Figure 1 It is the principle flow chart of the speech emotion recognition method for non-performing asset disposal in the present invention;
[0056] Figure 2 It is the method flow chart of forming the noise feature matrix N of the noise samples in the speech emotion recognition method for non-performing asset disposal in the present invention;
[0057] Figure 3 This is the flowchart of the method for dividing noise samples into n2 types of steady-state noise in the voice emotion recognition method for non-performing asset disposal of the present invention;
[0058] Figure 4 This is the flowchart of the method for constructing a noise-emotion misjudgment correlation matrix in the voice emotion recognition method for non-performing asset disposal of the present invention;
[0059] Figure 5 This is the flowchart of the method for determining the emotion type to be enhanced in the current voice signal in the voice emotion recognition method for non-performing asset disposal of the present invention;
[0060] Figure 6 This is the flowchart of the method for obtaining the current voice signal with enhanced emotion in the voice emotion recognition method for non-performing asset disposal of the present invention;
[0061] Figure 7 This is the flowchart of the method for obtaining the acoustic similarity coefficient, prosody similarity coefficient, and semantic similarity coefficient between the current voice signal with enhanced emotion and each emotion type in the voice emotion recognition method for non-performing asset disposal of the present invention;
[0062] Figure 8 This is the functional module diagram of the voice emotion recognition device for non-performing asset disposal in the present invention. Detailed implementation manners
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] Embodiment 1
[0065] Please refer to Figure 1 As shown, this embodiment provides a voice emotion recognition method for non-performing asset disposal, including:
[0066] Step S1000, collect non-speech segment noise samples of historical non-performing asset disposal calls, construct a Gaussian mixture model-hidden Markov chain joint model according to the noise samples, and generate a noise-emotion misjudgment correlation matrix M ne ; obtain the current voice signal, according to the constructed Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment correlation matrix M ne, obtain the probability of misjudgment of each emotion type caused by the noise in the current voice signal; according to the probability of misjudgment of each emotion type caused by the noise in the current voice signal, perform emotion feature enhancement on the current voice signal to obtain the current voice signal after emotion enhancement;
[0067] Further, step S1000 includes:
[0068] Step S1100, collect non-speech segment noise samples of historical non-performing asset disposal calls, construct a Gaussian mixture model-hidden Markov chain joint model according to the noise samples, and generate a noise-emotion misjudgment correlation matrix M ne ;
[0069] Further, step S1100 includes:
[0070] Step S1110, construct a Gaussian mixture model-hidden Markov chain joint model according to the noise samples;
[0071] Further, step S1110 includes:
[0072] Step S1111, extract non-speech noise samples from historical non-performing asset disposal calls, perform frame splitting on each noise sample to form a noise feature matrix N of the noise samples;
[0073] Specifically, during the process of non-performing asset disposal, the telephone communication scenario is often accompanied by various complex background noises, which seriously interfere with the accuracy of speech emotion recognition. Extract from more than 1000 historical non-performing asset disposal calls. A large number of telephone samples are selected because the noise situations in the non-performing asset disposal scenario are complex and diverse, and factors such as different call environments and equipment conditions will cause differences in noise characteristics. By collecting a sufficient number of telephone samples, as many noise types and change situations as possible can be covered, making subsequent analysis and modeling more representative and reliable. The sampling rate of these calls is 16 kHz and 16-bit quantization. A sampling rate of 16 kHz means that the sound signal is sampled 16,000 times per second. A higher sampling rate can capture the details of the sound signal more precisely and restore the true characteristics of the sound. 16-bit quantization means that the amplitude value of each sampling point is encoded in 16-bit binary, which can provide richer amplitude information, making the resolution of the collected sound signal in amplitude higher and reducing quantization errors. The duration of each non-speech noise sample is ≥200 ms. Such a duration requirement is set because the characteristics of the noise often need to be presented completely within a certain time range. If the sample duration is too short, the key characteristics of the noise may not be included, resulting in misjudgment of the noise type or inaccurate analysis. Extracting non-speech noise samples aims to specifically obtain those noise signals that interfere with speech emotion recognition, separate them from the speech signal, and thus analyze the characteristics and laws of the noise targeted.
[0074] Further, as Figure 2 shown, step S1111 includes:
[0075] Step S11111, perform frame segmentation on each noise sample, set the frame length to t1 and the frame shift to t2, to obtain a plurality of segmented signals;
[0076] Step S11112, perform fast Fourier transform on each segmented signal, convert the time-domain signal to the frequency domain, and generate a time-frequency matrix;
[0077] Step S11113, according to the time-frequency matrix of each segmented signal, extract the spectral centroid, the energy ratio of 20 Mel bands, and the zero-crossing rate of each segmented signal;
[0078] Step S11114, combine the spectral centroid, the energy ratio of 20 Mel bands, and the zero-crossing rate of each segmented signal into a noise feature vector, and form a noise feature matrix N of the noise sample from the noise feature vectors of each segmented signal.
[0079] Specifically, for the non-speech segment noise samples extracted from more than 1000 historical bad asset disposal calls, preferably, the frame length t1 is set to 20 ms and the frame shift t2 is set to 10 ms. The purpose of frame processing is to discretize the continuous noise signal and divide the long-time continuous noise signal into shorter and more manageable small segments, that is, frame signals. This is because the subsequent analysis of the noise signal requires precision at a smaller time scale, and continuous signals are not conducive to a detailed analysis of their characteristics. Through frame processing, each frame signal can be processed individually, improving the accuracy and precision of the analysis. In the scenario of bad asset disposal calls, steady-state noise persists, and frame processing can more clearly capture the changing characteristics of the noise at different moments. For example, if there is continuous mechanical noise in the disposal site, after framing, the changes in characteristics such as the frequency and amplitude of the mechanical noise can be observed in each short period of time, providing basic data for the subsequent accurate identification and analysis of the noise type. The fast Fourier transform is an efficient algorithm used to transform a signal represented in the time domain to the frequency domain for analysis. The noise signal is manifested as physical quantities such as voltage or sound pressure that change over time in the time domain, making it difficult to intuitively analyze its frequency components. Through the fast Fourier transform, these time-domain signals can be transformed into frequency-domain signals, showing the energy distribution of the signal with frequency as the variable, thus generating a time-frequency matrix. Each element in the time-frequency matrix corresponds to the signal energy intensity at different times and frequencies, clearly showing the energy distribution of each frame signal at different frequencies and providing a data basis for subsequent feature extraction. For example, for a framed noise signal containing line current noise, after the fast Fourier transform, obvious energy peaks can be seen in the time-frequency matrix in the frequency range of 1000 - 3000 Hz, which matches the frequency range of the line current noise, helping to identify the type of this noise.
[0080] The spectral centroid reflects the center of gravity of the signal frequency component and is obtained by calculating the weighted average of each frequency component in the signal spectrum. For each frame signal, the spectral centroid is calculated based on its spectrum. This feature can reflect the concentration trend of the signal frequency. The spectral centroid will show different characteristics for different types of steady-state noise. For example, the spectral centroid of high-frequency noise is relatively high, and the spectral centroid of low-frequency noise is relatively low. In the bad asset disposal scenario, the mechanical sound frequency band of the disposal site is concentrated in 200-800Hz, which is low-frequency noise, and its spectral centroid is relatively low; while the line current sound frequency range is 1000-3000Hz, which is high-frequency noise, and its spectral centroid is relatively high. By analyzing the spectral centroid, different types of noise can be preliminarily distinguished. The frequency range is divided into 20 Mel bands. Mel frequency is a nonlinear frequency scale based on the auditory characteristics of the human ear, which can better reflect the difference in the human ear's perception of sounds of different frequencies. The signal energy in each Mel band is calculated, and its proportion in the total signal energy is calculated. Since different types of steady-state noise have different energy distributions in each Mel frequency band, for example, the mechanical sound in the disposal site has a high proportion of energy in the low-frequency Mel frequency band, while the line current sound has a prominent proportion of energy in specific medium and high-frequency Mel frequency bands. By analyzing these energy distribution characteristics, different types of noise can be effectively distinguished. The zero-crossing rate indicates the number of times a signal crosses the zero level per unit time. For each frame signal, the number of times the signal amplitude crosses the zero level from positive to negative or from negative to positive within a frame length is counted, and then divided by the frame length to obtain the zero-crossing rate. The zero-crossing rate can reflect the intensity of the signal change. Noise with obvious suddenness, such as keyboard tapping, has a relatively high zero-crossing rate, while the zero-crossing rate of continuous and stable mechanical sound is relatively low. By extracting these three features, the characteristics of the noise frame signal can be fully described, providing rich feature information for subsequent noise classification and modeling.
[0081] The spectral centroid, the energy ratio of 20 Mel frequency bands, and the zero-crossing rate calculated for each sub-frame signal are combined into a feature vector. After each noise sample undergoes frame segmentation and feature extraction, multiple such feature vectors will be obtained. These feature vectors are arranged in sequence to form the noise feature matrix N of the noise sample. For example, if a noise sample undergoes frame segmentation to obtain 100 frame signals, and the dimension of the feature vector calculated for each sub-frame signal is 22 (1 spectral centroid + 20 energy ratios of Mel frequency bands + 1 zero-crossing rate), then the noise feature matrix N of this noise sample is a 100×22 matrix. The noise feature matrices N of multiple noise samples together constitute the overall noise feature dataset for subsequent analysis. This method integrates multiple features together to form a structured data matrix, facilitating subsequent operations such as clustering analysis and model construction. By constructing the noise feature matrix, the complex noise signal features can be presented in matrix form, and using matrix operations and data analysis methods, the noise data can be processed and analyzed more efficiently, thereby improving the ability to identify and process noise in the non-performing asset disposal scenario, laying a foundation for accurately extracting speech emotion features subsequently, and effectively solving the problem that the accuracy of speech emotion recognition is affected by complex background noise.
[0082] Step S1112: According to the noise feature matrix N of the noise sample, use the spectral clustering-Mahalanobis distance metric algorithm to perform clustering analysis on the noise sample, and divide the noise sample into n2 types of steady-state noise;
[0083] Furthermore, as Figure 3 shown, step S1112 includes:
[0084] Step S11121: Calculate the Mahalanobis distance matrix D of the noise feature matrix N M ;
[0085] Step S11122: Calculate the median σ = median(D M ) in D, and calculate the similarity matrix W according to the median σ and the Mahalanobis distance matrix D M ; M Step S11123: Calculate the degree matrix D based on the constructed similarity matrix W, and calculate the Laplacian matrix L according to the degree matrix D and the similarity matrix W;
[0086] Step S11124: Perform eigen-decomposition on the obtained Laplacian matrix L to obtain n1 eigenvectors, and take the first n2 largest eigenvectors to form a dimensionality reduction space; where, n1 > n2;
[0087]
[0088] Step S11125, using the constructed dimensionality reduction space, the noise samples are divided into n2 types of steady-state noise by the K-means++ algorithm, and the steady-state noise includes mechanical sounds at the disposal site, line current sounds, and keyboard tapping sounds.
[0089] Specifically, step S1112 aims to perform clustering analysis on the noise samples according to the noise feature matrix N of the noise samples by using the spectral clustering-Mahalanobis distance metric algorithm, and divide the noise samples into n2 types of steady-state noise, so as to achieve effective classification of complex noises in the non-performing asset disposal scenario, and lay a foundation for accurately identifying and processing the interference of noises on speech emotion recognition subsequently.
[0090] In step S11121, the Mahalanobis distance matrix DM of the noise feature matrix N is calculated. The noise feature matrix N is composed of feature vectors of multiple noise samples, and each feature vector contains key features reflecting noise characteristics such as spectral centroid, energy ratio of 20 Mel bands, and zero-crossing rate. Mahalanobis distance is a distance metric method that considers the data distribution characteristics, and its calculation formula is based on the noise feature matrix N. The elements in the Mahalanobis distance matrix DM represent the feature vector of the th noise sample in the noise feature matrix N and the th noise sample The Mahalanobis distance between the feature vectors. Different from the traditional Euclidean distance, the Mahalanobis distance considers the covariance structure of the data and can more accurately measure the degree of difference between the characteristics of different noise samples. For example, in the non-performing asset disposal scenario, the feature vectors of the mechanical sound at the disposal site and the line current sound may not fully reflect the essential difference between them under the Euclidean distance, but through the Mahalanobis distance calculation, the distribution differences of these two types of noises in features such as spectral centroid and Mel band energy ratio can be comprehensively considered, and a more reasonable distance metric can be given. By calculating the Mahalanobis distance matrix DM, a quantitative basis is provided for subsequent analysis of the similarity and difference between noise samples, which helps to gather noise samples with similar characteristics together, improve the accuracy of noise classification, and further improve the accuracy of speech emotion recognition in complex background noises.
[0091] In the noise environment of non-performing asset disposal, the Mahalanobis distance matrix DM contains numerous distance data, and the median σ represents the typical level of these Mahalanobis distances. Use the formula to construct the similarity matrix W. In this formula, the exponential function converts the Mahalanobis distance into similarity. When the Mahalanobis distance between two noise samples is It will approach 1, indicating that the characteristics of these two noise samples are similar; conversely, if the Mahalanobis distance is large, the similarity will approach 0, meaning that the characteristics of the two noise samples are quite different. For example, suppose there are noise samples A and B. Sample A is the mechanical sound generated by an office printer, and sample B is the sound generated by another similar printer. They are quite similar in characteristics such as spectral centroid and Mel band energy ratio. The calculated Mahalanobis distance is small, and the similarity calculated through the formula will be high, and they are more likely to be classified into the same category in subsequent clustering analysis. Construct a similarity matrix W to convert the Mahalanobis distance into a similarity form that is easier to understand and analyze, providing an important basis for determining the degree of closeness of the association between noise samples, helping to cluster noise samples more accurately, thereby improving the ability of the speech emotion recognition system to distinguish different noise types and enhancing the accuracy of speech emotion recognition in a complex noise environment.
[0092] Calculate the degree matrix D based on the constructed similarity matrix W, and then calculate the Laplacian matrix L according to the degree matrix D and the similarity matrix W. The degree matrix D is a diagonal matrix, and the elements on its diagonal , represent the sum of the weights of all edges connected to the th noise sample (i.e., the sum of similarities with other noise samples), reflecting the relative importance of this sample in the entire set of noise samples. For example, if a certain noise sample has high similarities with multiple other noise samples, it indicates that it is closely associated with other samples in the noise sample set, and the corresponding diagonal element value in the degree matrix D will be large. Then, use the formula to calculate the Laplacian matrix L. The Laplacian matrix L plays an important role in the noise clustering scenario of non-performing asset disposal. It can capture the local structural information between noise samples and convert the similarity matrix W into a form more suitable for eigen-decomposition and clustering analysis. By calculating the degree matrix D and the Laplacian matrix L, the internal connections between noise samples are further explored, providing strong support for more efficient noise clustering in the future, helping to improve the processing ability of the speech emotion recognition system for complex noises, and thus enhancing the reliability of speech emotion recognition in a complex background noise.
[0093] When dealing with bad asset disposal telephone noise sample data with a high dimension, directly processing these data will face the problems of huge computational complexity and possible redundant information. Eigenvalue decomposition is a mathematical method that decomposes the Laplacian matrix L into multiple eigenvectors and corresponding eigenvalues through eigenvalue decomposition. The magnitude of the eigenvalue reflects the contribution degree of the corresponding eigenvector to the information of the original matrix. Selecting the first n2 (e.g., 8) largest eigenvectors to form the dimensionality reduction space is because the eigenvalues corresponding to these eigenvectors are larger, representing the main change directions and structural information in the data. For example, among many noise sample data, the eigenvectors corresponding to these larger eigenvalues may reflect the key feature differences between different types of steady-state noises (such as mechanical sounds at the disposal site, line current sounds, etc.). By selecting these key eigenvectors to construct the dimensionality reduction space, the data dimension can be reduced while retaining the main feature information. Subsequent clustering will be performed in this dimensionality reduction space, which not only improves the computational efficiency but also avoids the computational complexity and redundant information interference problems that may be brought about by processing high-dimensional data, helps to cluster the noise samples more accurately, and thus improves the efficiency and accuracy of speech emotion recognition in complex background noises.
[0094] The K-means++ algorithm is a commonly used clustering algorithm. Its core idea is to divide data into different categories according to the distances between data points. In this embodiment, the K-means++ algorithm can select relatively reasonable initial clustering centers based on the data distribution in the dimensionality-reduced space. For example, in the dimensionality-reduced space, the algorithm will analyze the feature distribution of noise samples and select those points that are more reasonably distributed in the data space as the initial clustering centers, avoiding the problem of unstable clustering results that may be caused by randomly selecting initial clustering centers. Through continuous iteration, the algorithm gradually divides the noise samples into different categories. In the scenario of non-performing asset disposal, these n2 categories respectively correspond to different types of steady-state noises such as mechanical sounds in the disposal site and line current sounds. The intensity range of the mechanical sounds in the disposal site is roughly 40 - 60 dB, and its frequency band is concentrated in 200 - 800 Hz. For example, the continuous buzzing sounds generated when devices such as printers and copiers in the office are running belong to this category. Through long-term monitoring and analysis of these mechanical sounds, we found that their frequency characteristics are relatively stable and show a certain regularity. The frequency range of the line current sounds is between 1000 - 3000 Hz, and it shows the characteristics of periodic pulses. In actual calls, factors such as line aging and signal interference may cause the appearance of current sounds. We have made detailed records of the current sounds under different line conditions and found that their pulse periods and amplitudes are related to the line quality to a certain extent. In this way, clustering the complex noise samples into n2 categories according to their feature similarity helps to more accurately identify and analyze different types of steady-state noises in the process of non-performing asset disposal. After accurately identifying different types of steady-state noises, subsequent targeted processing can be carried out on the impact of different types of noises on speech emotion recognition, such as constructing a more accurate noise-emotion misjudgment correlation matrix, so as to improve the accuracy of speech emotion recognition in complex background noises and solve the problem that the recognition accuracy of current speech emotion recognition technology drops significantly in high-noise environments.
[0095] Step S1113, for each type of steady-state noise, construct a Gaussian mixture model-hidden Markov chain joint model, and a total of n2 Gaussian mixture model-hidden Markov chain joint models are constructed.
[0096] Step S1113 aims to construct a Gaussian mixture model-hidden Markov chain joint model (GMM-HMC) for the n2 types of steady-state noises divided in step S1112. A total of n2 such joint models are constructed to accurately identify and compensate for the noises in the non-performing asset disposal scenario, thereby improving the accuracy of speech emotion recognition in complex background noises. During the construction process, the Gaussian mixture model (GMM) is used to model the time-frequency characteristics of each type of noise to generate the static distribution of the noise. GMM is a probability model that assumes the data is composed of a mixture of multiple Gaussian distributions. For each type of steady-state noise, the energy distribution at different frequencies and times shows a specific pattern. GMM fits these time-frequency characteristics through a weighted combination of multiple Gaussian distributions. For example, for the mechanical sound in the disposal site, its frequency is concentrated in the range of 200 - 800 Hz, and the intensity range is approximately 40 - 60 dB. GMM adjusts the parameters of each Gaussian distribution, such as the mean, covariance, and weight, according to these characteristics, so that the combined model can accurately describe the time-frequency characteristics of the mechanical sound, thereby generating the static distribution of this type of noise. In this way, GMM can capture the frequency and energy distribution characteristics of the noise at a certain moment, providing a basis for subsequent analysis. At the same time, the hidden Markov chain (HMC) is used to model the time-evolution characteristics of the noise to generate the dynamic change trajectory of the noise. HMC is a stochastic process that includes a set of states and the transition probabilities between states. In noise modeling, each state can represent the characteristic state of the noise at different moments, and the state transition matrix describes the probability of transitioning from one state to another. For example, for the line current sound, which has the characteristics of periodic pulses, HMC simulates the changes in the current sound over time by setting appropriate initial state probabilities and state transition matrices. As time goes by, the pulse period and amplitude of the current sound may change, and HMC can update the state according to these changes, thereby generating a trajectory reflecting its dynamic changes. This helps to capture the change trend of the noise at different moments and is crucial for analyzing the time characteristics of the noise.
[0097] Finally, a joint Gaussian Mixture Model - Hidden Markov Chain (GMM - HMC) model for each type of noise is generated. In this joint model, the GMM part is responsible for describing the static characteristics of the noise, and the HMC part is responsible for describing the dynamic changes of the noise. The two are combined to comprehensively characterize the characteristics of the noise. Specifically, each HMC state in the GMM part corresponds to a GMM, enabling the model to describe the time - frequency characteristics of the noise according to the corresponding GMM in different states. When training the model parameters, the Baum - Welch algorithm is used. This algorithm is an iterative algorithm that maximizes the likelihood of the model for the training data by continuously adjusting the model parameters (such as the means, covariances, weights of the GMM, and the state transition matrix and initial state probabilities of the HMC). For example, when training on a large number of line current noise samples, the Baum - Welch algorithm will gradually optimize the model parameters according to the time - frequency characteristics and time series of the samples, making the GMM - HMC joint model more accurately fit the characteristics of the line current noise.
[0098] The beneficial effects of this step are significant. By constructing the GMM - HMC joint model, different types of steady - state noises in the non - performing asset disposal scenario can be modeled and analyzed more accurately. An accurate noise model helps to subsequently accurately identify the noise type in the current voice signal, and then calculate the probability of misjudgment of each emotion type caused by the noise. Based on these accurate probabilities, more effective enhancement of the emotion features of the current voice signal is carried out to improve the accuracy of voice emotion recognition. In a complex background noise environment, accurate noise recognition and compensation can avoid the interference of noise on emotion features, more accurately extract the emotion features in the voice, thereby improving the performance of the entire voice emotion recognition system and meeting the requirements for accurate voice emotion recognition in the non - performing asset disposal scenario.
[0099] Step S1120: Based on the collected noise samples, construct a noise - emotion misjudgment correlation matrix;
[0100] Furthermore, as Figure 4 shown, step S1120 includes:
[0101] Step S1121: Statistically obtain n3 types of emotion types, and obtain the probability distribution of misjudgment of different emotion types caused by different types of steady - state noises in the noise samples to obtain an emotion misjudgment probability matrix P; the element p i'j' in the emotion misjudgment probability matrix P represents the probability distribution of misjudgment of the j'th emotion type caused by the i'th type of steady - state noise; where i' is the index of the steady - state noise type, i' = 1, 2,..., n2, and j' is the index of the emotion type, j' = 1, 2,..., n3;
[0102] Step S1122: The element p i'j'Normalize to obtain the influence coefficient m of the i'-th type of steady-state noise on the j'-th type of emotion. i'j' , which consists of m i'j' to form the noise-emotion misjudgment correlation matrix M ne .
[0103] Specifically, in the telephone communication scenario of non-performing asset disposal, common emotion types include anger, happiness, anxiety, etc., which constitute n3 types of emotion types. For different types of steady-state noise, such as the mechanical sound in the disposal site and the line current sound, their misjudgment effects on different emotion types are different. Taking the current sound as an example, since its frequency is in the range of 1000 - 3000 Hz, which belongs to high-frequency noise, it is easily misjudged as a sharp tone, thus causing the emotion recognition system to misjudge the normal emotion as anger. For every 1 dB increase in the current sound, the threshold of the anger emotion will be lowered by 0.15. For the mechanical sound, its frequency band is concentrated in 200 - 800 Hz, which belongs to low-frequency noise. When the duration exceeds 5 seconds, it may mask the anxiety features such as the tremor in the speaker's voice, reducing the confidence level of the anxiety emotion by 0.2.
[0104] Collect a large number of voice samples containing different types of steady-state noise and corresponding clear emotion annotations. In these samples, determine the original emotion type through manual annotation or by using an existing high-precision emotion recognition system for preliminary emotion recognition. Then, after adding various steady-state noises with different intensities and durations, use the same emotion recognition system for recognition again, compare the emotion recognition results before and after adding the noise, and count the situations of misjudgment of various emotion types caused by different types of steady-state noise. For example, select 1000 voice samples containing different emotions such as anger, happiness, and anxiety, add different intensities of current sound respectively, record the number of samples whose emotion recognition results change after adding the current sound, as well as the specific misjudgment direction, so as to calculate the probability of misjudgment of each emotion caused by the current sound. Through statistical analysis of a large number of samples, obtain the emotion misjudgment probability matrix P, where the element p i'j' represents the probability distribution of misjudgment of the j'-th type of emotion (j' = 1, 2,..., n3) caused by the i'-th type of steady-state noise (i' = 1, 2,..., n2).
[0105] Normalization is a data processing method whose purpose is to map data to a specific interval, making different data comparable. In step S1122, normalizing the emotional misjudgment probability matrix P can eliminate the dimensional differences in probability values among different noises and emotional types, and more intuitively reflect the relative interference degrees of different noises on various emotions. For example, assume that the probability of mechanical sound in the disposal site causing misjudgment of anger emotion is 0.3, and the probability of causing misjudgment of anxiety emotion is 0.2. Through normalization, the influence coefficients on anger emotion and anxiety emotion can be obtained, and these coefficients can directly compare the differences in the interference degrees of mechanical sound on different emotions. Convert the misjudgment probabilities of each steady-state noise on different emotional types into relative influence coefficients, and all the influence coefficients form the noise-emotion misjudgment correlation matrix M ne For example, m 12 represents the influence degree of mechanical sound in the disposal site on anxiety emotion. Through this matrix, the interference degrees of different noises on various emotions can be clearly quantified. In the subsequent sound field compensation process, the voice signals affected by noise can be processed specifically according to this matrix. For example, when identifying voice emotions, according to the noise-emotion misjudgment correlation matrix, adjust and correct the emotional features that may be misjudged due to noise interference, improve the accuracy of voice emotion recognition, and solve the problem of decreased emotion recognition accuracy under complex background noise.
[0106] Step S1200: Obtain the current voice signal, and according to the constructed Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment correlation matrix, obtain the probabilities of each type of emotional misjudgment caused by noise in the current voice signal;
[0107] Furthermore, step S1200 includes:
[0108] Step S1210: Perform frame segmentation on the current voice signal to form the feature matrix V of the current voice signal;
[0109] Furthermore, step S1210 includes:
[0110] Step S1211: Perform frame segmentation on the current voice signal, set the frame length to t1, and the frame shift to t2, to obtain multiple voice frame signals;
[0111] Step S1212: Perform short-time Fourier transform on each voice frame signal to generate the time-frequency matrix of the voice frame;
[0112] Step S1213: According to the time-frequency matrix of each voice frame, obtain the feature matrix V of the current voice signal.
[0113] Specifically, in the scenario of telephone communication for non-performing asset disposal, similar to when dealing with noise samples, in order to facilitate subsequent fine analysis of the speech signal, the continuous speech signal needs to be discretized. Preferably, the frame length t1 is set to 20 ms, and the frame shift t2 is set to 10 ms. Through such frame segmentation processing, the longer speech signal is divided into shorter speech frame signals. For example, a speech signal with a duration of 1 second can be divided into approximately 100 speech frame signals according to the above parameters. The purpose of this operation is to convert the continuous speech signal into discrete and easily processed units because subsequent feature extraction and analysis of the speech signal often need to be carried out on shorter time segments, and frame segmentation processing can more finely capture the feature changes of the speech signal at different moments. Through frame segmentation processing, more accurate analysis can be carried out for each speech frame signal subsequently, laying a foundation for accurately identifying speech emotions and helping to improve the accuracy of speech emotion recognition in complex background noise. The short-time Fourier transform is a method for converting a time-domain signal into a frequency-domain signal, and it is used in this step to analyze the energy distribution of the speech frame signal at different frequencies. For each speech frame signal, it appears as physical quantities such as voltage or sound pressure that change over time in the time domain, and it is difficult to obtain its frequency component information directly from the time-domain observation. Through the short-time Fourier transform, the change of each speech frame signal in the time domain can be converted to the frequency domain, and the energy distribution of the signal is displayed with frequency as the variable, thus generating a time-frequency matrix. For example, for a speech frame signal containing human voice and background noise, after the short-time Fourier transform, the time-frequency matrix can clearly show in which frequency ranges the human voice energy is higher and which frequencies are more affected by background noise. This provides a data basis for subsequent extraction of the features of the speech frame signal, enabling the analysis of the speech signal from the frequency perspective, helping to distinguish the effective information and noise components in the speech, and improving the accuracy of speech emotion recognition.
[0114] According to the time-frequency matrix of each speech frame, extract the spectral centroid, the energy proportion of 20 Mel bands, and the zero-crossing rate of each speech frame, and combine these features into a speech feature vector. The feature matrix V of the current speech signal is formed by the speech feature vectors of each speech frame. The spectral centroid reflects the centroid position of the signal frequency components, which is obtained by calculating the weighted average of the frequency components in the signal spectrum, and it can reflect the concentration trend of the speech signal frequency. For speech signals with different emotions, their spectral centroids may vary. For example, the spectral centroid of a speech signal with an angry emotion may be relatively high. Divide the frequency range into 20 Mel bands. The Mel frequency is a non-linear frequency scale based on the auditory characteristics of the human ear, which can better reflect the perceptual differences of the human ear for sounds of different frequencies. Calculate the energy of the signal in each Mel band and find its proportion in the total energy of the entire signal. The energy distributions of speeches with different emotions in each Mel band are different. For example, the energy proportion of a speech with a happy emotion in some Mel bands may be significantly different from that of an anxious emotion, which helps to distinguish speeches with different emotions from the perspective of human ear perception. The zero-crossing rate represents the number of times the signal crosses the zero level per unit time. For each speech frame signal, count the number of times the signal amplitude crosses the zero level from positive to negative or from negative to positive within a frame length, and then divide by the frame length to obtain the zero-crossing rate. The zero-crossing rate can reflect the severity of signal changes. For example, the zero-crossing rate of a speech with strong emotional fluctuations may be relatively high. Combine the spectral centroid, the energy proportion of 20 Mel bands, and the zero-crossing rate of each speech frame into a speech feature vector. After each speech signal is framed and feature-extracted, multiple such feature vectors will be obtained. Arrange these feature vectors in order to form the feature matrix V of the speech signal. For example, if a speech signal is framed into 80 speech frame signals, and the dimension of the feature vector calculated for each speech frame signal is 22 (1 spectral centroid + 20 Mel band energy proportions + 1 zero-crossing rate), then the feature matrix V of the speech signal is an 80×22 matrix. The feature matrices V of multiple speech signals constitute the speech feature dataset for subsequent analysis. By extracting these features and constructing the feature matrix V, the features of the speech signal are comprehensively described, providing rich information for subsequent determining the noise type in the speech signal and calculating the misjudgment probability of the noise on the emotion type, which helps to improve the accuracy of speech emotion recognition in complex background noise.
[0115] Step S1220, input the feature matrix V of the current speech signal into the constructed n2 Gaussian mixture model-hidden Markov chain joint models to obtain the probability q that the current speech signal belongs to each type of steady-state noise i' , where i' = 1, 2,..., n2;
[0116] Specifically, in the telephone communication scenario of non-performing asset disposal, n2 GMM-HMC joint models have been constructed for different types of steady-state noises. These models respectively model the characteristics of steady-state noises such as mechanical sounds and line current sounds at the disposal site. When the feature matrix V of the current voice signal is input into these models, each GMM-HMC joint model will match and analyze the feature matrix V according to its own learning and modeling results of different noise characteristics. Taking the GMM-HMC joint model corresponding to the mechanical sound at the disposal site as an example, this model matches the time-frequency features in the feature matrix V through its Gaussian mixture model (GMM) part to judge the similarity between the time-frequency features of the current voice signal and the mechanical sound at the disposal site. At the same time, the hidden Markov chain (HMC) part will consider the changes in the feature matrix V in the time series, and combine the state transition matrix and the initial state probability to analyze whether the voice signal conforms to the dynamic change law of the mechanical sound at the disposal site in the time dimension. Through the comprehensive analysis of these two parts, the model will output a probability value, that is, the probability q that the current voice signal belongs to this type of steady-state noise of the mechanical sound at the disposal site i' (assuming that i' corresponds to the category index of the mechanical sound at the disposal site). Similarly, the other n2-1 GMM-HMC joint models will respectively calculate the probability that the current voice signal belongs to their respective corresponding steady-state noise categories
[0117] The beneficial effect of this step is that by inputting the feature matrix of the voice signal into multiple pre-constructed GMM-HMC joint models, the possibility of various steady-state noises contained in the current voice signal can be accurately judged. Accurately identifying the noise type in the voice signal is a key step in calculating the misjudgment probability of noise on the emotion type subsequently. Only by clarifying the noise type and its probability in the voice signal can the misjudgment impact of noise on various emotion types be calculated more precisely according to the noise-emotion misjudgment correlation matrix, thereby improving the accuracy of voice emotion recognition under complex background noises and meeting the requirements for accurate voice emotion recognition in the non-performing asset disposal scenario
[0118] Step S1230, according to the probability q that the current voice signal belongs to each type of steady-state noise i' and the noise-emotion misjudgment correlation matrix M ne , calculate the probability u that the noise in the current voice signal causes misjudgment of each type of emotion type j' .
[0119] Specifically, the probability q that the current voice signal belongs to each type of steady-state noise i' is obtained through step S1220, while the noise-emotion misjudgment correlation matrix M ne is constructed in step S1120, and the element m in the matrix i'j'It represents the influence coefficient of the i'-th type of steady-state noise on the j'-th type of emotion. Calculate the probability u of misjudgment of each type of emotion caused by the noise in the current speech signal. j' When j' , for each type of emotion j' (j' = 1, 2,..., n3), all categories of steady-state noise i' (i' = 1, 2,..., n2) will be traversed. According to the calculation principle of probability, use the formula to calculate. For example, assume that the probability q1 of the current speech signal belonging to the mechanical sound in the disposal site is 0.3, and the probability q2 of belonging to the line current sound is 0.2. The influence coefficient m 11 of the mechanical sound in the disposal site on the angry emotion is 0.1, and the influence coefficient m 21 of the line current sound on the angry emotion is 0.2. Then for the angry emotion, the probability u1 of misjudgment caused by the noise is 0.3×0.1 + 0.2×0.2 = 0.07. In this way, by comprehensively considering the probabilities of various types of steady-state noise in the speech signal and the influence degree of each type of steady-state noise on different types of emotions, the probability of misjudgment of each type of emotion caused by the noise in the current speech signal is calculated.
[0120] The beneficial effect of this step is reflected in that by accurately calculating the probability uj' of misjudgment of each type of emotion caused by the noise, it provides a key basis for subsequent enhancement of the emotion features of the current speech signal. Under complex background noise, different types of noise have different influences on the misjudgment of different types of emotions. Accurately calculating uj' can clarify the degree of interference of each type of emotion by the noise. Based on these probability values, the emotion features of the speech signal can be enhanced targeted, compensating for the emotion features interfered by the noise, thereby improving the accuracy of speech emotion recognition, solving the problem of the decline in emotion recognition accuracy under complex background noise, and meeting the requirements for accurate speech emotion recognition in the non-performing asset disposal scenario.
[0121] Step S1300, according to the probability u j' of misjudgment of each type of emotion caused by the noise in the current speech signal, perform emotion feature enhancement on the current speech signal to obtain the current speech signal with enhanced emotions.
[0122] Furthermore, step S1300 includes:
[0123] Step S1310, extract speech segment data from historical non-performing asset disposal calls as historical speech data, and determine the emotion types that need to be enhanced in the current speech signal according to the probability of misjudgment of each type of emotion caused by the noise in the current speech signal and the historical speech data;
[0124] Furthermore, as Figure 5 shown, step S1310 includes:
[0125] Step S1311: Statistically analyze the emotional characteristics of the voice signals of each emotional type in the historical voice data to obtain the probability distribution of the emotional characteristics of each emotional type; the emotional characteristics include the first acoustic characteristics and the first prosodic characteristics.
[0126] Step S1312: According to the probability u of misjudgment of each emotional type caused by the noise in the current voice signal j' and the probability distribution of the emotional characteristics of each emotional type, calculate the posterior probability that the current voice signal belongs to each emotional type.
[0127] Step S1313: Determine the n4 emotional types with the largest posterior probabilities that the current voice signal belongs to each emotional type as the emotional types to be enhanced in the current voice signal, where n4 ≥ 1.
[0128] Specifically, the first acoustic characteristics cover elements such as fundamental frequency, energy, formants, etc. The voices of different emotions vary in these aspects. For example, the voice of an angry emotion usually has a higher fundamental frequency and greater energy; while the voice of a happy emotion may have a unique pattern in the formant distribution. The first prosodic characteristics include speech rate, pause duration, intonation, etc. For example, an anxious emotion may cause an increase in speech rate and a shortening of the pause duration. By analyzing a large amount of historical voice data, the occurrence frequencies and variation ranges of these characteristics under each emotional type are statistically analyzed, so as to obtain the probability distribution of the emotional characteristics of each emotional type. For example, for 1000 pieces of historical voice data labeled as angry emotion, the average value, standard deviation of the fundamental frequency, and the distribution of the speech rate are statistically analyzed, etc., so as to determine the probability distribution of the emotional characteristics of the angry emotion. This step provides the basic data for calculating the possibility that the current voice signal belongs to various emotions in the follow-up, helps to mine the characteristic laws of different emotions from historical data, and provides strong support for accurately judging the emotional type of the current voice signal.
[0129] The posterior probability is to re-evaluate the probability that the current voice signal belongs to a certain emotional type based on the existing evidence (that is, the noise situation in the current voice signal and the probability distribution of the emotional characteristics in the historical voice data). It is calculated using Bayes' formula. Assuming there are n3 emotional types in total, for the j'-th emotional type (j' = 1, 2,..., n3), its calculation formula is: posterior probability = (probability u of misjudgment of the j'-th emotional type caused by the noise j' × probability distribution of the emotional characteristics of the j'-th emotional type) / ∑(probability u of misjudgment of each emotional type caused by the noise j' × probability distribution of the emotional characteristics of each emotional type) (j' ranges from 1 to n3). For example, it is known that the probability u of misjudgment of the angry emotion caused by the noise in the current voice signal j'It is 0.3. The probability distribution of the emotional characteristics of the angry emotion obtained from historical voice data is 0.2. At the same time, the sum of (the probability of emotional misjudgment caused by noise × the probability distribution of emotional characteristics) for all emotion types is calculated to be 0.5. Then, the posterior probability that the current voice signal belongs to the angry emotion = (0.3 × 0.2) / 0.5 = 0.12. By calculating the posterior probability, the noise interference and historical emotional characteristic information are comprehensively considered, which more accurately reflects the possibility that the current voice signal belongs to each emotion type, providing a quantitative basis for determining the emotion types that need to be enhanced subsequently.
[0130] Determine the n4 emotion types with the largest posterior probabilities that the current voice signal belongs to each emotion type as the emotion types that need to be enhanced in the current voice signal, where n4 ≥ 1. This is because the emotion type with the largest posterior probability is more likely to be the emotion actually contained in the current voice signal. However, due to factors such as noise interference, its emotional characteristics may be weakened, so it needs to be enhanced. For example, if the calculated posterior probabilities that the current voice signal belongs to the angry, happy, and anxious emotion types are 0.4, 0.3, and 0.2 respectively, and n4 is set to 2, then the angry and happy emotion types are determined as the emotion types that need to be enhanced. In this way, it is possible to focus on the most likely emotion types, enhance the emotional characteristics in a targeted manner, avoid unnecessary processing of irrelevant emotion types, improve the processing efficiency, and at the same time help to highlight the key emotional characteristics in the voice signal that are interfered by noise, laying a foundation for accurately identifying emotions subsequently, thereby improving the accuracy of voice emotion recognition in complex background noise.
[0131] Step S1320, extract the emotional characteristics corresponding to the emotion type j that needs to be enhanced from the probability distribution of the emotional characteristics of each emotion type. According to the emotion type j that needs to be enhanced in the current voice signal and the emotional characteristics corresponding to the emotion type j, perform emotional characteristic enhancement on the current voice signal to obtain the current voice signal after emotion enhancement; where j = 1,..., n4.
[0132] Furthermore, as Figure 6 shown, step S1320 includes:
[0133] Step S1321, adjust the spectrum of the current voice signal according to the first acoustic characteristics corresponding to the emotion type j that needs to be enhanced, enhance the first acoustic characteristics of the emotion type j, and obtain the voice signal after acoustic characteristic enhancement;
[0134] Step S1322, adjust the fundamental frequency curve and rhythm of the current voice signal according to the first prosodic characteristics corresponding to the emotion type j that needs to be enhanced, enhance the first prosodic characteristics of the emotion type j, and obtain the voice signal after prosodic characteristic enhancement;
[0135] Step S1323: Fuse the speech signal after acoustic feature enhancement and prosodic feature enhancement to obtain the current speech signal with enhanced emotion.
[0136] Specifically, step S1320 aims to enhance the emotion features of the current speech signal according to the emotion type j to be enhanced in the current speech signal and the emotion features corresponding to the emotion type j, so as to obtain the current speech signal with enhanced emotion, improve the accuracy of speech emotion recognition, and solve the problem that emotion features are interfered under complex background noise. In the telephone communication scenario of non-performing asset disposal, there are significant differences in the acoustic features of speeches with different emotion types. For example, the speech with an angry emotion has higher energy in certain frequency bands, and the formant distribution has a specific pattern. This step strengthens the acoustic features related to the target emotion by targeted adjustment of the speech signal spectrum, reduces the masking of emotion features by noise, and then improves the accuracy of speech emotion recognition under complex background noise. In actual operation, it is first necessary to clarify the frequency components involved in the first acoustic feature corresponding to the emotion type j to be enhanced. For example, if the current emotion type to be enhanced is anger, according to the first acoustic feature of the angry emotion, it is found that its energy is relatively concentrated in the mid-high frequency band. Signal processing techniques such as filtering can be used to increase the energy in the mid-high frequency band, adjust the spectrum shape, and make the features of the speech signal more prominent in these key frequency bands, thereby enhancing the first acoustic feature of the angry emotion. The purpose of this is to strengthen the acoustic features related to the target emotion in the speech signal, make it more conform to the acoustic pattern of this emotion type, reduce the masking of emotion features by noise, improve the sensitivity of the emotion recognition system to the target emotion, and then improve the accuracy of speech emotion recognition under complex background noise.
[0137] Taking the angry emotion as an example, its energy is relatively concentrated in the mid-high frequency band (assuming the frequency range is f1 - f2), where f1 represents the starting frequency of the mid-high frequency band concerned for enhancing the target emotion features such as the angry emotion, and f2 represents the ending frequency of this frequency band. These two values are determined based on the research and analysis of the first acoustic features of emotions such as anger, and the unit is Hertz (Hz). To enhance this feature, filtering technology can be used to adjust the spectrum of the speech signal. Here, a band-pass filter is used, and the transfer function H(f) of the band-pass filter can be expressed as: .
[0138] Suppose the representation of the current speech signal in the frequency domain is , and the frequency domain signal obtained after filtering is:
[0139] ;
[0140] Through this band-pass filter, only the frequency components in the medium and high frequency band (f1 - f2) are retained, effectively suppressing the signals in other frequency bands and highlighting the frequency characteristics related to angry emotions. However, simply retaining the frequency components is not enough; the energy of these key frequency bands also needs to be enhanced. Assume that the energy in the medium and high frequency band needs to be increased by k' times , where k' is the multiple of energy enhancement, then the enhanced frequency-domain signal is:
[0141] ;
[0142] Finally, the enhanced frequency-domain signal is converted back to the time domain through the inverse Fourier transform to obtain the speech signal with enhanced acoustic features :
[0143] ;
[0144] where represents the inverse Fourier transform operation. Through such processing, the energy of the speech signal in the medium and high frequency band is enhanced, the spectral shape is changed, making it more in line with the acoustic pattern of angry emotions, strengthening the acoustic features related to angry emotions, improving the sensitivity of the emotion recognition system to angry emotions, and helping to more accurately identify angry emotions under complex background noise.
[0145] In step S1322, according to the first prosodic features corresponding to the emotion type j to be enhanced, the fundamental frequency curve and rhythm of the current speech signal are adjusted to enhance the first prosodic features of the emotion type j, obtaining a speech signal with enhanced prosodic features. The fundamental frequency curve in the first prosodic features reflects the pitch change of the speech, and the rhythm involves aspects such as speech rate and pause duration. The prosodic features of different emotions are significantly different. For example, the anxious emotion may be manifested as an accelerated speech rate and larger fluctuations in the fundamental frequency; while the calm emotion has a relatively slow speech rate and a more stable fundamental frequency curve. For the emotion type j to be enhanced, analyze its first prosodic features, and then enhance these features by adjusting the fundamental frequency curve and rhythm of the speech signal. For example, if the prosodic features of the anxious emotion are to be enhanced, detect the fundamental frequency curve and rhythm of the current speech signal, appropriately increase the speech rate, and increase the fluctuation amplitude of the fundamental frequency to make the speech more in line with the prosodic characteristics of the anxious emotion. In this way, the prosodic features related to the target emotion in the speech signal are further strengthened, making the emotion expression more obvious, helping the emotion recognition system to more accurately judge the emotion type, and improving the reliability of speech emotion recognition under complex background noise.
[0146] Taking the adjustment of the fundamental frequency curve to enhance the prosodic features of the anxious emotion as an example. First, detect the fundamental frequency curve F0(t) of the current speech signal. Assume that the average fundamental frequency of the original speech signal is , and the standard deviation of the fundamental frequency fluctuation is To enhance the characteristics of anxious emotions, it is necessary to appropriately increase the speech rate and the amplitude of the fundamental frequency fluctuations.
[0147] The increase in speech rate can be achieved by compressing the time scale. Let the time compression ratio be , and the adjusted time axis is ×t. Under the new time axis, the fundamental frequency curve becomes . To increase the amplitude of the fundamental frequency fluctuations, assume that it is necessary to increase the standard deviation of the fundamental frequency fluctuations by m' times , then the adjusted fundamental frequency curve is:
[0148]
[0149] where is a random noise function with a mean of 0 and a standard deviation of 1, used to simulate the fluctuations of the fundamental frequency. Through such processing, the fluctuations of the fundamental frequency curve become more intense, conforming to the characteristics of anxious emotions. This formula is based on the average fundamental frequency , combined with the adjusted standard deviation of the fundamental frequency fluctuations and the random noise function , to achieve the enhancement of the amplitude of the fundamental frequency fluctuations, making the fundamental frequency curve more in line with the characteristics of target emotions such as anxious emotions.
[0150] In terms of rhythm, taking anxious emotions as an example, it is necessary to increase the speech rate and adjust the pause duration. Assume that the total duration of the original speech signal is , the total pause duration is T pause , and the average speech rate is ' ( is the number of syllables in the speech signal). To increase the speech rate, increase the speech rate by times , and the new average speech rate . At the same time, shorten the total pause duration. Let the shortening ratio be , and the new total pause duration .
[0151] According to the new speech rate and total pause duration, rearrange the syllable and pause distribution in the speech signal. Assume that the starting time of the i''-th syllable in the original speech signal is t start,i'' , the ending time is t end,i'' , and the pause time is t pause,i'' . The adjusted starting time t' start,i'' and ending time t' end,i'' can be calculated according to the new speech rate and pause arrangement. For example:
[0152] Calculate the adjusted starting time t' start,i'' :
[0153] Starting from the first syllable, .
[0154] For the case of, .
[0155] In this formula, as the index variable for cumulative calculation, starts taking values from 1, increments by 1 each time, until . During the cumulative process, sequentially obtain the syllable duration ( ) corresponding to each index and the pause time ( ), add them together, and finally multiply by the time compression ratio to determine the adjusted start time of the th syllable.
[0156] Calculate the adjusted end time t' end,i'' : ( is the adjusted start time of this syllable, represents the duration after scaling the original duration of the th syllable according to the time compression ratio . Adding the two together gives the adjusted end time.
[0157] The adjusted pause time : , since the total pause duration shortening ratio is q , so each pause time is shortened according to this ratio.
[0158] Through the above adjustments to the fundamental frequency curve and rhythm, the speech signal is made to be more in line with the characteristics of anxious emotions in prosodic features, further strengthening the prosodic features related to anxious emotions, making the emotion expression more prominent, helping the emotion recognition system to more accurately judge anxious emotions, and improving the reliability of speech emotion recognition under complex background noise.
[0159] After enhancing the acoustic features and prosodic features respectively in the previous two steps, fusing the two is to comprehensively improve the expression of the target emotion in the speech signal. The fusion process can adopt methods such as weighted summation, assign different weights according to the importance of the acoustic features and prosodic features in expressing emotions, and then merge the enhanced acoustic features and prosodic features. Through fusion, the speech signal is made to be more prominent in both acoustic and prosodic aspects in terms of the target emotion features, effectively compensating for the interference of noise on emotion features, providing a better-quality signal for subsequent accurate recognition of speech emotions, thereby improving the accuracy of speech emotion recognition under complex background noise and meeting the requirements for accurate recognition of speech emotions in the non-performing asset disposal scenario.
[0160] Step S2000: Perform emotion recognition on the current speech signal after emotion enhancement to obtain the emotion type of the current speech signal after emotion enhancement.
[0161] Furthermore, step S2000 includes:
[0162] Step S2100: Perform emotion annotation and semantic analysis on historical speech data, extract semantic features related to emotions in the historical speech data after emotion annotation, and construct semantic templates for each emotion type according to the extracted semantic features;
[0163] Specifically, the main purpose of step S2100 is to construct semantic templates for each emotion type by processing historical speech data, so as to be used for subsequent matching degree calculation with the current speech signal, thereby improving the accuracy of speech emotion recognition in complex background noise. In the telephone communication scenario of non-performing asset disposal, the semantic content of speech is closely related to emotion expression, and different emotion types are often reflected by specific words, phrases and sentence structures. In actual operation, first classify and annotate the semantic content of each historical speech data according to common emotion types, such as anger, happiness, anxiety, calmness, etc. Taking the angry emotion as an example, statements with obvious angry tendencies such as "Don't rush anymore" and "You are harassing us" will be marked; for the happy emotion, expressions such as "I will handle it as soon as possible" and "The problem has been solved" will be marked. After the annotation is completed, extract semantic features related to emotions from these annotated historical speech data. Specifically, count the frequently occurring words and phrases under different emotion types, and analyze the grammar structure and semantic logic of the sentences. For example, through a large amount of data statistics, it is found that the angry emotion is often accompanied by negative words (such as "no", "don't") and imperative sentences (such as sentences led by "must", "right away", etc.); while the happy emotion mostly contains positive words (such as "good", "satisfied") and affirmative expressions (such as "yes", "that's right").
[0164] Based on these extracted semantic features, a semantic template is constructed. The semantic template includes content such as a set of key semantic vocabulary and typical sentence structure patterns. For the angry emotion, its semantic template is set to include specific negative words, high-frequency negative evaluation vocabulary, and the sentence structure is mostly short and powerful imperative or accusatory. Suppose after statistics, words such as "don't" and "garbage" frequently appear in the angry emotion, and sentence structures such as "don't + verb" and "you + negative evaluation vocabulary" are relatively typical, then these elements constitute part of the semantic template of the angry emotion. For the happy emotion, the semantic template includes positive words, sentence structures expressing commitment or satisfaction, such as words like "will" and "satisfied", and structures such as "I will + verb" and "very satisfied". Each semantic template corresponds to an emotion type. Subsequently, when identifying the emotion of the current speech signal, the semantic features of the current speech signal can be calculated for matching with these templates, providing a semantic-level basis for emotion recognition, making up for the deficiency of relying only on acoustic and prosodic features for recognition, and improving the accuracy and reliability of speech emotion recognition in complex background noise.
[0165] Step S2200, according to each emotion type, construct a semantic template and the emotion features of the speech signals of each emotion type in the historical speech data, and calculate the acoustic similarity coefficient, prosodic similarity coefficient, and semantic similarity coefficient between the current speech signal with enhanced emotion and each emotion type.
[0166] Furthermore, as Figure 7 shown, step S2200 includes:
[0167] Step S2210, extract the acoustic features of the current speech signal with enhanced emotion, denoted as the second acoustic features, and calculate the acoustic similarity coefficient α between the first acoustic features corresponding to emotion type j' and the second acoustic features j' ;
[0168] Step S2220, extract the prosodic features of the current speech signal with enhanced emotion, denoted as the second prosodic features, and calculate the prosodic similarity coefficient β between the first prosodic features corresponding to emotion type j' and the second prosodic features j' ;
[0169] Step S2230, extract the semantic features of the current speech signal with enhanced emotion, and according to the semantic templates of different emotion types, calculate the semantic similarity coefficient γ between the semantic features and the semantic template of emotion type j' j' .
[0170] Specifically, the acoustic features cover important elements such as fundamental frequency, energy, formants, etc. The speech of different emotion types has significant differences in acoustic features. For example, the speech of the angry emotion usually has a higher fundamental frequency, greater energy, and a unique pattern in the distribution of formants. Calculate the acoustic similarity coefficient αj' When taking the fundamental frequency as an example, first obtain the average fundamental frequency in the first acoustic feature corresponding to the emotion type j' respectively and the average fundamental frequency in the second acoustic feature of the current speech signal after emotion enhancement . Calculate the similarity degree by calculating the relative error between the two. The formula is . In this formula, represents the absolute value of the difference between the two average fundamental frequencies, represents the larger value of the two average fundamental frequencies. Through such calculation, the closer the obtained value is to 1, the more similar the fundamental frequencies of the two are. For other acoustic features such as energy and formants, specific algorithms are also used to calculate their respective similarity scores. Assume that the energy similarity score is , and the formant similarity score is . Combining these scores, calculate the acoustic similarity coefficient by weighted average. By calculating the acoustic similarity coefficient , the similarity degree between the current speech signal after emotion enhancement and the emotion type j' can be quantified from the acoustic perspective, providing a key acoustic basis for emotion recognition and helping to more accurately judge the speech emotion under complex background noise.
[0171] Prosodic features mainly include aspects such as speech rate, pause duration, intonation, etc. The prosodic features of different emotions are significantly different. For example, an anxious emotion may be manifested as an increased speech rate, a shortened pause duration, and a larger intonation fluctuation; while a calm emotion has a relatively slow speech rate, a longer pause duration, and a more stable intonation. When calculating the prosodic similarity coefficient , taking the speech rate as an example, first determine the average speech rate in the first prosodic feature corresponding to the emotion type j' and the average speech rate in the second prosodic feature of the current speech signal after emotion enhancement, and calculate the speech rate similarity through the formula . Similarly, for the pause duration and intonation, appropriate algorithms are used to calculate the similarity scores respectively. Assume that the average pause duration similarity score is , and obtain the intonation similarity score by comparing the change trend and amplitude of the intonation. Combining these scores to obtain the prosodic similarity coefficient . By calculating the prosodic similarity coefficient , the similarity degree between the current speech signal after emotion enhancement and the emotion type j' is quantified from the prosodic perspective, providing an important prosodic reference for emotion recognition and further improving the accuracy of speech emotion recognition under complex background noise.
[0172] Convert the speech signal into text, and then extract semantic information such as keywords and sentiment tendencies. Taking the semantic template of anger emotion as an example, assume that the template contains key semantic words such as "don't" and "garbage", as well as typical sentence structures such as "don't + verb" and "you + negative evaluation word". For the text converted from the current speech signal after emotion enhancement, count the frequency of the key semantic words that appear. Assume that the frequency of "don't" appearing is , and the frequency of "garbage" appearing is , then the similarity score of the key semantic words can be calculated by the formula , where is the total number of words in the text. For the sentence structure, judge the proportion of the number of sentences that conform to the typical sentence structure in the total number of sentences. Assume that the number of sentences that conform to the "don't + verb" structure is , and the total number of sentences is , then the similarity score of the sentence structure. Combine these scores to obtain the semantic similarity coefficient . By calculating the semantic similarity coefficient , the matching degree between the current speech signal after emotion enhancement and the emotion type is evaluated at the semantic level. Combining the acoustic and prosodic similarity coefficients can more comprehensively and accurately determine the emotion type of the current speech signal, and improve the accuracy and reliability of speech emotion recognition in complex background noise.
[0173] Step S2300, construct a feature similarity matrix S according to the acoustic similarity coefficient, prosodic similarity coefficient and semantic similarity coefficient, and obtain the confidence score of the current speech signal after emotion enhancement belonging to each emotion type according to the feature similarity matrix S and the pre-constructed multi-modal emotion recognition model;
[0174] Furthermore, step S2300 includes:
[0175] Step S2310, construct a feature similarity matrix S according to the acoustic similarity coefficient, prosodic similarity coefficient and semantic similarity coefficient. The matrix element S kj' represents the similarity coefficient of the current speech signal after emotion enhancement to the j'-th emotion type in the k-th feature dimension; where k = 1, 2, 3, representing acoustic features, prosodic features and semantic features respectively;
[0176] Step S2320, obtain the confidence score of the current speech signal after emotion enhancement belonging to each emotion type according to the feature similarity matrix S and the pre-constructed multi-modal emotion recognition model.
[0177] Specifically, the core purpose of step S2300 is to obtain the confidence scores of the current speech signal with enhanced emotions belonging to each emotion type by constructing a feature similarity matrix S and combining a pre-constructed multi-modal emotion recognition model, providing a quantitative basis for finally determining the emotion type of the speech signal, and thus improving the accuracy of speech emotion recognition under complex background noise. The acoustic similarity coefficient α j' , the prosodic similarity coefficient β j' , and the semantic similarity coefficient γ j' respectively reflect the similarity degrees between the current speech signal with enhanced emotions and each emotion type from different dimensions. The matrix element S kj' represents the similarity coefficient of the current speech signal with enhanced emotions to the j'-th emotion type in the k-th feature dimension, where k = 1, 2, 3, representing acoustic features, prosodic features, and semantic features respectively; j' represents different emotion types, such as anger, happiness, anxiety, etc. For example, assuming there are three emotion types currently (anger, happiness, anxiety), then j' takes values of 1, 2, 3. For a certain enhanced speech signal, its similarity coefficient with the angry emotion in the acoustic feature dimension is α1, in the prosodic feature dimension is β1, and in the semantic feature dimension is γ1. Then in the feature similarity matrix S, S 11 = α1, S 21 = β1, S 31 = γ1, and so on, to complete the construction of the entire matrix. By constructing such a matrix, the multi-dimensional similarity information is integrated, facilitating subsequent comprehensive analysis of the correlation degree between the speech signal and various emotion types, providing comprehensive data support for determining the emotion type of the speech signal, and helping to improve the accuracy and reliability of emotion recognition.
[0178] The multi-modal emotion recognition model is trained based on a large amount of speech data, and it can learn the complex relationships between different feature combinations and various emotion types. In this embodiment, the model can be a support vector machine (SVM) model, etc. Taking the SVM model as an example, its training process uses a large amount of speech data with labeled emotion types as the training data set, and these data cover various emotion types and speech samples in different noise environments. During training, the acoustic, prosodic, and semantic features of the speech data are used as inputs, and the corresponding emotion types are used as outputs. By adjusting the model parameters, the model can accurately classify the input features. After obtaining the feature similarity matrix S, it is used as an input and passed to the multi-modal emotion recognition model. The model will analyze and process the data in the matrix according to the previously learned relationship between features and emotion types. For example, the model will comprehensively consider the similarity coefficients in the dimensions of acoustic, prosodic, and semantic features, and through a series of calculations and judgments, output the confidence scores of the current speech signal with enhanced emotion belonging to each emotion type. These confidence scores reflect the likelihood of the speech signal belonging to different emotion types. For example, the model outputs that the confidence score of the speech signal belonging to the angry emotion is 0.7, the confidence score of the happy emotion is 0.2, and the confidence score of the anxious emotion is 0.1, which indicates that the model believes that the speech signal is most likely the angry emotion. In this way, using the multi-modal emotion recognition model to process the feature similarity matrix, the obtained confidence scores provide an important quantitative reference for accurately judging the emotion type of the speech signal in the future, and further improve the accuracy and reliability of speech emotion recognition in complex background noise.
[0179] Step S2400: According to the confidence scores of the current speech signal with enhanced emotion belonging to each emotion type, and combining with the context semantics, comprehensively judge the emotion type of the current speech signal with enhanced emotion to obtain the emotion type of the current speech signal with enhanced emotion.
[0180] Further, step S2400 includes:
[0181] Step S2410: Obtain the previous and next N2 speech segments of the current speech signal, denoted as context speech segments; according to the constructed multi-modal emotion recognition model, obtain the emotion type sequence E of the context speech segments;
[0182] Step S2420: Statistically analyze the transition probabilities of different emotion types in the historical speech data, and construct an emotion transition probability matrix P ec , and the matrix element represents the probability of transitioning from emotion type to emotion type ;
[0183] Step S2430, based on the confidence scores of the current voice signal with enhanced emotions belonging to each emotion type, the emotion type sequence E of the context voice segments, and the emotion transition probability matrix P ec , calculate the most likely emotion type of the current voice signal with enhanced emotions through the Viterbi algorithm.
[0184] Specifically, in the actual telephone communication scenario for non-performing asset disposal, the emotional expression of speech often has coherence and context relevance. Obtaining the N2 voice segments before and after the current voice signal is to utilize the context information to assist in judging the emotion type of the current voice signal. For example, when N2 = 3, the first 3 and the last 3 voice segments of the current voice signal are obtained as context voice segments. For these context voice segments, feature extraction is performed in a similar manner to that of the current voice signal. First, each context voice segment is framed to obtain multiple voice frame signals, then the short-time Fourier transform is performed on each voice frame signal to generate a time-frequency matrix, and then features such as spectral centroid, energy ratio of 20 Mel bands, and zero-crossing rate are extracted, and combined to form the feature matrix of each context voice segment. Then, these feature matrices are input into the constructed multi-modal emotion recognition model, and the model will output the emotion type of each context voice segment according to the relationship between the features learned during previous training and the emotion types, and these emotion types are arranged in order to form the emotion type sequence E. By considering the emotion types of the context voice segments, it is possible to more comprehensively understand the emotional context in which the current voice signal is located, provide more information for accurately judging the emotion type of the current voice signal, and help improve the accuracy of speech emotion recognition in complex background noise.
[0185] In a large amount of historical voice data, there are certain transfer rules between different emotion types. For example, in the telephone communication for non-performing asset disposal, the customer may start communicating with a relatively calm emotion, but as the problem is discussed, the emotion may gradually turn into anxiety or anger. By counting the number of times of transfer from one emotion type to another emotion type , and then dividing by the total number of occurrences of the emotion type , the corresponding transfer probability can be obtained. Suppose that in 1000 segments of historical voice data, the situation of emotion transfer from calm to anxiety appears 50 times, and the total number of occurrences of the calm emotion is 200 times, then the probability of emotion transfer from calm to anxiety (where represents the calm emotion, represents the anxiety emotion) is 50÷200 = 0.25. By statistically analyzing all possible emotion type transfer situations, the emotion transition probability matrix P is constructed. ecThis matrix reflects the transfer trends between different emotion types, providing an important reference basis for comprehensively judging the emotion type of the current speech signal by combining the confidence score of the current speech signal and the emotion types of the context speech segments. It helps to more accurately identify speech emotions and improve the reliability of speech emotion recognition in complex background noises.
[0186] The Viterbi algorithm is a dynamic programming algorithm for finding the optimal path. In the speech emotion recognition scenario of this embodiment, the confidence scores of the current speech signal enhanced by emotions belonging to each emotion type are used as the initial probabilities of each state, the emotion type sequence E of the context speech segments is used as the known state sequence, and the emotion transfer probability matrix P ec is used as the state transition probability. For example, assume that the confidence scores of the current speech signal belonging to the three emotion types of anger, happiness, and anxiety are 0.6, 0.2, and 0.2 respectively, and the previous segment in the emotion type sequence E of the context speech segments is the anger emotion. According to the emotion transfer probability matrix P ec , the probabilities of transferring from the anger emotion to different emotion types are known. Based on this information, the Viterbi algorithm will calculate an optimal path, that is, the most likely emotion type sequence, through dynamic programming, and the last state of this sequence is the most likely emotion type of the current speech signal enhanced by emotions. In this way, by comprehensively considering the characteristics of the speech signal itself, the context information, and the emotion transfer probability, the emotion type of the current speech signal can be determined more accurately, effectively improving the accuracy of speech emotion recognition in complex background noises and meeting the requirements for accurate recognition of speech emotions in the non-performing asset disposal scenario.
[0187] Embodiment 2
[0188] Based on Embodiment 1, this embodiment provides a speech emotion recognition device for non-performing asset disposal, as Figure 8 shown, including:
[0189] Noise modeling module: used to collect non-speech segment noise samples of historical non-performing asset disposal calls, construct a Gaussian mixture model-hidden Markov chain joint model according to the noise samples, and generate a noise-emotion misjudgment correlation matrix M ne ;
[0190] Misjudgment probability calculation module: used to obtain the current speech signal, and obtain the probability of misjudging each emotion type caused by the noise in the current speech signal according to the constructed Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment correlation matrix M ne ,
[0191] Emotion Enhancement Module: Enhance the emotion features of the current speech signal according to the probability of misjudgment of each emotion type caused by noise in the current speech signal, and obtain the current speech signal with enhanced emotions;
[0192] Emotion Recognition Module: Used to recognize the emotion of the current speech signal with enhanced emotions, and obtain the emotion type of the current speech signal with enhanced emotions.
[0193] In the noise modeling module, the generation of the noise-emotion misjudgment correlation matrix M ne includes:
[0194] Step S11121, calculate the Mahalanobis distance matrix D of the noise feature matrix N M ;
[0195] Step S11122, calculate the median σ = median(D M ) in the Mahalanobis distance matrix D, and calculate the similarity matrix W according to the median σ and the Mahalanobis distance matrix D M ; M Calculate the degree matrix D based on the constructed similarity matrix W, and calculate the Laplacian matrix L according to the degree matrix D and the similarity matrix W;
[0196] Step S11123, perform eigen-decomposition on the obtained Laplacian matrix L to obtain n1 eigenvectors, and take the first n2 largest eigenvectors to form a dimensionality reduction space; where, n1 > n2;
[0197] Step S11124, divide the noise samples into n2 types of steady-state noise by using the constructed dimensionality reduction space and the K-means++ algorithm. The steady-state noise includes mechanical sounds in the disposal site, line current sounds, and keyboard tapping sounds.
[0198] In the misjudgment probability calculation module, the probability of misjudgment of each emotion type caused by the noise in the current speech signal includes:
[0199] Step S1210, perform frame processing on the current speech signal to form the feature matrix V of the current speech signal;
[0200] Step S1220, input the feature matrix V of the current speech signal into the constructed n2 Gaussian mixture model-hidden Markov chain joint models to obtain the probability q that the current speech signal belongs to each type of steady-state noise
[0201] , where, i' = 1, 2,..., n2; i' ;
[0202] Step S1230, according to the probability q that the current speech signal belongs to each type of steady-state noise i' and the noise-emotion misjudgment correlation matrix Mne , calculate the probability u of misjudgment of each emotion type caused by noise in the current speech signal j' .
[0203] In the emotion enhancement module, the obtained current speech signal after emotion enhancement includes:
[0204] Step S1321: According to the first acoustic feature corresponding to the emotion type j to be enhanced, adjust the spectrum of the current speech signal, enhance the first acoustic feature of the emotion type j, and obtain the speech signal with enhanced acoustic features;
[0205] Step S1322: According to the first prosodic feature corresponding to the emotion type j to be enhanced, adjust the fundamental frequency curve and rhythm of the current speech signal, enhance the first prosodic feature of the emotion type j, and obtain the speech signal with enhanced prosodic features;
[0206] Step S1323: Fuse the speech signals after acoustic feature enhancement and prosodic feature emotion enhancement to obtain the current speech signal after emotion enhancement.
[0207] In the emotion recognition module, the emotion recognition of the current speech signal after emotion enhancement includes:
[0208] Step S2100: Perform emotion annotation and semantic analysis on the historical speech data, extract the semantic features related to emotions in the historical speech data after emotion annotation, and construct semantic templates for each emotion type according to the extracted semantic features;
[0209] Step S2200: Calculate the acoustic similarity coefficient, prosodic similarity coefficient, and semantic similarity coefficient between the current speech signal after emotion enhancement and each emotion type according to the semantic templates constructed for each emotion type and the emotion features of the speech signals of each emotion type in the historical speech data;
[0210] Step S2300: Construct a feature similarity matrix S according to the acoustic similarity coefficient, prosodic similarity coefficient, and semantic similarity coefficient, and obtain the confidence scores of the current speech signal after emotion enhancement belonging to each emotion type according to the feature similarity matrix S and the pre-constructed multi-modal emotion recognition model;
[0211] Step S2400: Make a comprehensive judgment on the emotion type of the current speech signal after emotion enhancement according to the confidence scores of the current speech signal after emotion enhancement belonging to each emotion type, combined with the context semantics, to obtain the emotion type of the current speech signal after emotion enhancement.
[0212] The methods and apparatuses of the present application can be implemented in many ways. For example, the methods and apparatuses of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the methods is only for illustration purposes, and the steps of the methods of the present application are not limited to the specific order described above, unless otherwise specifically stated.
[0213] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0214] As described above in the specific embodiments, the purpose, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A speech emotion recognition method for non-performing asset disposal, characterized in that: The method comprises: Collect non-speech noise samples of historical non-performing asset disposal calls, build a Gaussian mixture model-hidden Markov chain joint model based on the noise samples, and generate the noise-emotion misjudgment correlation matrix M ne ; The noise-emotion misjudgment correlation matrix M is generated ne The process includes: obtaining n3 emotion types by statistics, obtaining the probability distribution of different emotion types misjudged by different types of steady-state noise in the noise sample, and obtaining the emotion misjudgment probability matrix P; the element p in the emotion misjudgment probability matrix P is i'j' represents the probability distribution of the i'th type of steady-state noise leading to the j'th emotion type misjudgment; where i' is the index of the steady-state noise type, i'=1, 2, ..., n2, j' is the index of the emotion type type, j'=1, 2, ..., n3; the element p in the emotion misjudgment probability matrix P is i'j' Normalize and get the influence coefficient m of the i'th type of steady-state noise on the j'th emotion type i'j' , by m i'j' Construct the noise-emotion misjudgment correlation matrix M ne ; Where n2 is the number of steady-state noise types; Get the current speech signal, and construct the Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment correlation matrix M ne , get the probability of misjudgment of each type of emotion due to the noise in the current speech signal; According to the probability of misjudgment of each type of emotion caused by noise in the current speech signal, the emotion feature of the current speech signal is enhanced to obtain the current speech signal after emotion enhancement; Emotion recognition is performed on the current speech signal after emotion enhancement to obtain the emotion type of the current speech signal after emotion enhancement.
2. The speech emotion recognition method for non-performing asset disposal according to claim 1 is characterized in that: The Gaussian mixture model-hidden Markov chain joint model constructed according to the noise sample includes: Perform frame processing on each noise sample to form a noise feature matrix N of the noise sample; According to the noise feature matrix N of the noise samples, the spectral clustering-Mahalanobis distance metric algorithm is used to perform cluster analysis on the noise samples and divide the noise samples into n2 types of steady-state noise; For each type of steady-state noise, a Gaussian mixture model-hidden Markov chain joint model is constructed, and a total of n2 Gaussian mixture model-hidden Markov chain joint models are constructed.
3. The speech emotion recognition method for non-performing asset disposal according to claim 2 is characterized in that: The noise feature matrix N forming the noise sample includes: Each noise sample is framed, the frame length is set to t1, the frame shift is set to t2, and multiple framed signals are obtained; Perform fast Fourier transform on each frame signal, convert the time domain signal into the frequency domain, and generate a time-frequency matrix; According to the time-frequency matrix of each frame signal, the spectrum centroid, energy proportion and zero-crossing rate of 20 Mel frequency bands of each frame signal are extracted; The spectrum centroid of each frame signal, the energy proportion of 20 Mel frequency bands and the zero-crossing rate are combined into a noise feature vector, and the noise feature matrix N of the noise sample is formed by the noise feature vector of each frame signal.
4. The speech emotion recognition method for non-performing asset disposal according to claim 3 is characterized in that: The classification of the noise samples into n2 types of steady-state noise includes: Calculate the Mahalanobis distance matrix D of the noise feature matrix N M ; Calculate the Mahalanobis distance matrix D M Medianσ=median(D M ), according to the median σ and the Mahalanobis distance matrix D M Calculate the similarity matrix W; The degree matrix D is calculated based on the constructed similarity matrix W, and the Laplace matrix L is calculated based on the degree matrix D and the similarity matrix W; Perform eigendecomposition on the obtained Laplace matrix L to obtain n1 eigenvectors, and take the first n2 largest eigenvectors to form a reduced dimension space; where n1>n2; By using the constructed dimensionality reduction space, the K-means++ algorithm is used to divide the noise samples into n2 types of steady-state noise, which include mechanical noise in the disposal site, line current noise and keyboard tapping noise.
5. The speech emotion recognition method for non-performing asset disposal according to claim 2 is characterized in that: The probability of misjudging each type of emotion due to noise in the current speech signal includes: Perform frame processing on the current speech signal to form a feature matrix V of the current speech signal; The feature matrix V of the current speech signal is input into the constructed n2 Gaussian mixture model-hidden Markov chain joint model to obtain the probability q that the current speech signal belongs to each type of steady-state noise i' , where i'=1, 2, ..., n2; According to the probability q that the current speech signal belongs to each type of steady-state noise i' and the noise-emotion misjudgment correlation matrix M ne , calculate the probability u that the noise in the current speech signal causes misjudgment of each type of emotion j' .
6. The speech emotion recognition method for non-performing asset disposal according to claim 5 is characterized in that: The emotional feature enhancement of the current speech signal comprises: Extract voice segment data from historical non-performing asset disposal calls as historical voice data, count the emotional features of voice signals of each emotional type in the historical voice data, and obtain the emotional feature probability distribution of each emotional type; determine the emotional type that needs to be enhanced in the current voice signal based on the probability of misjudgment of each emotional type caused by noise in the current voice signal and the historical voice data, and the number of emotional types that need to be enhanced is n4; The emotional features corresponding to the emotion type j that needs to be enhanced are extracted from the emotional features probability distribution of each emotion type. According to the emotion type j that needs to be enhanced in the current speech signal and the emotional features corresponding to the emotion type j, the emotional features of the current speech signal are enhanced to obtain the current speech signal after emotion enhancement; wherein j=1,...,n4.
7. The speech emotion recognition method for non-performing asset disposal according to claim 6 is characterized in that: The emotion feature includes a first acoustic feature and a first prosodic feature; Determining the emotion type that needs to be enhanced in the current speech signal includes: The probability u of misjudging each type of emotion based on the noise in the current speech signal j' and the probability distribution of emotional features of each emotional type, and calculate the posterior probability that the current speech signal belongs to each emotional type; The n4 emotion types with the largest posterior probability that the current speech signal belongs to each emotion type are determined as the emotion types that need to be enhanced in the current speech signal.
8. The speech emotion recognition method for non-performing asset disposal according to claim 7 is characterized in that: The current speech signal after the emotion enhancement is obtained includes: According to the first acoustic feature corresponding to the emotion type j to be enhanced, the frequency spectrum of the current speech signal is adjusted to enhance the first acoustic feature of the emotion type j, thereby obtaining a speech signal after the acoustic feature is enhanced; According to the first prosodic feature corresponding to the emotion type j to be enhanced, the fundamental frequency curve and the rhythm of the current speech signal are adjusted to enhance the first prosodic feature of the emotion type j, thereby obtaining a speech signal with enhanced prosodic features; The speech signals after acoustic feature enhancement and prosodic feature emotion enhancement are fused to obtain the current speech signal after emotion enhancement.
9. A speech emotion recognition device for non-performing asset disposal, which is used to implement the speech emotion recognition method for non-performing asset disposal according to any one of claims 1 to 8, characterized in that: The device comprises: Noise modeling module: used to collect non-speech noise samples of historical non-performing asset disposal calls, build a Gaussian mixture model-hidden Markov chain joint model based on the noise samples, and generate the noise-emotional misjudgment correlation matrix M ne ; Misjudgment probability calculation module: used to obtain the current speech signal, based on the constructed Gaussian mixture model-hidden Markov chain joint model and the noise-emotion misjudgment correlation matrix M ne , get the probability of misjudgment of each type of emotion due to the noise in the current speech signal; Emotion enhancement module: According to the probability of misjudgment of each type of emotion caused by noise in the current speech signal, the emotion feature of the current speech signal is enhanced to obtain the current speech signal after emotion enhancement; Emotion recognition module: used to perform emotion recognition on the current speech signal after emotion enhancement, and obtain the emotion type of the current speech signal after emotion enhancement.
Citation Information
Patent Citations
Speech emotion recognition system and recognition method
CN109243492A
Voice emotion recognition method and device, equipment and storage medium
CN117612569A
Robust speech emotion recognition method based on compressive sensing
CN103021406A
Emotional identification device, method and program
JP2010054568A