AI speech emotion analysis method and platform of deep learning architecture

Through the short-time Fourier transform noise reduction and vocalprint feature extraction of deep learning architecture, combined with Euclidean distance clustering and multi-layer perceptron classification, the problem of noise and multi-speaker separation in speech sentiment analysis is solved, and high-precision sentiment analysis is achieved.

CN120279949AInactive Publication Date: 2025-07-08SHANDONG QIANGBI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510637183.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing speech emotion analysis technology has low accuracy in complex noise environments, making it difficult to separate the voice signals of multiple speakers, resulting in inaccurate emotion analysis.

Method used

The deep learning architecture is adopted to extract the vocalprint features through short-term Fourier transform noise reduction and linear prediction of cepspectral coefficients, and the speaker is separated by Euclidean distance clustering, and multi-layer perceptrons are used for emotional classification.

Benefits of technology

Effectively remove noise and accurately separate speakers, improve the accuracy and adaptability of speech emotion analysis, be able to identify new speakers, enhance the system's adaptability and recognition ability, and improve the accuracy and reliability of emotional classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279949A_ABST
    Figure CN120279949A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech sentiment analysis, and discloses an AI speech sentiment analysis method and platform of a deep learning architecture, and the platform comprises a signal collection unit, a noise reduction processing unit, a speech separation unit, a feature extraction unit, and a sentiment classification unit. In the signal processing link, analog-to-digital conversion is carried out according to the Nyquist sampling theorem, and short-time Fourier transform is utilized to respectively process noisy voice and noise, so that background interference is effectively eliminated; in the speaker separation stage, voiceprint features are extracted by means of an LPCC method, and the Euclidean distance clustering technology is combined, so that the voice of a known speaker can be accurately distinguished, and a new sounding individual can be recognized; the emotion feature extraction layer is used for mining emotion information in the speech from multiple dimensions by calculating the speech speed and analyzing intonation changes; in a sentiment classification link, the multi-layer perceptron model carries out deep processing and probability calculation on sentiment features through collaborative operation of an input layer, a hiding layer and an output layer, and it is ensured that a sentiment prediction result is scientific and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech emotion analysis, and particularly to an AI speech emotion analysis method and platform based on a deep learning architecture. Background Art

[0002] With the rapid development of artificial intelligence technology, speech emotion analysis has broad application prospects in fields such as intelligent customer service, mental health assessment, and human-computer interaction. By analyzing the emotional information in speech, machines can better understand user intentions and perceive user emotions, thus achieving a more intelligent and user-friendly interaction experience.

[0003] However, there are still many problems to be solved in the current speech emotion analysis technology. In the speech signal acquisition link, there are often complex noise interferences in actual application scenarios, such as the noisy voices in public places and the roaring sounds of equipment operation. These noises will seriously affect the quality of speech signals, resulting in a significant reduction in the accuracy of subsequent emotion analysis. Traditional noise reduction methods are difficult to effectively remove these complex noises and cannot meet the actual application requirements.

[0004] In terms of speaker separation, when multiple speakers speak simultaneously, it is difficult for the existing technology to accurately distinguish the speech signals of different speakers. Without accurately separating the speakers, the extracted emotional features will be confused with each other, and accurate emotion analysis cannot be performed for specific speakers, which greatly limits the application of the speech emotion analysis system in multi-speaker scenarios.

[0005] Therefore, there is an urgent need for a method that can effectively overcome the above problems and achieve high-precision speech emotion analysis to promote the further development of speech emotion analysis technology in practical applications. Summary of the Invention

[0006] The purpose of the present invention is to provide an AI speech emotion analysis method and platform based on a deep learning architecture, which solves the technical problems proposed in the background art.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] The AI speech emotion analysis method based on a deep learning architecture includes the following steps:

[0009] Speech signal acquisition: Use a microphone device to collect an analog speech signal S(t), and convert it into a digital speech signal S d [n] through an analog-to-digital converter, where t represents time and n is a discrete time series index;

[0010] Noise reduction processing: Perform noise reduction processing on the speech signal S containing noise d[n] Perform short-time Fourier transform to obtain the speech spectrum X(m,k). During the period when no one is speaking, extract the digital noise signal and perform short-time Fourier transform to obtain the noise spectrum D(m,k). Obtain the denoised speech spectrum Y(m,k) through a noise reduction method based on spectrum subtraction, and perform inverse short-time Fourier transform on Y(m,k) to obtain the denoised speech signal S dn [n];

[0011] Speaker separation based on voiceprint features: Use the linear prediction cepstral coefficient method to extract voiceprint features from the denoised speech signal S dn [n] to obtain the voiceprint feature vector C. Use a clustering method based on Euclidean distance to perform speaker clustering on the voiceprint feature vector C, and classify the new speech signal into the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs;

[0012] Emotion feature extraction: For each speaker, extract a speech segment of a specified duration TS, and calculate the speech rate feature VS and the intonation change feature FB;

[0013] Emotion classification: Use a multi-layer perceptron as the emotion classification model, which includes an input layer, a hidden layer, and an output layer. The input layer receives the emotion feature vector, the hidden layer performs feature transformation and abstraction processing on it, and the output layer calculates the emotion category probability. Determine the emotion prediction result by comparing the output values of the neurons in the output layer.

[0014] As a further solution of the present invention: Among them, the conversion process of the analog-to-digital converter is based on the Nyquist sampling theorem, and its sampling frequency fs satisfies fs≥2f max , and f max is the highest frequency component in the speech signal.

[0015] As a further solution of the present invention: The specific method of noise reduction processing is as follows:

[0016] First, perform short-time Fourier transform on the speech signal S d [n] containing noise;

[0017] The short-time Fourier transform formula is:

[0018] In the formula, X(m,k) is the speech spectrum containing noise, m represents the frame index, k represents the frequency index, N is the frame length, H is the frame shift, w[n] is the window function, and j represents the imaginary unit, which is used to describe the phase information of the signal in frequency domain analysis;

[0019] At the same time, during the period when no one is speaking, extract the digital noise signal containing only noise and no human voice, and perform short-time Fourier transform on it at the same time to obtain the noise spectrum D(m,k);

[0020] Then, through:

[0021] the denoised speech spectrum Y(m,k) is calculated;

[0022] where θ X (m,k) is the phase of the noisy speech spectrum X(m,k);

[0023] After that, the inverse short-time Fourier transform is performed on Y(m,k) to obtain the denoised speech signal S dn [n].

[0024] As a further solution of the present invention, the voiceprint feature extraction method is as follows:

[0025] The linear prediction cepstral coefficient (LPCC) method is used to extract the voiceprint features of the denoised speech signal S dn [n]; the specific method is as follows:

[0026] First, perform linear prediction analysis on the speech signal;

[0027] By solving: the linear prediction coefficients a i ;

[0028] where i = 1, 2,..., p, p is the prediction order, and a0 = 1;

[0029] Then, calculate the cepstral coefficients c through the following recurrence formula k ;

[0030]

[0031] where k = 2, 3,..., p;

[0032] Subsequently, the cepstral coefficients c k constitute the voiceprint feature vector C = [c1, c2,..., c p .

[0033] As a further solution of the present invention, the speaker clustering method is as follows:

[0034] The voiceprint feature vectors of the pre-recorded known speakers are used as the voiceprint feature templates CM r , r = 1, 2,..., M, M represents the number of known speakers;

[0035] For the newly extracted voiceprint feature vector C, through: calculate its Euclidean distance D from each voiceprint feature template r ;

[0036] where C[k] and CM r[k] are respectively the voiceprint feature vector C and the voiceprint feature template CM r of the i-th element;

[0037] Then, classify the new speech signal corresponding to the new voiceprint feature vector C to the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs.

[0038] As a further solution of the present invention: wherein, if the Euclidean distance values between the new voiceprint feature vector C and all known voiceprint feature templates are greater than a preset distance threshold, then it is considered that the new voiceprint feature vector C is a new speaker not existing among the known speakers.

[0039] As a further solution of the present invention, the calculation method of the speech rate feature is as follows:

[0040] For each speaker, extract a speech segment with a specified duration TS, and then use the word segmentation technology in natural language processing to count the number of words NW in this speech segment;

[0041] Then through: Calculate the speech rate feature VS of the speaker.

[0042] As a further solution of the present invention, the extraction method of the intonation change feature is as follows:

[0043] For each speaker, extract a speech segment with a specified duration TS, and then divide this speech segment into several sub-segments, and use the autocorrelation function method to extract the fundamental frequency of each sub-segment; and according to the time trend, record the fundamental frequencies of the speech signals corresponding to each sub-segment as F0(v) in turn, where v = 1, 2,... f, and f represents the number of sub-segments;

[0044] Then through: Calculate the average change rate FB of the fundamental frequencies of each adjacent sub-segment, which is the intonation change feature.

[0045] As a further solution of the present invention: in the extraction of the intonation change feature, the extraction method of the fundamental frequency is as follows:

[0046] Select a sub-segment and mark the speech signal of this sub-segment as S y [n];

[0047] Then through: Calculate the autocorrelation function R(m) of the speech signal corresponding to this sub-segment;

[0048] In the formula, L is the length of the sub-segment;

[0049] Then based on the autocorrelation function R(m), extract the position of the first peak value and record it as m0, and the fundamental period T0 = m0;

[0050] Through: Calculate the fundamental frequency F0 of the speech signal corresponding to this sub - segment;

[0051] And so on, calculate the fundamental frequencies of the speech signals corresponding to each sub - segment.

[0052] As a further scheme of the present invention: The feature transformation and abstraction processing method of the hidden layer are as follows:

[0053] The neurons in the hidden layer use a pre - set weight matrix W ih to perform weighted summation with the feature vector of the input emotion, and add the bias vector b of the hidden layer h , and then perform a non - linear transformation through an activation function to obtain the output value h of the hidden layer;

[0054] Its expression is: h = A(W ih ×X + b h );

[0055] where X represents the emotion feature vector, and A represents the activation function;

[0056] The activation function selects the sigmoid function, and

[0057] As a further scheme of the present invention: The calculation method of the emotion category probability of the output layer is as follows:

[0058] Perform weighted calculation on the output value h of the hidden layer through a pre - set weight matrix W ho , and add the bias vector b of the output layer o to obtain the output result y;

[0059] Its expression is: y = W ho ×h + b o .

[0060] As a further scheme of the present invention: Among them, each neuron in the output layer corresponds to an emotion category, and its output value represents the probability that the speech signal belongs to this emotion category;

[0061] By comparing the output values of each neuron in the output layer, select the emotion category corresponding to the neuron with the largest value as the emotion prediction result of the emotion classification model for the input speech signal.

[0062] An AI speech emotion analysis platform based on a deep learning architecture, which is used for the AI speech emotion analysis method based on a deep learning architecture. This platform includes:

[0063] A signal acquisition unit: used to collect analog speech signals and convert them into digital speech signals through an analog - to - digital converter;

[0064] Noise reduction processing unit: It is used to perform short-time Fourier transform on the noisy speech signal to obtain the speech spectrum. At the same time, during the period when no one is speaking, it extracts the digital noise signal and performs short-time Fourier transform to obtain the noise spectrum. Through the noise reduction method based on spectrum subtraction, it obtains the noise-reduced speech spectrum and performs inverse short-time Fourier transform on it to obtain the noise-reduced speech signal;

[0065] Speech separation unit: It is used to extract the voiceprint features from the noise-reduced speech signal to obtain the voiceprint feature vector, and uses the clustering method based on Euclidean distance to cluster the voiceprint feature vectors of speakers, and classify the new speech signal into the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs;

[0066] Feature extraction unit: For each speaker, it extracts a speech segment of a specified duration and extracts its speech rate feature and intonation change feature from it;

[0067] Emotion classification unit: The input layer of the multi-layer perceptron is used to receive the emotion feature vector, its hidden layer performs feature transformation and abstraction processing on it, and its output layer calculates the emotion category probability, and determines the emotion prediction result by comparing the output values of the neurons in the output layer.

[0068] Advantages of the present invention:

[0069] In the present invention, the noisy speech signal and the noise signal are respectively processed by short-time Fourier transform, the noise-reduced speech spectrum is calculated, and inverse short-time Fourier transform is performed to obtain the noise-reduced speech signal, effectively removing the noise in the speech signal, improving the quality of the speech signal, and laying a good foundation for subsequent analysis.

[0070] In the present invention, the linear prediction cepstral coefficient (LPCC) method is used to extract the voiceprint features, and then the clustering method based on Euclidean distance is used to cluster the voiceprint feature vectors, which can accurately separate the speech signals of different speakers and can be applied to the speech emotion analysis scenario of multiple speakers, making the subsequent emotion analysis more targeted and accurate. And when the Euclidean distance values between the new voiceprint feature vector and all known voiceprint feature templates are greater than the pre-set distance threshold, a new speaker can be identified, enhancing the adaptability and recognition ability of the system to unknown speakers.

[0071] In the present invention, emotion features such as speech rate features and intonation change features are extracted. The speech rate feature is calculated by counting the number of words in a speech segment of a specified duration; the intonation change feature is obtained by dividing the speech segment into sub-segments, extracting the fundamental frequency of each sub-segment and calculating the average change rate of the fundamental frequencies of adjacent sub-segments. The extraction of multi-dimensional emotion features can more comprehensively reflect the emotion information contained in the speech and provide richer feature vectors for emotion classification.

[0072] In the present invention, a multi-layer perceptron is used as an emotion classification model. The input layer receives an emotion feature vector, the hidden layer performs feature transformation and abstraction processing on it, and the output layer calculates the emotion category probability. In this way, the input speech signal can be effectively emotion-classified, and the emotion prediction result is determined by comparing the output values of the neurons in the output layer, improving the accuracy and reliability of emotion classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The present invention will be further described below with reference to the accompanying drawings.

[0074] Figure 1 is a system block diagram of an AI speech emotion analysis platform with a deep learning architecture according to the present invention.

[0075] Figure 2 is a flowchart of an AI speech emotion analysis method with a deep learning architecture according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0077] Embodiment 1

[0078] Please refer to Figure 1 and Figure 2 As shown, the present invention is an AI speech emotion analysis method with a deep learning architecture, including the following steps:

[0079] Speech signal acquisition:

[0080] First, a microphone device is used to collect an analog speech signal, denoted as S(t), where t represents time; then the collected analog speech signal S(t) is converted into a digital speech signal Sd[n] through an analog-to-digital converter, where n is a discrete time series index;

[0081] Among them, the collected speech signal is the original analog signal;

[0082] In this embodiment, for example, in a multi-person meeting room scenario, the microphone will collect a sound signal mixed with the voices of all people, and this signal contains the speech information of different speakers and environmental noise, etc.;

[0083] Among them, the conversion process of the analog-to-digital converter is based on the Nyquist sampling theorem, and its sampling frequency fs satisfies fs≥2f max ,and fmax is the highest frequency component in the voice signal.

[0084] In this embodiment, assuming that the highest frequency of the voice signal is 8 kHz, then the sampling frequency is at least 16 kHz; after analog-to-digital conversion, the continuous analog signal becomes a discrete digital signal sequence, which is convenient for subsequent digital signal processing;

[0085] Speaker separation based on voiceprint features:

[0086] Step K1, voiceprint feature extraction:

[0087] Use the linear predictive cepstral coefficient (LPCC) method to extract voiceprint features from the noise-reduced voice signal S dn [n]; the specific method is as follows:

[0088] First, perform linear predictive analysis on the voice signal;

[0089] By solving: Calculate the linear prediction coefficient a i ;

[0090] In the formula, i = 1, 2,..., p, p is the prediction order, where a0 = 1;

[0091] Then, calculate the cepstral coefficient c through the following recurrence formula k ;

[0092]

[0093] In the formula, k = 2, 3,..., p;

[0094] Subsequently, the cepstral coefficient c k constitutes the voiceprint feature vector C = [c1, c2,..., c p ;

[0095] Step K2, speaker clustering:

[0096] Use the clustering method based on Euclidean distance to perform speaker clustering on the extracted voiceprint feature vector C, and the specific method is as follows:

[0097] Use the voiceprint feature vectors of the known speakers pre-recorded as the voiceprint feature templates CM r , r = 1, 2,..., M, M represents the number of known speakers;

[0098] For the newly extracted voiceprint feature vector C, through: Calculate its Euclidean distance D from each voiceprint feature template r ;

[0099] Among them, C[k] and CMr [k] are the i-th elements of the voiceprint feature vector C and the voiceprint feature template CM respectively r ;

[0100] Then, classify the new voice signal corresponding to the new voiceprint feature vector C into the speaker to whom the voiceprint feature template with the minimum Euclidean distance value belongs;

[0101] Among them, if the Euclidean distance values between the new voiceprint feature vector C and all known voiceprint feature templates are all greater than a pre-set distance threshold, then it is considered that the new voiceprint feature vector C is a new speaker not existing among the known speakers;

[0102] Emotional feature extraction:

[0103] Step G1, Speech rate feature calculation:

[0104] For each speaker, extract a speech segment with a specified duration TS, and then use the word segmentation technology in natural language processing to count the number of words NW in this speech segment;

[0105] In this embodiment, the word segmentation technology in natural language processing is a prior art, so it will not be elaborated here;

[0106] Then through: Calculate the speech rate feature VS of the speaker;

[0107] Step G2, Intonation change feature extraction:

[0108] For each speaker, extract a speech segment with a specified duration TS, and then divide this speech segment into several sub-segments, and use the autocorrelation function method to extract the fundamental frequency of each sub-segment; and in the time direction, record the fundamental frequencies of the speech signals corresponding to each sub-segment as F0(v) in turn, where v = 1, 2,... f, and f represents the number of sub-segments;

[0109] Then through: Calculate the average change rate FB of the fundamental frequencies of each adjacent sub-segment, which is the intonation change feature;

[0110] The extraction method of the fundamental frequency is as follows:

[0111] Select a sub-segment and mark the speech signal of this sub-segment as S y [n];

[0112] Then through: Calculate the autocorrelation function R(m) of the speech signal corresponding to this sub-segment;

[0113] In the formula, L is the length of the sub-segment;

[0114] Subsequently, based on the autocorrelation function R(m), the position of the first peak is extracted and denoted as m0, and the fundamental period T0 = m0;

[0115] By: Calculate the fundamental frequency F0 of the speech signal corresponding to this sub - segment;

[0116] And so on, calculate the fundamental frequencies of the speech signals corresponding to each sub - segment;

[0117] Emotion classification:

[0118] Based on the extracted emotion feature vectors, perform emotion classification according to the pre - set emotion classification rules, specifically as follows:

[0119] When VS > VSa and FB > FBa, it is determined that the speaker's emotion is excited;

[0120] When VS < VSb and FB < FBb, it is determined that the speaker's emotion is sad;

[0121] Wherein, VSa and VSb are pre - set speech rate thresholds; FBa and FBb are pre - set fundamental frequency change rate thresholds.

[0122] The AI speech emotion analysis method of the deep learning architecture proposed in Embodiment 1 collects analog speech signals through a microphone and converts them into digital signals. Based on the Nyquist sampling theorem, the accuracy of signal conversion is ensured, laying a foundation for subsequent processing. The linear prediction cepstral coefficient method is used to extract voiceprint features, and the clustering method based on Euclidean distance is combined to separate speakers, which can effectively distinguish different speakers and avoid interference caused by the mixing of multiple voices. In the emotion feature extraction link, by calculating the speech rate feature and intonation change feature, emotion information is extracted from dimensions such as speech rate and intonation fluctuation, and emotion classification is carried out according to the preset rules, realizing the preliminary analysis of speech emotion. This embodiment provides a complete and clear speech emotion analysis process, and all links from speech collection to emotion classification are closely connected, providing a basic framework and core idea for more complex speech emotion analysis schemes in the future.

[0123] Embodiment 2

[0124] Please refer to Figure 1 and Figure 2 As shown in

[0125] As the second embodiment of the present invention, when the present application is specifically implemented, compared with Embodiment 1, the technical solution of this embodiment is only different from that of Embodiment 1 in that this embodiment further includes a step of noise reduction processing:

[0126] First, for the noisy speech signal S d[n] Perform short-time Fourier transform;

[0127] The short-time Fourier transform formula is:

[0128] Where X(m,k) is the speech spectrum with noise, m represents the frame index, k represents the frequency index, N is the frame length, H is the frame shift, w[n] is the window function; j represents the imaginary unit, which is used to describe the phase information of the signal in frequency-domain analysis, helping to achieve the transformation from the time domain to the frequency domain and related processing;

[0129] Meanwhile, during the period when no one is speaking, extract the digital noise signal that only contains noise and no human voice, and perform short-time Fourier transform on it simultaneously to obtain the noise spectrum D(m,k);

[0130] Then, through:

[0131] Calculate the denoised speech spectrum Y(m,k);

[0132] Where θ X (m,k) is the phase of the speech spectrum with noise X(m,k);

[0133] After that, perform inverse short-time Fourier transform on Y(m,k) to obtain the denoised speech signal S dn [n].

[0134] Example 2 adds a noise reduction processing step on the basis of Example 1, and adopts a noise reduction method based on spectral subtraction. This method performs short-time Fourier transform on the noisy speech signal and the noise signal respectively, processes them in the frequency domain, and then obtains the denoised speech signal through inverse short-time Fourier transform. This process can effectively remove interference factors such as environmental noise in the speech signal and improve the quality of the speech signal. For the speech signal after noise reduction processing, when performing operations such as speaker separation and voiceprint feature extraction in the subsequent process, it can reduce the influence of noise on the feature extraction and classification results, make the extracted voiceprint features and emotional features more accurate, thereby improving the reliability and accuracy of speech emotion analysis, and enhancing the applicability of the entire speech emotion analysis method in the actual complex environment.

[0135] Example 3

[0136] Please refer to Figure 1 and Figure 2As shown in the figure, as the third embodiment of the present invention, in the specific implementation of this application, compared with the first and second embodiments, the technical solution of this embodiment lies in combining the solutions of the above-mentioned first and second embodiments. The difference between the technical solution of this embodiment and the first and second embodiments is only that in this embodiment, the sentiment classification step also uses a multi-layer perceptron as the sentiment classification model to classify the sentiment of the speaker;

[0137] Among them, the sentiment classification model includes an input layer, a hidden layer, and an output layer;

[0138] The input layer is used to receive the sentiment feature vectors extracted from the speech signal, that is, the speech rate feature and the intonation change feature;

[0139] Among them, the input layer contains multiple neurons, and each neuron corresponds to a different sentiment feature vector;

[0140] The hidden layer is used to perform feature transformation and abstraction processing on the sentiment feature vectors transmitted from the input layer. The feature transformation and abstraction processing methods are as follows:

[0141] The neurons in the hidden layer use a pre-set weight matrix W ih to perform weighted summation with the feature vectors of the input sentiment, and add the bias vector b of the hidden layer h , and then perform a non-linear transformation through an activation function to obtain the output value h of the hidden layer;

[0142] Its expression is: h = A(W ih ×X + b h );

[0143] Among them, X refers to the sentiment feature vector, and A refers to the activation function;

[0144] The activation function selects the sigmoid function, and

[0145] The output layer is used to calculate the sentiment category probability of the processing result after the feature transformation and abstraction processing from the hidden layer;

[0146] Among them, the output layer contains multiple neurons, and each neuron corresponds to a different sentiment category;

[0147] In this embodiment, the sentiment categories include happy, sad, angry, etc.;

[0148] The sentiment category probability calculation method is as follows:

[0149] Through a pre-set weight matrix W ho to perform weighted calculation on the output value h of the hidden layer, and add the bias vector b of the output layer o , to obtain the output result y;

[0150] Its expression is: y = W ho ×h+b o ;

[0151] Among them, each neuron in the output layer corresponds to an emotion category, and its output value represents the probability that the speech signal belongs to this emotion category;

[0152] Finally, by comparing the output values ​​of each neuron in the output layer, the emotion category corresponding to the neuron with the largest value is selected as the emotion prediction result of the emotion classification model for the input speech signal, thereby completing the task of speech emotion analysis;

[0153] In this embodiment, for example, if the output layer has 3 neurons, corresponding to happiness, sadness, and anger, respectively, and the output result of the output layer is y=[0.2, 0.6, 0.1], the predicted emotion is sadness.

[0154] Embodiment 3 introduces a multilayer perceptron as a sentiment classification model based on Embodiment 1 and Embodiment 2. The input layer of the multilayer perceptron receives sentiment feature vectors such as speech rate features and intonation change features. The hidden layer performs feature transformation and abstract processing through weighted summation, adding bias and activation function. The output layer calculates the probabilities of different sentiment categories and obtains the final sentiment classification result. Compared with the preset rule classification in Embodiment 1, this sentiment classification method based on deep learning can automatically learn and mine the complex relationship between sentiment features and sentiment categories, optimize model parameters through large-scale data training, and improve the accuracy and flexibility of sentiment classification. For complex and diverse speech emotion expressions, the multilayer perceptron can better adapt to different emotion patterns and effectively improve the performance and generalization ability of the speech emotion analysis system.

[0155] Embodiment 4

[0156] As the fourth embodiment of the present invention, when the present application is specifically implemented, compared with the first, second and third embodiments, the technical solution of this embodiment is to combine and implement the solutions of the above-mentioned first, second and third embodiments.

[0157] See also Figure 1 and Figure 2As shown in the figure, Example 4 integrates the technical solutions of Example 1, Example 2, and Example 3 to form a comprehensive and optimized AI speech emotion analysis method. It includes basic processes such as speech acquisition, digital signal conversion, and speaker separation based on voiceprint features. It also improves the quality of speech signals through noise reduction processing and uses the powerful deep learning ability of multi-layer perceptrons to achieve accurate emotion classification. This combined implementation method gives full play to the advantages of each part of the technology, enabling the entire speech emotion analysis system to achieve good results in various links such as signal processing, feature extraction, and emotion classification. It can effectively handle complex speech environments and diverse emotion expressions, improve the accuracy, reliability, and generalization ability of speech emotion analysis, has higher practical value and application prospects, and provides a more perfect and efficient solution for speech emotion analysis tasks in actual scenarios.

[0158] Please refer to Figure 1 and Figure 2 As shown in the figure, the present invention also provides an AI speech emotion analysis platform for a deep learning architecture, which is used for the AI speech emotion analysis method of the deep learning architecture. The platform includes:

[0159] Signal acquisition unit: used to acquire analog speech signals and convert them into digital speech signals through an analog-to-digital converter;

[0160] Noise reduction processing unit: used to perform short-time Fourier transform on the speech signal containing noise to obtain the speech spectrum, and at the same time extract the digital noise signal during the period when no one is speaking and perform short-time Fourier transform to obtain the noise spectrum. The noise-reduced speech spectrum is obtained through a noise reduction method based on spectrum subtraction, and the inverse short-time Fourier transform is performed on it to obtain the noise-reduced speech signal;

[0161] Speech separation unit: used to extract voiceprint features from the noise-reduced speech signal to obtain voiceprint feature vectors, and use a clustering method based on Euclidean distance to cluster the voiceprint feature vectors of speakers, and classify the new speech signal into the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs;

[0162] Feature extraction unit: for each speaker, extract a speech segment of a specified duration and extract its speech rate feature and intonation change feature from it;

[0163] Emotion classification unit: The input layer of the multi-layer perceptron is used to receive emotion feature vectors, its hidden layer performs feature transformation and abstraction processing on them, and its output layer calculates the emotion category probability, and determines the emotion prediction result by comparing the output values of the neurons in the output layer.

[0164] It should be noted that: all user data collected in this application is collected with the consent and authorization of the users, and the uses of the user data are legal and compliant, and the use and processing of the user data comply with the relevant laws, regulations and standards of the relevant regions.

[0165] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0166] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An AI voice emotion analysis method for a deep learning architecture, characterized in that, It includes the following steps: Speech signal acquisition: Use a microphone device to collect the analog speech signal S(t), and convert it into a digital speech signal S d [n], where t represents time and n is the discrete time series index; Noise reduction processing: For the speech signal S d [n] with noise, perform short-time Fourier transform to obtain the speech spectrum X(m,k). During the period when no one is speaking, extract the digital noise signal and perform short-time Fourier transform to obtain the noise spectrum D(m,k). Obtain the noise-reduced speech spectrum Y(m,k) through a noise reduction method based on spectrum subtraction, and perform inverse short-time Fourier transform on Y(m,k) to obtain the noise-reduced speech signal S dn [n]; Speaker separation based on voiceprint features: The linear prediction cepstral coefficient method is used to extract the voiceprint features of the noise-reduced speech signal S dn [n], obtaining the voiceprint feature vector C. A clustering method based on Euclidean distance is used to cluster the speakers of the voiceprint feature vector C, and the new speech signal is classified into the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs; Emotional feature extraction: For each speaker, extract a speech segment of a specified duration TS, and calculate the speech rate feature VS and the intonation change feature FB; Emotional classification: Use a multi-layer perceptron as the emotional classification model, which includes an input layer, a hidden layer, and an output layer. The input layer receives the emotional feature vector, the hidden layer performs feature transformation and abstraction processing on it, and the output layer calculates the probability of the emotional category. Determine the emotional prediction result by comparing the output values of the neurons in the output layer.

2. The AI voice emotion analysis method of the deep learning architecture according to claim 1, characterized in that, The specific method of noise reduction is as follows: First, perform a short-time Fourier transform on the noisy speech signal S d [n]. The short-time Fourier transform formula is as follows: In the formula, X(m,k) is the speech spectrum containing noise, m represents the frame index, k represents the frequency index, N is the frame length, H is the frame shift, w[n] is the window function, and j represents the imaginary unit, which is used to describe the phase information of the signal in frequency domain analysis; At the same time, during the period when no one is speaking, extract the digital noise signal that only contains noise and no human voice, and perform short-time Fourier transform on it at the same time to obtain the noise spectrum D(m,k); Then, by: Calculate the denoised speech spectrum Y(m,k); where θ X (m,k) is the phase of the noisy speech spectrum X(m,k); After that, perform the inverse short-time Fourier transform on Y(m,k) to obtain the denoised speech signal S dn [n].

3. The AI voice emotion analysis method of the deep learning architecture according to claim 2, wherein, The method of voiceprint feature extraction is as follows: The linear predictive cepstral coefficient (LPCC) method is used to extract the voiceprint features of the noise-reduced speech signal S dn [n]; the specific method is as follows: First, perform linear prediction analysis on the speech signal; By solving: the linear prediction coefficients a are calculated i ; In the formula, i = 1, 2,... p, p is the prediction order, where a0 = 1; Then, the cepstrum coefficients c are calculated by the following recurrence formula k ; In the formula, k = 2, 3,... p; Subsequently, the cepstral coefficients c k constitute the voiceprint feature vector C = [c1, c2,..., c p .

4. The AI voice emotion analysis method of the deep learning architecture according to claim 3, characterized in that, The method of speaker clustering is as follows: Use the pre-recorded voiceprint feature vector of the known speaker as the voiceprint feature template CM r , where r = 1, 2, …… M, and M represents the number of known speakers; For the newly extracted voiceprint feature vector C, by: Calculating its Euclidean distance D from each voiceprint feature template r ; Among them, C[k] and CM r [k] are the i-th elements of the voiceprint feature vector C and the voiceprint feature template CM r respectively; Then classify the new speech signal corresponding to the new voiceprint feature vector C into the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs; Among them, if the Euclidean distance value between the new voiceprint feature vector C and all known voiceprint feature templates is greater than the pre-set distance threshold, it is considered that the new voiceprint feature vector C is a new speaker that does not exist among the known speakers.

5. The AI voice emotion analysis method of the deep learning architecture according to claim 4, characterized in that, Speech rate feature VS Pass: Calculated; Among them, TS is the duration of the speech segment extracted for the speaker, and NW is the number of words in the speech segment.

6. The AI voice emotion analysis method of the deep learning architecture according to claim 5, characterized in that, The method of intonation change feature extraction is as follows: For each speaker, extract a speech segment of a specified duration TS, and then divide the speech segment into several sub-segments, and use the autocorrelation function method to extract the fundamental frequency of each sub-segment; and in the time direction, record the fundamental frequency of the speech signal corresponding to each sub-segment as F0(v) in turn, where v = 1, 2,... f, and f represents the number of sub-segments; Followed by: Calculate the average change rate FB of the pitch frequencies of each adjacent sub-segment, which is the intonation change feature.

7. The AI voice emotion analysis method of the deep learning architecture according to claim 6, characterized in that, In the extraction of intonation change features, the method of extracting the fundamental frequency is as follows: Select a sub - segment and label the speech signal of this sub - segment as S y [n]; Followed by: Calculating the autocorrelation function R(m) of the voice signal corresponding to the sub-segment; In the formula, L is the length of the sub-segment; Then, based on the autocorrelation function R(m), extract the position of the first peak value and record it as m0, and the fundamental period T0 = m0; Adopted by: Calculating the fundamental frequency F0 of the voice signal corresponding to the sub-fragment; And so on, calculate the fundamental frequency of the speech signal corresponding to each sub-segment.

8. The AI voice emotion analysis method of the deep learning architecture according to claim 6, characterized in that, The method of feature transformation and abstraction processing in the hidden layer is as follows: The neurons in the hidden layer use a preset weight matrix W ih to perform a weighted sum with the feature vector of the input sentiment, and add the bias vector b of the hidden layer h , and then perform a non-linear transformation through the activation function to obtain the output value h of the hidden layer; Its expression is: h = A(W ih ×X + b h ); Among them, X is referred to as the emotional feature vector, and A is referred to as the activation function; The activation function is selected as the sigmoid function, and 9. The AI voice emotion analysis method of the deep learning architecture according to claim 8, characterized in that, The calculation method of the emotional category probability in the output layer is as follows: Through the pre-set weight matrix W ho Perform weighted calculation on the output value h of the hidden layer and add the bias vector b of the output layer o , to obtain the output result y; Its expression is: y = W ho ×h + b o ; Among them, each neuron in the output layer corresponds to an emotional category, and its output value represents the probability that the speech signal belongs to this emotional category; By comparing the output values of each neuron in the output layer, select the emotional category corresponding to the neuron with the largest value as the emotional prediction result of the emotional classification model for the input speech signal.

10. An AI voice emotion analysis platform with a deep learning architecture, which is used to implement the AI voice emotion analysis method with the deep learning architecture described in any one of claims 1-9, and is characterized in that, The platform includes: Signal acquisition unit: used to acquire analog speech signals and convert them into digital speech signals through an analog-to-digital converter; Noise reduction processing unit: It is used to perform short-time Fourier transform on the speech signal containing noise to obtain the speech spectrum. At the same time, during the period when no one is speaking, it extracts the digital noise signal and performs short-time Fourier transform to obtain the noise spectrum. It obtains the denoised speech spectrum through a noise reduction method based on spectrum subtraction, and performs inverse short-time Fourier transform on it to obtain the denoised speech signal; Speech separation unit: It is used to extract the voiceprint features from the denoised speech signal to obtain the voiceprint feature vector. It uses a clustering method based on Euclidean distance to cluster the speakers of the voiceprint feature vector, and classifies the new speech signal into the speaker to which the voiceprint feature template with the smallest Euclidean distance value belongs; Feature extraction unit: For each speaker, it is used to extract a speech segment of a specified duration and extract its speech rate feature and intonation change feature from it; Emotion classification unit: The input layer of the multi-layer perceptron is used to receive the emotion feature vector. Its hidden layer performs feature transformation and abstraction processing on it. Its output layer calculates the emotion category probability, and determines the emotion prediction result by comparing the output values of the neurons in the output layer.

Citation Information

Cited By

  • Voice content sentiment analysis method and system based on deep learning

    CN122392578A

  • A speech content emotion analysis method and system based on deep learning

    CN122392578B