Echo cancellation and noise reduction method based on deep learning

By combining time-frequency domain transformation and deep learning methods of emotion recognition networks, the problem of insufficient emotion recognition in echo cancellation technology is solved, the emotional information of speech is retained during the echo cancellation process, and the naturalness and emotion transmission ability of speech are improved.

CN120612952AInactive Publication Date: 2025-09-09SHENZHEN ZHILIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511114300.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing echo cancellation technology ignores the emotional component of speech, resulting in a lack of emotional color in speech after echo cancellation, affecting user experience and emotional communication. In particular, it is difficult to accurately capture and convey the user's emotional state in application scenarios that require emotional understanding.

Method used

Through a deep learning-based method, combined with time-frequency domain transformation, emotion recognition network and deep neural network, the time-frequency features of the echo speech signal are extracted, and the emotional fluctuations are analyzed using convolutional neural network and LSTM model, generating multimodal feature representation, and designing a joint loss function to optimize echo cancellation and speech emotion retention, and perform speech quality and emotion consistency evaluation.

Benefits of technology

It achieves the retention of emotional information of speech during the echo cancellation process, improves the naturalness and intelligibility of speech, ensures the accurate transmission of emotions, and improves the emotion recognition and response capabilities of the voice interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612952A_ABST
    Figure CN120612952A_ABST
Patent Text Reader

Abstract

The invention discloses an echo cancellation noise reduction method based on deep learning, relates to the technical field of echo noise reduction optimization, and is used for solving the problem of poor emotion capture in an echo cancellation process. By combining time-frequency domain transformation, an emotion recognition network and a deep learning model, emotion information of voice is effectively reserved, time-frequency features of echo voice signals are extracted, emotion fluctuation is analyzed by using a convolutional neural network and LSTM, and multi-modal feature representation is generated, so that it is ensured that the voice emotion is accurately expressed while echoes are removed, and the voice emotion recognition efficiency is improved. Based on echo signal estimation and emotion perception adjustment of a deep neural network, balance of echo cancellation and emotion consistency is realized, a joint loss function is designed to optimize echo cancellation and emotion retention, definition and emotion accuracy of output voice are ensured through quality evaluation, an echo cancellation effect is improved, emotion information transmission is enhanced, and the user experience is improved. And the emotion recognition and response capability of the voice interaction system in man-machine conversation is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of echo cancellation and noise reduction, and more specifically, to an echo cancellation and noise reduction method based on deep learning. Background Art

[0002] Echo occurs when a user's voice signal is captured by a microphone and then returned to the receiving end after a period of time due to sound wave reflection or changes in the transmission path during a call, resulting in a repeated audio signal. This phenomenon commonly occurs in scenarios such as remote communications, video calls, and conferencing systems. The presence of echo seriously affects voice quality, causing delays in information transmission and even unclear communication between the two parties. The goal of echo cancellation is to remove or reduce echo signals to improve call quality during voice communications.

[0003] Deficiencies in existing technologies: Echo cancellation technology usually focuses on removing echoes or noise interference in audio signals, but ignores the emotional component of speech. This makes the speech become mechanical and lacking in emotion after echo cancellation, seriously affecting the user's auditory experience and emotional communication. Especially in some application scenarios that require emotional understanding (such as voice assistants, customer service, mental health monitoring, etc.), the conflict between echo cancellation and emotion recognition makes it difficult for the system to accurately capture and convey the user's emotional state, which in turn affects the quality and effectiveness of voice interaction. Existing echo cancellation technologies mostly use traditional single-task learning models and fail to effectively combine the optimization requirements of the two tasks of echo cancellation and emotion recognition. Therefore, they have great limitations in multi-task processing and multi-task coordination. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the following solution is provided to solve the problem of poor emotion capture in the echo cancellation process in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions: A deep learning-based echo cancellation and noise reduction method comprises the following steps: The time-frequency features of the echo speech signal are extracted through time-frequency domain transformation, and the emotional state is identified by combining the emotion recognition network. The local features of the signal are extracted using a convolutional neural network, and the emotional fluctuations are analyzed through the LSTM model. The emotional label is then embedded in the echo cancellation network to generate a multimodal feature representation. A deep neural network is trained based on multimodal features to estimate the echo signal, and the echo removal signal is adjusted through an emotion perception network to preserve the emotional information of the speech. Based on the designed joint loss function, the system optimizes echo cancellation and speech emotion preservation, and evaluates the echo cancellation and noise reduction quality of the output speech signal. The output speech signal is evaluated for voice quality and emotional consistency, and different signals are generated. The quality of echo cancellation and noise reduction is determined based on the different signals.

[0006] In a preferred embodiment, the time-frequency features of the echo speech signal are extracted based on the time-frequency domain transformation, and the emotional state is identified in combination with the emotion recognition network. The local features of the signal are extracted using a convolutional neural network, the emotional fluctuations are analyzed through the LSTM model, and the emotional label is embedded in the echo cancellation network to generate a multimodal feature representation. The specific steps are as follows: For the input speech signal with echo, the time-frequency domain is transformed, the time-frequency features of the speech signal are extracted using short-time Fourier transform, and the spectrum is obtained according to the short-time Fourier transform extraction; Use the LSTM-based emotion classification model as the emotion recognition network to identify the emotional state of the speech signal, combine the time-frequency features with the emotion markers in the speech, and output the emotion label for each time period; The emotion label of each time period is combined with the spectrum graph and transformed into the time-frequency domain. The resulting time-frequency graph is used as input and the local features of the time-frequency graph are extracted using a convolutional neural network. The convolution layer extracts different local features of the time-frequency signal through convolution operations, while retaining the key features of the speech signal through pooling operations.

[0007] In a preferred embodiment, emotional fluctuations are analyzed by an LSTM model, and emotional labels are embedded in an echo cancellation network to generate a multimodal feature representation. The specific steps include: An LSTM network is used to capture the temporal relationships in speech signals and model the emotional fluctuations in speech. The input of the emotion recognition network is the local features of the time-frequency graph extracted from the convolutional neural network. The emotional state in the speech is identified, and each time point corresponds to an emotional state category, resulting in a sequence of emotional state labels. Convert the emotion label into an emotion vector and embed it into the network input of the echo cancellation; The emotion embedding vector is combined with the time-frequency features extracted from the convolutional neural network for feature splicing and fusion. The emotion features and time-frequency features are spliced ​​along the time axis to generate a multimodal feature representation.

[0008] In a preferred embodiment, a deep neural network is trained based on multimodal features to estimate the echo signal, and the echo removal signal is adjusted by an emotion perception network to retain the speech emotion information. The specific steps are as follows: A deep neural network is used as the echo estimation model, with the input being the fused features after multimodal feature extraction and emotion embedding, and the output being the time series of the echo part. Use the emotion perception network to adjust the speech signal after echo removal, and perform denoising and feature extraction preprocessing on the input echo estimation signal and emotion features; The emotion perception network processes the preprocessed echo signal to generate an echo-removed signal, combines the echo-removed signal and the emotion feature through weighted sum, and uses it to adjust the speech signal after echo removal to retain the speech emotion information.

[0009] In a preferred embodiment, the optimization of echo cancellation and speech emotion preservation according to the designed joint loss function includes the following steps: The joint loss function includes echo cancellation loss and emotion preservation loss; Echo cancellation loss minimizes the difference between the speech signal after echo removal and the ideal echo-free signal, obtaining the ideal echo-free speech signal , calculate the speech signal after removing the echo The distance between the ideal speech signal and the echo cancellation loss function is calculated using the minimum mean square error: ; The emotion preservation loss is to minimize the emotional difference between the speech after echo cancellation and the original speech. The degree of preservation of the emotional features is determined by the cosine similarity of the emotional features. The emotion preservation loss function is: ,in, Calculate the cosine similarity function; The joint loss function combines the echo cancellation loss with the emotion preservation loss in a weighted manner. The joint loss function is: ,in, and are the weight coefficients of echo cancellation loss and emotion preservation loss respectively.

[0010] In a preferred embodiment, performing echo cancellation and noise reduction quality evaluation on the output speech signal includes the following steps: Evaluate from two aspects: speech quality evaluation and emotion consistency evaluation, and obtain the speech quality index generated by the speech quality evaluation process and the emotion consistency index generated by the emotion consistency evaluation process; The obtained speech quality index and emotion consistency index are combined to generate a noise reduction quality coefficient; The voice quality index and emotion consistency index are both proportional to the noise reduction quality coefficient.

[0011] In a preferred embodiment, performing voice quality assessment and emotion consistency assessment analysis on the output voice signal and generating different signals, and determining the echo cancellation noise reduction quality according to the different signals, includes the following steps: Compare the generated noise reduction quality coefficient with the set echo noise reduction threshold; If the noise reduction quality coefficient is greater than or equal to the echo noise reduction threshold, an echo noise reduction stable signal is generated without additional noise reduction processing; If the noise reduction quality coefficient is less than the echo noise reduction threshold, an echo noise reduction fluctuation control signal is generated to trigger adaptive adjustment, provide real-time feedback on the current noise reduction status, and start the noise reduction adjustment mechanism.

[0012] The technical effects and advantages of the deep learning-based echo cancellation and noise reduction method of the present invention are as follows: The present invention realizes efficient echo cancellation and speech emotion retention by combining time-frequency domain transformation, emotion recognition network and deep learning model. By extracting the time-frequency features of the echo speech signal and using convolutional neural network and LSTM to analyze emotion fluctuations, a multimodal feature representation is generated, thereby retaining the emotional information of the speech during the echo cancellation process. While removing the echo, it can ensure the accurate expression of speech emotion and effectively improve the naturalness and intelligibility of the speech. At the same time, the echo signal estimation and emotion perception adjustment based on the deep neural network make the output speech achieve a balance between echo cancellation and emotion consistency. Then, by designing a joint loss function, the echo cancellation and emotion retention are optimized, and the quality of the speech signal is evaluated to ensure the clarity and emotional accuracy of the output signal. This not only improves the effect of echo cancellation, but also enhances the transmission of emotional information, which helps to improve the emotion recognition and response capabilities of the voice interaction system in human-computer dialogue. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a flow chart of an echo cancellation and noise reduction method based on deep learning in the present invention. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0015] In order to achieve the above objectives, Figure 1 A structural diagram of an echo cancellation and noise reduction method based on deep learning is given in the present invention, which specifically includes the following steps: The time-frequency features of the echo speech signal are extracted through time-frequency domain transformation, and the emotional state is identified by combining the emotion recognition network. The local features of the signal are extracted using a convolutional neural network, and the emotional fluctuations are analyzed through the LSTM model. The emotional label is then embedded in the echo cancellation network to generate a multimodal feature representation. A deep neural network is trained based on multimodal features to estimate the echo signal, and the echo removal signal is adjusted through an emotion perception network to preserve the emotional information of the speech. Based on the designed joint loss function, the system optimizes echo cancellation and speech emotion preservation, and evaluates the echo cancellation and noise reduction quality of the output speech signal. The output speech signal is evaluated for voice quality and emotional consistency, and different signals are generated. The quality of echo cancellation and noise reduction is determined based on the different signals.

[0016] Step 1: Perform multimodal feature extraction and emotion embedding. The primary task in processing echo speech signals is to transform the speech signals into time-frequency domains, converting the time-domain signals into more expressive time-frequency features. Because echoes usually appear as fluctuations in frequency components, the frequency distribution and variation patterns of the speech and echo signals are determined through time-frequency transformation. The specific steps are as follows: Perform preprocessing and time-frequency feature extraction on the input speech signal. For the input speech signal with echo, perform time-frequency domain transformation to extract more representative frequency domain features from the time domain signal. Use short-time Fourier transform (STFT) or Mel-frequency cepstral coefficient (MFCC) to extract the time-frequency features of the speech signal. STFT maps the speech signal from the time domain to the time-frequency domain to capture the distribution of the signal at different times and frequencies and obtain the instantaneous spectrum of the speech signal. Given the input signal x(t), the calculation process of short-time Fourier transform is as follows: , where x(t) is the input time domain signal, is a sliding window function, f is a frequency variable, X(t, f) represents the local characteristics of the signal at time t and frequency f, Indicates the position or moment of the time window, that is, the center position of the window function; j is an integer, indicating the frequency component in the frequency domain; STFT can be used to extract spectrograms as time-frequency features for subsequent neural network input. These features can effectively represent the dynamic changes of speech signals in the frequency domain and capture the effects of echo and noise.

[0017] Embedding emotional information into the echo cancellation process requires identifying the emotional state of the speech signal. First, the emotional state of the speech signal is identified using an emotion recognition network, such as a long short-term memory (LSTM)-based emotion classification model. The LSTM-based emotion classification model combines time-frequency features with the emotional markers in the speech to output an emotion label for each time segment. The emotion labels for each time period are combined with the spectrogram, and the resulting time-frequency graph is transformed into the time-frequency domain. This graph is then used as input for further processing using a convolutional neural network (CNN). CNN is a powerful tool that can automatically extract spatial features of input signals. By convolving the signal in local regions using convolution kernels, it can extract local features from speech signals and capture deeper speech information through multi-layer convolution. The convolution operation extracts the local features of the time-frequency graph by convolving the convolution kernel with the input signal. Let the time-frequency graph be S(t, f). The convolution operation process is: ,in, is the weight of the convolution kernel, K and L are the sizes of the convolution kernel, and Z(t, f) is the local feature of the time-frequency graph obtained after the convolution operation. The convolution layer extracts different local features of the time-frequency signal through multiple convolution operations, and reduces the dimension and computational complexity of the feature map through pooling operations, further retaining the key features of the speech signal. Speech signals carry not only echo information but also emotional information. In practical applications, the emotional state of speech has a significant impact on the naturalness and clarity of speech. Therefore, in the process of echo cancellation, emotional information needs to be preserved. The input of the emotion recognition network is the high-dimensional features from the CNN layer (local features of the time-frequency graph). The network obtains the emotional state of the speech by modeling the temporal information of the speech signal. The LSTM network is used to capture the temporal relationship in the speech signal, thereby modeling the emotional fluctuations in the speech. The input of the emotion recognition network is the local features Z(t, f) of the time-frequency graph extracted from the CNN. The LSTM model can identify the emotional state in the speech (such as anger, happiness, sadness, etc.). Based on the input features Z(t, f), its output is a sequence of emotional states, with each time point corresponding to an emotional state category. The recursive formula of LSTM is: Input gate: , forget gate: , output gate: , cell state update : , the output at the current moment: ,in, is the output of the LSTM unit at time t, is the feature from CNN, is the weight matrix of LSTM, As a bias term, the modeling of time series features through the LSTM network can effectively capture the emotional fluctuations in speech and generate an emotional state label sequence; Use the sentiment embedding layer to label the emotional state Convert it into an emotion vector and embed it into the echo cancellation network input. The emotion embedding layer converts the emotion label into a low-dimensional vector representation E(t), so that the echo cancellation network can not only focus on the time-frequency characteristics of the speech, but also combine the emotion information to remove the echo. The sentiment embedding process can be achieved through the embedding matrix: ,in, Indicates emotional labels The sentiment feature vector E(t) is generated by mapping it into a low-dimensional space through the embedding matrix; The emotion embedding vector is fused with the time-frequency features extracted from CNN. The fusion method uses feature splicing or weighted fusion to jointly represent the features of the two and generate a multimodal feature representation for the joint expression of time-frequency features and emotion features. The splicing method splices the emotion features and time-frequency features along the time axis to form an expanded feature vector: ,in, It represents the final fused feature vector, which contains the time-frequency information and emotional features of the speech. It serves as the input of the subsequent echo cancellation network to ensure that the influence of emotional factors can be considered in the echo cancellation process.

[0018] Through multimodal feature extraction and emotion embedding, this step provides richer input features for the echo cancellation task. By combining time-frequency features with emotion features, echo cancellation can not only remove echoes more accurately, but also preserve the emotional fluctuations of speech, avoiding the loss of emotional information.

[0019] Step 2: Perform echo cancellation and dynamic adjustment of emotion perception. That is, after completing multimodal feature extraction and emotion embedding, generate an echo estimation model to separate the echo component from the input echoed speech signal. The specific steps are as follows: A deep neural network (DNN) is used as the echo estimation model. The input of the model is the fusion feature (multimodal feature) after multimodal feature extraction and emotion embedding, which contains both time-frequency features and emotion information. By training the network, an echo signal is estimated. , which represents the echo portion removed from the echoed input signal. The output of the echo estimation model is the time series of the echo portion: ,in, is the echo estimation function of the deep neural network, It is a fusion feature. After being input into the DNN model, the network learns the echo component in the signal through multi-layer nonlinear mapping. The estimation of this echo component not only depends on the time-frequency features, but also appropriately retains the emotional fluctuations in the speech according to the emotional characteristics. Perform dynamic echo adjustment combined with emotion perception, introduce the dynamic adjustment mechanism of emotion perception, and fine-tune the speech signal after echo removal through an emotion-aware network (EAN). The input of the emotion-aware network is the echo signal. and the emotional feature E(t) extracted in the previous step, the output of which is the fine-tuned echo removal signal , that is, the speech signal that still maintains emotional consistency after removing the echo; The processing of the emotion perception network is as follows: Preprocess the echo signal and emotional features, including denoising and feature extraction; The task of the echo cancellation and emotion perception fusion network is to remove the echo component from the echo signal while retaining the emotional information in the speech signal. To this end, the network first processes the echo signal to generate a signal with the echo removed, and combines the extracted emotional features with the signal to maintain the emotional information of the speech. The signal after the echo is removed and the emotional features are combined by weighted sum. Assume is the characteristic representation of the signal after echo removal, It is the representation of emotional characteristics, and the final output signal It can be expressed by the following formula: ,Among them, the SignalProcessing operation may be based on a weighting mechanism that ,combines the features of the echo signal and the emotional ,features to adjust the speech signal after echo removal; We optimize the balance of emotion preservation to ensure that the emotional characteristics of the speech (such as pitch and tone) remain consistent while the echo is removed. We introduce a loss function to calculate the emotional similarity between the speech signal after echo removal and the original speech signal. The loss function for emotion preservation may include the following: ,in, is the original speech signal, the loss function Evaluate the emotional retention of the speech signal after echo removal and the original speech signal, It is a metric function used to measure the difference in speech emotional features, and emotional features can be extracted by emotion recognition models (such as emotion classifiers); In this process, the joint optimization of echo cancellation and emotion perception ensures that the final output not only removes the echo, but also effectively preserves the emotional information of the speech. For example, if the input echo signal is a speech signal with anxiety, the goal of the emotion perception network is to remove the echo while maintaining the anxious tone, pitch, and emotional characteristics. In this way, the network can more accurately preserve the emotion in the speech while eliminating unwanted echo noise.

[0020] Step 3: perform joint loss function optimization and adaptive training analysis of the echo; In the deep learning tasks of echo cancellation and emotion preservation, only focusing on the clarity of the speech signal while ignoring the influence of the preservation of emotional features can easily lead to excessive noise reduction in the echo cancellation process. Therefore, a joint loss function is designed to simultaneously optimize the echo removal effect and the emotion preservation effect. The joint loss function consists of two main parts: Echo cancellation loss: This loss is used to measure the effectiveness of echo cancellation. The goal is to minimize the difference between the speech signal after echo removal and the ideal echo-free signal. To calculate this loss, you first need to obtain the ideal echo-free speech signal. , and then calculate the speech signal after removing the echo The distance from the ideal speech signal. For example, the signal reconstruction error or the minimum mean square error (MSE) can be used as a metric: ,in, is the de-echoed speech signal generated by the model, The target anechoic signal is the MSE loss function, which minimizes the echo portion of the speech by continuously optimizing the model parameters, thereby improving the clarity and intelligibility of the speech signal. This loss function encourages the model to learn to remove the echo and make the output signal as close to the ideal anechoic speech as possible. Emotion-preserving loss: The emotion-preserving loss is designed to minimize the emotional difference between the echo-cancelled speech and the original speech. To this end, the difference in emotional characteristics between the speech signal after echo removal and the ideal speech can be measured through the embedding layer of the emotional feature E(t). This vector represents the emotional color of the speech, such as happiness, anger, sadness, etc. Specifically, the similarity measure of sentiment features (such as cosine similarity or sentiment prediction error) is used to measure the degree of preservation of sentiment features. The loss function can be defined as: , the calculation formula of cosine similarity is: ,This loss function ensures that the speech signal after echo cancellation retains the emotional characteristics of the original speech; The joint loss function combines the echo removal loss and the emotion preservation loss in a weighted manner, thereby achieving a balance between echo removal and emotion preservation during training. The specific joint loss function is as follows: ,in, and These are the weight coefficients for echo cancellation loss and emotion preservation loss, respectively. By adjusting these two coefficients, the model can find the optimal balance between echo removal and emotion preservation. Depending on the application scenario, these two coefficients can be dynamically adjusted to optimize the loss function, maximizing echo cancellation while minimizing emotion loss. Adaptive training analysis of the joint loss function is performed. This involves adjusting the weights of various components of the loss function to further optimize the effects of echo cancellation and emotion preservation. Since the difficulty of echo cancellation is closely related to factors such as the complexity of the speech signal, the intensity of emotion, and background noise, the training strategy and loss function weights need to be dynamically adjusted based on these factors during training. To improve the generalization ability of the model, we optimized the selection of training data. The training data should not only include speech signals with echoes, but also include variations in emotional tone, background noise, and speech speed. This diverse data can help the model better adapt to various real-world scenarios and improve its ability to eliminate echoes and retain emotion. During the training process, the weights in the joint loss function are adaptively adjusted according to the current echo cancellation and emotion retention effects of the model. By monitoring the emotional feature differences between the speech signal after echo removal and the original speech, combined with the effect of echo removal, the loss function is dynamically adjusted. and For example, in the early stages of training, you may need to focus more on echo removal (increasing ), and as the model performance improves, it can be appropriately increased To enhance the ability to retain emotions; To avoid gradient explosion and instability during optimization, gradient clipping is employed. This technique limits the maximum value of the gradient to a certain range, preventing excessive gradients from causing unstable model parameter updates and ensuring stable model convergence during training. Adaptive optimization algorithms, such as Adam or RMSprop, are introduced to dynamically adjust the learning rate, further accelerating model convergence. The selection of optimization algorithms and parameter adjustments can be updated in real time based on the model's training results, ensuring a balanced approach to echo cancellation and emotion preservation.

[0021] Step 4: Perform joint optimization of dynamic echo cancellation and emotion restoration. After performing echo cancellation and emotion restoration, in order to ensure that the model processing effect achieves the expected goal, it is necessary to conduct a quality assessment of the final output. The specific steps are as follows: The dynamic nature of echoes means they are time-varying. Especially in real-time speech processing scenarios, the generation and attenuation of echoes are closely related to environmental factors and speech content. As speech content changes, so too do emotional characteristics. Therefore, it is necessary to dynamically adjust the output signal's parameters, such as pitch, speech rate, and tone, based on the input emotional characteristics and the current echo removal signal. After joint optimization of dynamic echo estimation and emotion restoration, the final speech signal will complete the dual tasks of echo removal and emotion restoration. The final output signal is then evaluated for echo cancellation and noise reduction quality to ensure that the echo removal and emotion restoration effects meet expectations. Evaluate from two aspects: speech quality evaluation and emotion consistency evaluation, and obtain the speech quality index generated by the speech quality evaluation process and the emotion consistency index generated by the emotion consistency evaluation process; The Speech Quality Index (SQI) measures the quality of speech after echo cancellation or noise reduction. It quantifies the quality of speech signals primarily in terms of signal distortion and signal-to-noise ratio. Specifically, the SQI assesses the improvement or loss in clarity, naturalness, and intelligibility of speech signals after echo cancellation or noise reduction. Signal distortion measures the similarity of the signal before and after echo cancellation. Lower distortion indicates that the signal retains a higher level of original quality after echo removal. The signal-to-noise ratio (SNR) measures the clarity of the signal and the impact of noise. A higher SNR indicates a clearer voice signal and less impact from noise. The SQI comprehensively considers distortion and signal-to-noise ratio to quantify the clarity and naturalness of the voice signal after echo cancellation. The SQI can be used to compare the effectiveness of different echo cancellation or noise reduction algorithms and select the optimal solution. The logic for obtaining the voice quality index is as follows: Get the original speech signal before speech processing, get the noise components before and after speech echo cancellation, get the speech signal after speech echo cancellation, and calculate the signal-to-noise difference ratio: , where is the amplitude value of the original speech signal at time t, n(t) is the amplitude value of the noise signal at time t, and T is the time length of the signal or the number of sampling points; the signal distortion is calculated as follows: , where is the amplitude value of the speech signal after echo cancellation at time t. The speech quality index is calculated based on the signal-to-noise difference ratio and signal distortion: .

[0022] It should be noted that the original speech signal is the original speech signal before processing, which can be obtained through recording equipment or a speech database; the noise component refers to the echo part or background noise part before and after echo cancellation, which is usually obtained through signal separation technology. For example, algorithms such as adaptive filtering or minimum mean square error (LMS) can be used to separate the noise component from the original signal.

[0023] The Emotional Consistency Index (ECI) measures whether the emotional characteristics of speech signals (such as pitch, speaking rate, and intonation) are effectively preserved after echo cancellation or noise reduction, and whether they have undergone significant changes during the echo removal process. The Emotional Consistency Index primarily assesses the emotional consistency of speech after echo cancellation, ensuring that the emotional information in the speech is not lost or distorted. Emotional similarity refers to the degree of similarity between the emotional characteristics of the de-echoed speech and the original speech, and is usually measured by features such as pitch, volume, and speaking rate. Emotional distortion measures the degree of change in emotional characteristics during the echo cancellation process, and is mainly evaluated by the change in features such as pitch, volume, and speaking rate. The emotional consistency index directly reflects whether the emotional characteristics of the speech are preserved after echo cancellation. Echo cancellation processing should clarify the speech without losing the emotional characteristics. The emotional consistency index can quantify this effect. In voice applications, especially in scenarios involving emotional expression (such as customer service and smart assistants), emotional consistency is crucial to user experience. The emotional consistency index can be used as an emotional consistency monitoring tool to ensure that emotional information is not lost or distorted during speech processing. The logic for obtaining the sentiment consistency index is as follows: Get the instantaneous energy of the speech signal and calculate the volume value: , where V(t) is the volume value at time t, N is the total time length of the sliding window, and y(t+n) is the amplitude of signal y(t) at time point t+n; Get the spectrum of the speech signal and calculate the pitch: , where Y(f, t) is the spectrum amplitude of the signal at frequency f and time t; Get the speech rate of the voice. The calculation expression is: , where R(t) is the speaking speed at time t, is the number of syllables at time t, is the number of pronunciation units per unit time, is the length of the time frame analyzed; Get the emotional feature vector of the original speech and the speech emotion feature vector after echo cancellation , calculate the sentiment similarity: , calculate the sentiment consistency index: .

[0024] It should be noted that volume is typically assessed by calculating the instantaneous energy of a signal. This energy reflects the strength of a signal at a given moment, representing the square of the signal's amplitude; a larger energy indicates a higher volume. Pitch estimation is typically determined using a short-time Fourier transform (SFT), which performs time-frequency analysis on the signal's spectrum, extracting the dominant frequency component as pitch. Speech rate, typically calculated by phoneme detection algorithms to measure the number of pronunciation units per second, reflects the speed of speech and is generally based on the recognition of phonemes or syllables. A common method for calculating speech rate is to count the number of syllables or phonemes per unit time.

[0025] The obtained speech quality index and emotion consistency index are combined to generate the noise reduction quality coefficient. The noise reduction quality coefficient expression is: , where 、 is the preset proportional coefficient of the voice quality index and the emotional consistency index, and 、 Both are greater than 0.

[0026] The larger the voice quality index and the emotion consistency index, the larger the noise reduction quality coefficient generated by the joint process. The larger the noise reduction quality coefficient, the more effective the removal of echo and noise in the environment, reducing interference in the voice signal. As a result, the voice signal after echo cancellation is clearer, background noise is significantly reduced, and the purity of the voice is greatly improved. This means that users can hear the speech more easily and clearly, and the original characteristics of the voice (such as pronunciation clarity and sound quality) can be preserved as much as possible while removing the echo, avoiding obvious sound quality distortion caused by the noise reduction process. A higher noise reduction quality coefficient means that the echo removal process does not over-smoothe or reduce the key frequency components of the voice, ensuring the naturalness and clarity of the voice. A smaller noise reduction quality coefficient indicates that the echo cancellation algorithm is not performing effectively at removing echoes. Some echoes may still remain in the speech signal, resulting in a mixture of echo and noise. This residual echo can make speech sound unclear, affecting the listening experience. It also fails to effectively preserve the emotional characteristics of speech (such as pitch, speaking rate, and volume). Emotional variations in speech (such as tone and emotional intensity) may be lost or altered due to excessive noise reduction, leading to significant distortion of emotional information. This is particularly true in scenarios sensitive to emotional expression (such as customer service and emotional speech recognition). This can lead to poor emotional consistency in speech, making it difficult for users to hear clearly what is being said or causing deviations in emotional expression, resulting in a poor user experience. In emotionally interactive applications, it becomes even more difficult for users to understand speech and provide emotional feedback.

[0027] Comparing the generated noise reduction quality coefficient with a preset echo noise reduction threshold to generate an echo noise reduction stabilization signal and an echo noise reduction fluctuation control signal; After obtaining the noise reduction quality coefficient, compare the noise reduction quality coefficient with the echo noise reduction threshold; If the noise reduction quality coefficient is greater than or equal to the echo noise reduction threshold, an echo noise reduction stability signal is generated, indicating that the echo cancellation has successfully removed most of the echo components and the noise reduction effect has achieved the expected goal. The echo interference in the speech signal has been effectively eliminated, and the remaining echo components are small or completely eliminated, so no additional noise reduction processing is required. If the noise reduction quality coefficient is less than the echo noise reduction threshold, an echo noise reduction fluctuation control signal is generated. However, echo components may remain in the speech signal, and the denoising process may not completely remove the echo. This may be due to inaccurate noise or echo estimation, or strong echo interference, which prevents the noise reduction algorithm from effectively identifying and eliminating the echo. Oversmoothing or attenuation may occur during echo removal, resulting in degraded speech quality and a loss of emotional characteristics such as pitch, volume, or speaking speed. Excessive noise reduction may make speech sound unnatural or lack emotional expressiveness.

[0028] It is necessary to trigger adaptive adjustment through echo noise reduction fluctuation control signal, provide real-time feedback on the current noise reduction status of the system, and start the noise reduction adjustment mechanism. For example, the system can automatically optimize the noise reduction algorithm parameters, adjust the model training strategy or adjust the signal processing process based on the comparison results of the real-time calculated noise reduction quality coefficient and the echo noise reduction threshold to make the echo cancellation effect more stable and reliable, and guide the noise reduction algorithm to perform adaptive adjustment so that it can quickly respond to changes in the echo or noise environment and optimize the echo cancellation effect and signal quality.

[0029] It should be noted that the relevant threshold information in this embodiment is pre-set by professionals and will not be explained in detail here. Some parameter English letters in the embodiments have the same situation, but different meanings are explained when used, and will not be explained one by one here.

[0030] The present invention realizes efficient echo cancellation and speech emotion retention by combining time-frequency domain transformation, emotion recognition network and deep learning model. By extracting the time-frequency features of the echo speech signal and using convolutional neural network and LSTM to analyze emotion fluctuations, a multimodal feature representation is generated, thereby retaining the emotional information of the speech during the echo cancellation process. While removing the echo, it can ensure the accurate expression of speech emotion and effectively improve the naturalness and intelligibility of the speech. At the same time, the echo signal estimation and emotion perception adjustment based on the deep neural network make the output speech achieve a balance between echo cancellation and emotion consistency. Then, by designing a joint loss function, the echo cancellation and emotion retention are optimized, and the quality of the speech signal is evaluated to ensure the clarity and emotional accuracy of the output signal. This not only improves the effect of echo cancellation, but also enhances the transmission of emotional information, which helps to improve the emotion recognition and response capabilities of the voice interaction system in human-computer dialogue.

[0031] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0032] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0033] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0034] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0035] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0036] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A deep learning-based echo cancellation and noise reduction method, characterized by: The steps include: The time-frequency features of the echo speech signal are extracted through time-frequency domain transformation, and the emotional state is identified by combining the emotion recognition network. The local features of the signal are extracted using a convolutional neural network, and the emotional fluctuations are analyzed through the LSTM model. The emotional label is then embedded in the echo cancellation network to generate a multimodal feature representation. A deep neural network is trained based on multimodal features to estimate the echo signal, and the echo removal signal is adjusted through an emotion perception network to preserve the emotional information of the speech. Based on the designed joint loss function, the system optimizes echo cancellation and speech emotion preservation, and evaluates the echo cancellation and noise reduction quality of the output speech signal. The output speech signal is evaluated for voice quality and emotional consistency, and different signals are generated. The quality of echo cancellation and noise reduction is determined based on the different signals.

2. The echo cancellation and noise reduction method based on deep learning according to claim 1, characterized in that: The time-frequency features of the echo speech signal are extracted based on the time-frequency domain transformation. The emotional state is identified by combining the emotion recognition network. The local features of the signal are extracted using a convolutional neural network. The emotional fluctuations are analyzed using the LSTM model. The emotional label is then embedded in the echo cancellation network to generate a multimodal feature representation. The specific steps are as follows: For the input speech signal with echo, the time-frequency domain is transformed, the time-frequency features of the speech signal are extracted using short-time Fourier transform, and the spectrum is obtained according to the short-time Fourier transform extraction; Use the LSTM-based emotion classification model as the emotion recognition network to identify the emotional state of the speech signal, combine the time-frequency features with the emotion markers in the speech, and output the emotion label for each time period; The emotion label of each time period is combined with the spectrum graph and transformed into the time-frequency domain. The resulting time-frequency graph is used as input and the local features of the time-frequency graph are extracted using a convolutional neural network. The convolution layer extracts different local features of the time-frequency signal through convolution operations, while retaining the key features of the speech signal through pooling operations.

3. The echo cancellation and noise reduction method based on deep learning according to claim 2, characterized in that: The LSTM model is used to analyze emotional fluctuations and embed the emotional labels into the echo cancellation network to generate multimodal feature representations. The specific steps include: An LSTM network is used to capture the temporal relationships in speech signals and model the emotional fluctuations in speech. The input of the emotion recognition network is the local features of the time-frequency graph extracted from the convolutional neural network. The emotional state in the speech is identified, and each time point corresponds to an emotional state category, resulting in a sequence of emotional state labels. Convert the emotion label into an emotion vector and embed it into the network input of the echo cancellation; The emotion embedding vector is combined with the time-frequency features extracted from the convolutional neural network for feature splicing and fusion. The emotion features and time-frequency features are spliced ​​along the time axis to generate a multimodal feature representation.

4. The echo cancellation and noise reduction method based on deep learning according to claim 3, characterized in that: Based on multimodal features, a deep neural network is trained to estimate the echo signal, and the echo removal signal is adjusted through the emotion perception network to retain the speech emotion information. The specific steps are as follows: A deep neural network is used as the echo estimation model, with the input being the fused features after multimodal feature extraction and emotion embedding, and the output being the time series of the echo part. Use the emotion perception network to adjust the speech signal after echo removal, and perform denoising and feature extraction preprocessing on the input echo estimation signal and emotion features; The emotion perception network processes the preprocessed echo signal to generate an echo-removed signal, combines the echo-removed signal and the emotion feature through weighted sum, and uses it to adjust the speech signal after echo removal to retain the speech emotion information.

5. The echo cancellation and noise reduction method based on deep learning according to claim 4, characterized in that: According to the designed joint loss function, the optimization of echo cancellation and speech emotion preservation includes the following steps: The joint loss function includes echo cancellation loss and emotion preservation loss; Echo cancellation loss minimizes the difference between the speech signal after echo removal and the ideal echo-free signal, obtaining the ideal echo-free speech signal , calculate the speech signal after removing the echo The distance between the ideal speech signal and the echo cancellation loss function is calculated using the minimum mean square error: ; The emotion preservation loss is to minimize the emotional difference between the speech after echo cancellation and the original speech. The degree of preservation of the emotional features is determined by the cosine similarity of the emotional features. The emotion preservation loss function is: ,in, Calculate the cosine similarity function; The joint loss function combines the echo cancellation loss with the emotion preservation loss in a weighted manner. The joint loss function is: ,in, and are the weight coefficients of echo cancellation loss and emotion preservation loss respectively.

6. The echo cancellation and noise reduction method based on deep learning according to claim 5, characterized in that: The echo cancellation and noise reduction quality evaluation of the output voice signal is performed, including the following steps: Evaluate from two aspects: speech quality evaluation and emotion consistency evaluation, and obtain the speech quality index generated by the speech quality evaluation process and the emotion consistency index generated by the emotion consistency evaluation process; The obtained speech quality index and emotion consistency index are combined to generate a noise reduction quality coefficient; The voice quality index and emotion consistency index are both proportional to the noise reduction quality coefficient.

7. The echo cancellation and noise reduction method based on deep learning according to claim 6, characterized in that: The output speech signal is analyzed for speech quality and emotional consistency, and different signals are generated. The echo cancellation and noise reduction quality is determined based on the different signals, including the following steps: Compare the generated noise reduction quality coefficient with the set echo noise reduction threshold; If the noise reduction quality coefficient is greater than or equal to the echo noise reduction threshold, an echo noise reduction stable signal is generated without additional noise reduction processing; If the noise reduction quality coefficient is less than the echo noise reduction threshold, an echo noise reduction fluctuation control signal is generated to trigger adaptive adjustment, provide real-time feedback on the current noise reduction status, and start the noise reduction adjustment mechanism.