Voice enhancement method based on AI echo cancellation
By building an echo cancellation model based on convolutional neural network and long and short-term memory network, combining adaptive filters and deep convolutional autoencoder, the problems of echo cancellation and speech enhancement in complex environments are solved, and efficient echo cancellation and speech quality improvement are achieved.
Patent Information
- Application Number
- CN202510747000.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing echo cancellation techniques are poorly robust in complex environments, making it difficult to effectively remove echoes and improve the clarity and intelligibility of speech signals, especially in dynamic environments and nonlinear echo scenarios, and traditional methods often ignore further enhancement of speech quality.
The voice data sets of different echo environments are obtained through the recording device, time-frequency conversion is performed and echo type labels are classified. The echo cancellation model is constructed by combining convolutional neural networks and long-term memory networks, signal estimation is generated using adaptive filters, and speech enhancement is performed through deep convolutional autoencoder, and filter parameters are dynamically adjusted to optimize the echo cancellation effect.
It realizes efficient and stable echo cancellation and voice quality improvement in complex echo environments, adapts to the needs of different echo environments, and improves the clarity and intelligibility of voice signals.
Smart Images

Figure CN120472920A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a speech enhancement method based on AI echo cancellation. Background Art
[0002] With the rapid development of information and communication technologies, voice interaction has become an indispensable part of people's daily lives, widely used in fields such as smart assistants, remote communications, and speech recognition. However, in practical applications, voice signals are often affected by echoes, especially in call environments, conference systems, and smart home environments. Echo problems often lead to reduced voice quality and affect the user experience. Echo cancellation technology, as a key voice processing technology, is widely used in fields such as voice communications and audio conferencing. Its purpose is to eliminate echoes caused by environmental noise, device reflections, or other external factors, thereby improving the clarity and intelligibility of voice signals.
[0003] While existing echo cancellation technologies have addressed the echo problem in speech signals to some extent, traditional echo cancellation methods often rely on filter-based adaptive algorithms. These methods often perform poorly in complex environments, particularly those with strong echoes. Existing technologies are less robust when dealing with dynamic environments, nonlinear echoes, and non-stationary noise, which can easily lead to speech distortion or residual echo. With the development of deep learning technology, convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) have been introduced into echo cancellation models, improving performance to some extent. However, these methods typically require large amounts of training data and may suffer from latency issues during real-time processing. Furthermore, existing echo cancellation technologies often neglect further processing to enhance speech quality after echo removal, resulting in suboptimal signal quality after echo removal, impacting speech clarity and intelligibility. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a speech enhancement method based on AI echo cancellation, which solves the problems of the above-mentioned background technology.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a speech enhancement method based on AI echo cancellation, comprising the following steps: S1. obtaining speech data sets in different echo environments through a recording device, performing time-frequency conversion on the speech signal of the speech data set and extracting the speech spectrum, defining and classifying the echo type label; S2. extracting the time-frequency features of the speech spectrum through a convolutional neural network, combining the long short-term memory network to capture the temporal dependency of the signal, and constructing an echo cancellation model; S3. extracting the target echo signal characteristics according to the echo type label and the echo cancellation model, combining the adaptive filter to generate a signal estimate after echo removal, measuring the signal quality after echo removal, and calculating the echo cancellation effect index by comprehensively analyzing the signal quality after echo removal and the echo residual characteristics; S4. dynamically adjusting the filter parameters according to the echo cancellation effect index through an adaptive algorithm, and performing speech enhancement on the signal after echo removal through a deep convolutional autoencoder.
[0006] Furthermore, the speech signals in the speech dataset are converted into time-frequency domains and the speech spectrum is extracted. The specific process of defining and classifying echo type labels is as follows: short-time Fourier transform is performed on each speech signal in the speech dataset to convert the time domain signal into a time-frequency domain signal; the spectral features of the speech signal, including the amplitude spectrum and phase spectrum, are extracted, and different echo type labels are defined according to the type and characteristics of the echo; the speech signals in the speech dataset are classified according to the echo characteristics, and a corresponding echo type label is assigned to each data sample to construct a labeled speech dataset.
[0007] Furthermore, the time-frequency features of the speech spectrum are extracted through a convolutional neural network, and the specific process of combining the long short-term memory network to capture the temporal dependency of the signal is as follows: the spectral features of the speech signal are input into the convolutional neural network, and the local time-frequency features are extracted through the convolution layer to capture the spatial information in different frequency ranges; the extracted time-frequency features are passed into the pooling layer to further reduce the feature dimension, and the convolution and pooling features are passed to the long short-term memory network to capture the temporal dependency of the signal and retain the temporal information and dynamic changes of the signal; through the output of the long short-term memory network, high-level feature representations containing time-frequency features and temporal information are extracted.
[0008] Furthermore, the specific process of constructing the echo cancellation model is as follows: the extracted time-frequency features and timing information are input into the echo cancellation model as the input data of the model; the input data is fused with the echo type label to form a composite input feature containing speech signal features and echo features; based on the fused features, local features are extracted through the convolutional layer, and the time-frequency features and timing information are integrated through the fully connected layer to construct echo cancellation models for different echo types.
[0009] Furthermore, the specific process of extracting the characteristics of the target echo signal based on the echo type label and the echo cancellation model is as follows: based on the echo type label, the echo characteristic category of the target signal is determined, specifically including the time delay, frequency response characteristics, and echo intensity of the echo; the time-frequency characteristics and timing information of the input signal are extracted through the convolutional layer and long short-term memory network in the echo cancellation model, and combined with the echo type label to form a multidimensional feature vector; based on the combination of the echo type label and the time-frequency characteristics, the characteristics of the target echo are identified, including the time delay, intensity, and frequency response of the echo; the extracted target echo characteristics can be used as model output.
[0010] Furthermore, the specific process of generating an echo-removed signal estimate in combination with an adaptive filter is as follows: based on the extracted echo characteristics and the input speech signal, the signal is processed by an adaptive filtering algorithm to generate a preliminary estimated signal with the echo removed; the filter coefficients are adjusted by the adaptive filter to minimize the difference between the input signal and the output signal to remove the echo signal; the filter output is evaluated, and the generated echo-removed signal is used as the signal estimate after denoising.
[0011] Furthermore, by comprehensively analyzing the signal quality and echo residual characteristics after echo removal, the specific process of calculating the echo cancellation effect index is as follows: calculating the signal quality after echo removal, including the signal-to-noise ratio and speech distortion, to obtain the speech quality evaluation index; analyzing the echo residual characteristics, measuring the amplitude, delay, and spectral distribution of the residual echo, and comparing them with the initial echo characteristics; and calculating the echo cancellation effect index based on the evaluation results of the signal quality and echo residual characteristics.
[0012] Furthermore, the specific process of dynamically adjusting the filter parameters through an adaptive algorithm according to the echo cancellation effect index is as follows: according to the echo cancellation effect index, the adjustment direction and amplitude of the current filter parameters are determined; according to the difference between the signal quality and the echo residual, the target signal quality index is set; the error is calculated and the filter coefficient is adjusted through the adaptive filtering algorithm, and the learning rate and filter step size are adjusted according to the size of the error; the filter coefficient is dynamically updated, and the echo cancellation effect is optimized through multiple iterations until the expected echo cancellation effect index threshold is reached.
[0013] Furthermore, the specific process of further performing speech enhancement on the signal after echo removal through a deep convolutional autoencoder is as follows: the signal after echo removal is input into the deep convolutional autoencoder network as the input signal; in the encoder part of the autoencoder, the spatial features of the signal are extracted through the convolution layer to further compress the representation dimension of the signal; in the decoder part, the spatial information of the signal is gradually restored through the convolution layer, and the denoised signal is generated through the deconvolution layer; the network parameters are optimized through the reconstruction error of the autoencoder, and the convolution kernel and filter coefficients are adjusted during the training process to enhance the denoised speech quality, and the enhanced speech signal is output as the final speech enhancement result.
[0014] The present invention has the following beneficial effects: (1) This speech enhancement method based on AI echo cancellation can classify different types of echoes by collecting multiple speech data in an echo environment, performing time-frequency conversion on the speech signals, and extracting speech spectra. It can also provide accurate echo type labels for the subsequent echo cancellation model. By extracting the time-frequency features of the speech signal through a convolutional neural network and combining it with a long short-term memory network to capture the temporal dependencies of the signal, an echo cancellation model can be efficiently constructed, effectively providing an accurate signal processing solution for echo removal.
[0015] (2) This speech enhancement method based on AI echo cancellation generates an echo-removed signal estimate by combining an adaptive filter. It then calculates the echo cancellation effect index by comprehensively analyzing the signal quality and echo residual characteristics after echo removal. Based on this, it dynamically adjusts the filter parameters and optimizes the echo cancellation effect. At the same time, a deep convolutional autoencoder is used to perform speech enhancement on the echo-removed signal, further improving the clarity and quality of the speech signal, effectively improving the intelligibility of the speech signal, and adapting to the complex needs of different echo environments.
[0016] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a speech enhancement method based on AI echo cancellation according to the present invention. DETAILED DESCRIPTION
[0018] The embodiments of the present application solve the problem of echo interference in complex echo environments through a speech enhancement method based on AI echo cancellation. By combining the extraction and modeling of time-frequency features and timing information, this method can accurately identify and eliminate echoes, improve the quality and clarity of speech signals, and reduce the impact of residual echo on speech understanding. At the same time, through the combination of adaptive algorithms and deep convolutional autoencoders, the signal enhancement effect is further optimized, achieving efficient and stable speech echo cancellation and speech quality improvement, meeting the needs of speech enhancement in different echo environments.
[0019] The overall idea of the solution in the embodiments of this application is as follows: A speech dataset in different echo environments is obtained through recording equipment. The speech signal of the speech dataset is converted into time-frequency and the speech spectrum is extracted. The echo type label is defined and classified.
[0020] The time-frequency features of the speech spectrum are extracted through convolutional neural networks, and the temporal dependencies of the signals are captured by long short-term memory networks to construct an echo cancellation model.
[0021] Based on the echo type label and echo cancellation model, the target echo signal characteristics are extracted, and an adaptive filter is used to generate a signal estimate after echo removal. The signal quality after echo removal is measured, and the echo cancellation effect index is calculated by comprehensively analyzing the signal quality after echo removal and the echo residual characteristics.
[0022] According to the echo cancellation effect index, the filter parameters are dynamically adjusted through an adaptive algorithm, and the speech enhancement is performed on the signal after echo removal through a deep convolutional autoencoder.
[0023] See also Figure 1 , an embodiment of the present invention provides a technical solution: a speech enhancement method based on AI echo cancellation, comprising the following steps: S1. obtaining a speech data set in different echo environments through a recording device, performing time-frequency conversion on the speech signal of the speech data set and extracting the speech spectrum, defining and classifying the echo type label; S2. extracting the time-frequency features of the speech spectrum through a convolutional neural network, combining it with a long short-term memory network to capture the temporal dependency of the signal, and constructing an echo cancellation model; S3. extracting the target echo signal characteristics according to the echo type label and the echo cancellation model, combining it with an adaptive filter to generate a signal estimate after echo removal, measuring the signal quality after echo removal, and calculating the echo cancellation effect index by comprehensively analyzing the signal quality after echo removal and the echo residual characteristics; S4. dynamically adjusting the filter parameters according to the echo cancellation effect index through an adaptive algorithm, and performing speech enhancement on the signal after echo removal through a deep convolutional autoencoder.
[0024] In this implementation, S1. Recording equipment: This is used to capture voice signals in different echo environments. Voice data from different environments can be used to simulate echo phenomena that may be encountered in real-world scenarios. Time-frequency conversion: This is a signal processing technique that uses the Fourier transform to convert voice signals from the time domain to the frequency domain. The time domain represents how the signal changes over time, while the frequency domain displays the frequency components of the signal. A commonly used method is the short-time Fourier transform (STFT), which allows the frequency distribution of the signal to be analyzed within a specific time window. Speech spectrum: This is the spectral representation of the voice signal obtained after time-frequency conversion. The speech spectrum displays the distribution of the signal in time and frequency, which is very important for subsequent feature extraction. Echo type label: The characteristics of the echo signal are classified according to their different propagation paths, degree of reflection, and other factors. Defining echo type labels can assist in subsequent model training and echo cancellation. Echo types may include long echoes, short echoes, and echoes of different frequencies. S2. Convolutional neural network (CNN): A CNN is a deep learning model commonly used for feature extraction from images and signals. It uses convolutional layers to automatically extract important time-frequency features from the speech spectrum and identify patterns in the speech signal, such as phonemes and tones. Convolution operations effectively process local features of speech and extract spatial information from the signal. Long Short-Term Memory (LSTM): LSTM is a type of recurrent neural network (RNN) particularly well-suited for processing time series data. LSTM can capture long-term dependencies in the signal, helping the model understand the relationship between different time segments in the speech signal. LSTM is often used to address long-term dependencies that traditional RNNs cannot handle. Echo Cancellation Model: Combining the time-frequency features extracted by CNN with the temporal information captured by LSTM, a powerful echo cancellation model is constructed. This model automatically learns and removes echoes from input speech signals with echoes. S3. Target Echo Signal Feature Extraction: Based on the echo type label, the echo cancellation model is combined to extract echo signal features, such as echo delay, frequency range, and intensity. This facilitates further analysis of the echo characteristics contained in the signal. Adaptive Filter: An adaptive filter is a filter that automatically adjusts its parameters based on the input signal. It continuously adjusts the filter coefficients so that the signal output by the filter is closest to the target signal. In echo cancellation, the filter can be adjusted to remove the echo signal and reduce the residual echo. Signal quality measurement and residual echo analysis: Evaluate the signal quality after echo removal using certain indicators (such as signal-to-noise ratio and distortion). At the same time, analyze the residual echo characteristics and assess the impact of the echo that remains after echo removal. Echo Cancellation Effectiveness Index: Combines the analysis results of signal quality and residual echo to calculate an echo cancellation effectiveness index, which is a quantitative measure of echo cancellation effectiveness. A higher effectiveness index means better echo cancellation. S4. Adaptive Algorithm: An adaptive algorithm is an algorithm that can automatically adjust its parameters based on system feedback.In this step, an adaptive algorithm dynamically adjusts filter parameters (such as learning rate and step size) based on the calculated echo cancellation performance index to optimize the echo cancellation effect. Deep Convolutional Autoencoder (DCAE): A DCAE is a deep learning model that extracts features from the input signal through convolution operations and reconstructs an enhanced signal through a decoder. It effectively improves the clarity and naturalness of speech signals, reduces noise, and enhances speech audibility. In particular, DCAE can further enhance speech quality in echo-prone environments.
[0025] Specifically, the speech signals in the speech dataset are converted into time-frequency domains, the speech spectrum is extracted, and the specific process of defining and classifying echo type labels is as follows: short-time Fourier transform is performed on each speech signal in the speech dataset to convert the time domain signal into a time-frequency domain signal; the spectral features of the speech signal, including the amplitude spectrum and phase spectrum, are extracted, and different echo type labels are defined according to the type and characteristics of the echo; the speech signals in the speech dataset are classified according to the echo characteristics, and a corresponding echo type label is assigned to each data sample to construct a labeled speech dataset.
[0026] In this implementation, the Short-Time Fourier Transform (STFT) divides a long signal into multiple short time segments, then performs a Fourier transform on each segment to obtain frequency information within each segment. This transformation method effectively captures the signal's temporal and frequency variations and is particularly suitable for processing non-stationary signals (such as speech signals). The STFT converts a time-domain signal into a time-frequency domain signal, where each point represents a signal component at a specific time and frequency. The time-domain signal, which represents the signal's fluctuations over time, typically reflects the signal's original characteristics. The frequency-domain signal, which represents the signal's distribution over frequency, reveals its spectral characteristics, such as the intensity of different frequency components. Using the STFT, the speech signal is decomposed into a two-dimensional time-frequency plot, which simultaneously displays both the time and frequency domain characteristics of the signal, providing a foundation for subsequent analysis. The amplitude spectrum, which represents the intensity or amplitude of the signal at each frequency point, is obtained by calculating the amplitude of each frequency component and is typically represented as the "magnitude" portion of the spectrum. In speech signal processing, the amplitude spectrum reflects the primary frequency components of the speech signal, which are crucial for speech recognition and analysis. The phase spectrum, which describes the phase information of the signal at each frequency point, is also described. It reflects the temporal structure of the signal (e.g., waveform variations) and plays an important role in the reconstruction and enhancement of speech signals. In speech processing, while the amplitude spectrum is commonly used for feature extraction, the phase spectrum also plays a key role in echo cancellation and signal enhancement, as phase information helps restore the naturalness and clarity of speech signals. Echo type: Echo types are categorized based on factors such as the environmental characteristics of the environment in which they occur, the propagation path, and the intensity of the reflection. Different types of echoes cause varying degrees of signal interference and exhibit distinct spectral characteristics. Echo types include: Long echoes: Echoes of prolonged duration, typically caused by long signal reflection paths or complex environmental factors. Short echoes: Echoes of shorter duration, typically occurring in close-range or rapidly reflecting environments. High-frequency echoes: Echoes that are more pronounced in the high-frequency range, likely due to reflections from walls or other surfaces. Low-frequency echoes: Echoes with more pronounced low-frequency components, typically occurring in larger spaces or environments with strong low-frequency wave propagation. Echo type labels are defined to distinguish signals of different echo types for subsequent model training and echo cancellation. Echo characteristic classification: By analyzing the echo characteristics of each speech signal (such as echo delay, intensity, and frequency), speech signals in a speech dataset can be classified according to different echo types. This process uses STFT signal spectrum analysis and echo type definitions to determine the echo characteristics of each signal. Echo type label assignment: For each speech signal sample, a corresponding echo type label is assigned based on its echo characteristics. This ensures that each data sample has a clear echo characteristic classification, helping the model learn how to cancel different types of echoes.Constructing a labeled speech dataset: After echo feature analysis and label assignment, the entire speech dataset becomes a labeled dataset, with each speech signal labeled with its echo type. This labeled data is used to train the echo cancellation model, enabling it to automatically adjust and eliminate different echo types.
[0027] Specifically, the specific process of extracting the time-frequency features of the speech spectrum through a convolutional neural network and capturing the temporal dependency of the signal in combination with a long short-term memory network is as follows: the spectral features of the speech signal are input into the convolutional neural network, and the local time-frequency features are extracted through the convolution layer to capture the spatial information in different frequency ranges; the extracted time-frequency features are passed into the pooling layer to further reduce the feature dimension, and the convolutional and pooled features are passed to the long short-term memory network to capture the temporal dependency of the signal and retain the temporal information and dynamic changes of the signal; through the output of the long short-term memory network, high-level feature representations containing time-frequency features and temporal information are extracted.
[0028] In this implementation, the spectral features of speech signals can be obtained after processing speech signals using methods such as the short-time Fourier transform (STFT). Spectral features generally include the intensity distribution of the signal at different frequencies, typically expressed as an amplitude spectrum and a phase spectrum. Spectral features represent detailed information about the speech signal in the frequency domain and serve as the basis for subsequent processing and analysis. Convolutional layers: In convolutional neural networks, convolutional layers perform convolution operations on the input signal using different convolution kernels to extract important features of the signal in local regions (i.e., small frequency and time intervals). These features include the local frequency components of the speech signal, which can capture details within the frequency range of the signal. Local features: These refer to local information of the signal, reflecting changes in the signal at a specific time point or frequency band. For example, the changes in a specific frequency component over a period of time could represent a specific syllable or specific pitch in the signal. Spatial information within the frequency range: Convolutional layers not only capture the signal's time domain features but also identify the relationships between different frequency components, i.e., spatial information. This helps extract the different frequency components of the speech signal and their changes over different time periods. Pooling layer: The purpose of the pooling operation is to reduce the dimensionality of the feature map obtained after the convolution operation while retaining the most important feature information. Pooling layers typically use max pooling or average pooling, which reduces data complexity by selecting the maximum or average value within a small local region. Dimensionality reduction: Pooling layers not only reduce computational effort but also prevent feature overfitting and improve model generalization. By reducing the size of the feature map, pooling layers speed up subsequent computations and make the model more robust. Long short-term memory (LSTM): LSTM is a special type of recurrent neural network (RNN) that uses gating mechanisms (such as input, forget, and output gates) to capture long-term dependencies in time series data. Compared to standard RNNs, LSTM effectively addresses the vanishing gradient problem and is suitable for processing long time series data. Temporal dependencies: LSTMs, through their internal memory cells, can preserve temporal information in the input sequence. For example, the temporal dependencies between syllables, words, and sentences in speech signals are important for speech understanding and generation. LSTM can capture dynamic changes in signals by processing these temporal dependencies. Time series information: Speech signals are typical time series signals, and the meaning and context of each speech segment are closely related to the temporal relationships between them. Through LSTM processing, the model can preserve the signal's time series information and identify patterns in signal changes over time. Dynamic changes: Speech signals contain not only static spectral features but also dynamic features that change over time. LSTM can extract these dynamic features and process them as important components of the signal. This is crucial for enhancing the clarity and accuracy of speech signals.High-level feature representation: After a convolutional neural network extracts time-frequency features and an LSTM captures temporal information, the final model outputs a high-level feature representation that combines both time-frequency and temporal information. These features integrate all important information about the speech signal in both the time and frequency dimensions, providing a foundation for subsequent speech enhancement or echo cancellation tasks. Fusion of time-frequency and temporal information: The resulting high-level features not only preserve the spectral information of the speech signal but also incorporate its temporal characteristics. This feature representation is significantly effective in restoring signals contaminated by echo or noise, as well as in subsequent speech enhancement and clarity processing.
[0029] Specifically, the specific process of constructing the echo cancellation model is as follows: the extracted time-frequency features and timing information are input into the echo cancellation model as the input data of the model; the input data is fused with the echo type label to form a composite input feature containing speech signal features and echo features; based on the fused features, local features are extracted through the convolutional layer, and the time-frequency features and timing information are integrated through the fully connected layer to construct echo cancellation models for different echo types.
[0030] In this implementation, time-frequency features and time series information: In the previous step, we extracted the time-frequency features of the speech signal using a convolutional neural network (CNN) and captured the signal's temporal dependencies using a long short-term memory network (LSTM). These features represent the speech signal's frequency domain representation and its dynamic characteristics over time. Input data for the echo cancellation model: The extracted time-frequency features and time series information serve as input data for the echo cancellation model. This input data contains detailed spectral information of the speech signal and its dynamic changes over time, providing the necessary signal information for the echo cancellation model. Fusion features: To enable the echo cancellation model to process different echo types, the input time-frequency features, time series information, and echo type labels are combined to form a composite input feature. This fused feature incorporates the speech signal's time-frequency characteristics, temporal dependencies, and additional information about the echo type, providing more precise guidance for subsequent echo cancellation. Fully connected layer: The fully connected layer integrates local features with global information by weighting and concatenating the features output by the convolutional layer. The fully connected layer can integrate time-frequency features and timing information, allowing the model to more comprehensively understand the structure and patterns of the input data. Integrating time-frequency features and timing information: During the echo cancellation process, it is necessary to consider not only the frequency components of the signal but also the temporal change pattern of the signal. The fully connected layer can combine the time-frequency features extracted by the convolutional layer and the timing information extracted by the LSTM, allowing the model to simultaneously consider the frequency domain information and time domain dynamics of the signal, thereby better eliminating the echo. Echo cancellation model for echo type: Based on different echo types (such as the characteristic differences of echoes in different environments), the echo cancellation model will have the ability to adapt to different echo types. By integrating multi-dimensional features, the echo cancellation model can customize the processing of different echo features, thereby improving the effect of echo cancellation.
[0031] Specifically, the specific process of extracting the characteristics of the target echo signal based on the echo type label and the echo cancellation model is as follows: based on the echo type label, the echo characteristic category of the target signal is determined, including the echo time delay, frequency response characteristics, and echo intensity; the time-frequency characteristics and timing information of the input signal are extracted through the convolutional layer and long short-term memory network in the echo cancellation model, and combined with the echo type label to form a multidimensional feature vector; based on the combination of the echo type label and the time-frequency characteristics, the characteristics of the target echo are identified, including the echo time delay, intensity and frequency response; the extracted target echo characteristics can be used as model output.
[0032] In this embodiment, the echo characteristic category is determined based on the echo type label: In this step, the system first classifies different types of echoes by using the echo type label. Each echo type may have its own unique characteristics, such as echo delay, frequency response characteristics, and echo intensity. Delay: refers to the delay time of the echo signal compared to the original signal, and is usually used to represent the propagation delay of the sound wave from the source to the reflection and then to the receiver. Frequency response characteristics: The change of the echo signal at different frequencies relative to the original voice signal may appear as frequency attenuation or enhancement. Echo intensity: refers to the intensity of the echo, which is equivalent to the amplitude of the reflected signal. Time-frequency features and timing information are extracted through convolutional layers and long short-term memory networks: Convolutional layers are used to extract time-frequency features from the input voice signal. Convolutional neural networks use filters to extract local features of the signal, helping to identify signal patterns and characteristics within different frequency bands, which is particularly effective when processing time-frequency spectrograms. Long short-term memory networks (LSTMs) are used to capture the temporal dependencies of the signal, that is, how to predict the output at the current moment from the past signal state. The advantage of LSTM is its ability to process signals with long-term dependencies, which is particularly important for echo signals, as their generation often involves long time delays. Combining echo type labels with time-frequency features to form a multidimensional feature vector: The echo type label is combined with the time-frequency features extracted by the convolutional layer and LSTM to form a multidimensional feature vector. This combination enables the model to more accurately understand the echo type and associate it with the signal's time-frequency characteristics. The purpose of this step is to link the echo type (such as short echoes, long echoes, and echoes with different frequency responses) with specific signal characteristics for subsequent analysis. Identifying the characteristics of the target echo: Once these echo characteristics are extracted, they serve as the output of the echo cancellation model. These characteristics provide essential information for the subsequent echo cancellation process, enabling the system to more accurately remove echoes and enhance speech quality.
[0033] Specifically, the specific process of generating a signal estimate for echo removal in combination with an adaptive filter is as follows: based on the extracted echo characteristics and the input speech signal, the signal is processed by an adaptive filtering algorithm to generate a preliminary estimated signal for echo removal; the filter coefficients are adjusted by the adaptive filter to minimize the difference between the input signal and the output signal to remove the echo signal; the filter output is evaluated, and the generated echo-removed signal is used as the signal estimate after denoising.
[0034] In this embodiment, processing is performed based on echo characteristics and the input speech signal. First, based on the echo characteristics extracted in the previous step (such as delay, frequency response, and echo strength) and the original input speech signal, this information is used as input for the adaptive filtering algorithm. The echo characteristics provide the adaptive filter with essential clues for removing specific echo types. Using these characteristics, the system can identify the delay and other frequency-related characteristics of the echo, providing important reference data for filter design and adjustment. Adaptive filtering algorithm processing: An adaptive filter is a filter that automatically adjusts its filter coefficients based on input signal and environmental changes. The key to adaptive filtering is to remove echo by continuously adjusting the filter parameters to make the output signal as close to the target signal as possible. During the echo removal process, the adaptive filter first processes the input signal to generate a preliminary echo-removed estimate signal. This preliminary signal may not completely remove the echo, but it does remove most of the echo components. The filter output is evaluated to generate a denoised signal estimate. After adjusting the filter coefficients, the filter output signal is evaluated and verified. This echo-removed signal estimate is considered a denoised signal, which will be clearer and have significantly reduced echo components than the original signal. The resulting signal estimate is used as the basis for further optimization and speech enhancement.
[0035] Specifically, the echo cancellation effect index is calculated by comprehensively analyzing the signal quality and echo residual characteristics after echo removal. The specific process is as follows: the signal quality after echo removal, including the signal-to-noise ratio and speech distortion, is calculated to obtain the speech quality evaluation index; the echo residual characteristics are analyzed, the amplitude, delay, and spectral distribution of the residual echo are measured, and compared with the initial echo characteristics; the echo cancellation effect index is calculated based on the evaluation results of the signal quality and echo residual characteristics.
[0036] In this embodiment, the echo cancellation effect index formula is: ; Parameter description : Echo cancellation effect index, which measures the overall effect of the echo cancellation algorithm. A higher value indicates better echo cancellation effect. : The signal-to-noise ratio of the signal after echo removal, indicating the signal quality after echo removal. A higher signal-to-noise ratio indicates better signal quality after echo removal. : The signal-to-noise ratio of the signal before echo removal, used to compare with the signal-to-noise ratio after echo removal, reflecting the signal quality before echo removal. : Signal-to-noise ratio improvement ratio, which indicates the degree of improvement in signal quality before and after echo removal. A larger value indicates a more significant improvement in signal quality after echo removal. : Residual echo delay, which indicates the delay of the echo remaining in the system after echo removal. The smaller the delay, the better the echo cancellation effect. : Maximum allowed delay, indicating the upper limit of the echo residual delay that the system can tolerate, which is usually determined by the device and system response time. : Residual echo delay ratio, which measures the ratio of the residual echo delay to the maximum delay. A smaller ratio indicates a shorter residual echo delay and better cancellation. : Echo residual amplitude, which indicates the strength of the remaining echo after echo removal. The smaller the value, the better the echo cancellation effect. : Maximum allowed amplitude, indicating the maximum acceptable echo residual amplitude. : Echo residual amplitude ratio, which represents the ratio of the residual echo intensity to the maximum allowable residual echo intensity. The smaller the ratio, the better the echo cancellation effect. : The spectral integral of the echo residual signal, which represents the total energy of the echo residual signal in the frequency domain. It measures the degree of presence of the echo in the frequency domain. : The spectral integral of the original signal before echo removal, representing the frequency energy of the original signal. : Echo residual spectrum ratio, which indicates the proportion of the echo residual signal energy in the frequency domain. A smaller ratio indicates a cleaner spectrum after echo cancellation, indicating better cancellation effectiveness. : Weights determine the relative contributions of signal quality, residual echo amplitude, delay, and spectral distribution to the echo cancellation performance index. By adjusting these weights, you can flexibly optimize the echo cancellation performance evaluation criteria based on the requirements of the actual application scenario.
[0037] Specifically, the specific process of dynamically adjusting the filter parameters through an adaptive algorithm based on the echo cancellation effect index is as follows: according to the echo cancellation effect index, the adjustment direction and amplitude of the current filter parameters are determined; according to the difference between the signal quality and the echo residual, the target signal quality index is set; the error is calculated and the filter coefficient is adjusted through the adaptive filtering algorithm, and the learning rate and filter step size are adjusted according to the size of the error; the filter coefficient is dynamically updated, and the echo cancellation effect is optimized through multiple iterations until the expected echo cancellation effect index threshold is reached.
[0038] In this embodiment, according to the echo cancellation effect index First, determine whether the current filter cancellation effect meets the expectation. If not, calculate the direction and amplitude of adjustment to further optimize the echo cancellation effect. According to the difference between signal quality and echo residual, set the target signal quality index (SNR). or speech distortion ). This goal is used to measure whether the filter needs to be adjusted to ensure the best signal quality. Calculate the error and adjust the filter coefficients based on the error between the current signal quality and the target signal quality. , the filter coefficients are adjusted. The error reflects the gap between the current filtering effect and the expected effect: ;in: is the target signal quality. is the current signal quality. The larger the error, the more adjustments are needed for the filtering effect. According to the size of the error, adjust the learning rate ( ) and the filter step size ( ). The larger the error, the more appropriate the learning rate should be to correct the echo cancellation effect quickly. ;in: is the update amount of the filter coefficients. is the input signal (at the current moment). Is the learning rate, which is used to control the speed of filter update and usually needs to be adjusted dynamically. Dynamic update of filter coefficients: Dynamic update of filter coefficients through adaptive filtering algorithm to gradually optimize the echo cancellation effect. The update formula is: ;in: are the coefficients of the filter at the current moment. Is the updated filter coefficient. Multiple iterations to optimize the echo cancellation effect: This process updates the filter coefficients through multiple iterations to gradually optimize the echo cancellation effect. After each iteration, the signal quality is re-evaluated, the error is calculated, and the filter parameters are adjusted until the echo cancellation effect reaches the expected threshold. Once the echo cancellation effect index Reaching the target effect threshold (such as ), the filter update process is terminated, indicating that the echo cancellation effect has been optimized.
[0039] Specifically, the specific process of further performing speech enhancement on the signal after echo removal through a deep convolutional autoencoder is as follows: the signal after echo removal is input into the deep convolutional autoencoder network as the input signal; in the encoder part of the autoencoder, the spatial features of the signal are extracted through the convolution layer to further compress the representation dimension of the signal; in the decoder part, the spatial information of the signal is gradually restored through the convolution layer, and the denoised signal is generated through the deconvolution layer; the network parameters are optimized through the reconstruction error of the autoencoder, and the convolution kernel and filter coefficients are adjusted during the training process to enhance the denoised speech quality, and the enhanced speech signal is output as the final speech enhancement result.
[0040] In this implementation, input signal processing involves feeding the echo-removed signal (echo-cancelled signal) into a deep convolutional autoencoder network as the network input. The input signal can be a speech signal that has undergone echo removal, but may still contain noise or other distortion components and require further enhancement. Encoder: Extracting Spatial Features and Compressing Signal Representation: In the encoder portion of the autoencoder, the input signal is processed through multiple convolutional layers. Each convolutional layer helps extract the signal's spatial features (e.g., local patterns in frequency and time information), and pooling layers further compress the signal's representation dimensionality. Convolutional layers convolve the input signal with filters (convolution kernels) to extract local features, such as spectral characteristics and temporal patterns. In this way, the high-dimensional representation of the signal is gradually compressed to a lower dimension to capture important feature information. Decoder: Restore the Signal's Spatial Information and Denoise: In the decoder, convolutional layers gradually restore the signal's spatial information, reconstructing details of the input signal layer by layer. Convolutional layers gradually expand the low-dimensional feature representation into a higher-dimensional space to reconstruct signal details while preserving the signal's local characteristics. Next, the signal is upsampled through a deconvolution layer (also known as a transposed convolution layer or upsampling layer). The deconvolution layer expands the signal representation from a low-dimensional space back to the original spatial dimensions and restores the signal's detailed features. Deconvolution not only helps restore the spatial distribution of the signal but also removes noise and other distortions, enhancing speech clarity. Reconstruction Error Optimization: Network parameter updates: Throughout the training process, the reconstruction error is calculated by calculating the difference between the network output and the target signal (such as the denoised speech signal). The reconstruction error reflects the distortion introduced by the autoencoder network during signal reconstruction. The reconstruction error is used to optimize the network's convolution kernel and filter coefficients through backpropagation. This optimization process gradually improves the network's denoising performance and, by adjusting the parameters of the convolution kernel and deconvolution layers, gradually enhances the signal enhancement effect. Outputting an enhanced speech signal: During training, through backpropagation and parameter updates, the network ultimately outputs an enhanced speech signal. This signal is denoised, enhanced, and has more details restored. The final output, serving as the final speech enhancement result, exhibits higher quality and clarity.
[0041] In summary, this application has at least the following effects: A speech enhancement method based on AI-powered echo cancellation effectively improves echo cancellation accuracy and speech signal quality by combining convolutional neural networks, long short-term memory networks, and adaptive filters. By dynamically adjusting filter parameters and applying a deep convolutional autoencoder for speech enhancement, this method achieves adaptive optimization in varying echo environments, significantly improving speech clarity and naturalness. By comprehensively analyzing the signal's time-frequency characteristics, timing information, and residual echo properties, this method comprehensively evaluates and optimizes echo cancellation effectiveness, demonstrating strong adaptability and application value, particularly in the fields of speech enhancement and communications.
[0042] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0043] The present invention is described with reference to flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0044] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0045] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0046] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0047] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A speech enhancement method based on AI echo cancellation, characterized in that: The following steps are involved: S1. Obtain speech datasets in different echo environments through recording equipment, perform time-frequency conversion on the speech signals in the speech datasets, extract speech spectra, and define and classify echo type labels; S2. Use a convolutional neural network to extract the time-frequency features of the speech spectrum, combine it with a long short-term memory network to capture the temporal dependencies of the signal, and build an echo cancellation model. S3. Extract the target echo signal characteristics based on the echo type label and the echo cancellation model, generate an echo-removed signal estimate using an adaptive filter, measure the signal quality after echo removal, and calculate the echo cancellation effect index by comprehensively analyzing the signal quality and echo residual characteristics after echo removal. S4. Based on the echo cancellation effect index, the filter parameters are dynamically adjusted through an adaptive algorithm, and the speech enhancement is performed on the signal after echo removal through a deep convolutional autoencoder.
2. The speech enhancement method based on AI echo cancellation according to claim 1, characterized in that: The specific process of performing time-frequency conversion on the speech signal of the speech dataset, extracting the speech spectrum, and defining and classifying the echo type labels is as follows: Perform short-time Fourier transform on each speech signal in the speech dataset to convert the time domain signal into a time-frequency domain signal; Extract the spectral features of the speech signal, including the amplitude spectrum and phase spectrum, and define different echo type labels according to the type and characteristics of the echo; The speech signals in the speech dataset are classified according to the echo characteristics, and the corresponding echo type label is assigned to each data sample to construct the speech dataset.
3. The speech enhancement method based on AI echo cancellation according to claim 2, characterized in that: The specific process of extracting the time-frequency features of the speech spectrum through a convolutional neural network and combining it with a long short-term memory network to capture the temporal dependencies of the signal is as follows: The spectral features of the input speech signal are fed into the convolutional neural network, which extracts local time-frequency features through the convolutional layer to capture spatial information within different frequency ranges. The extracted time-frequency features are passed to the pooling layer to further reduce the feature dimension. The convolution and pooling features are then passed to the long short-term memory network to capture the temporal dependencies of the signals and preserve the timing information and dynamic changes of the signals. Through the output of the long short-term memory network, high-level feature representations containing time-frequency features and timing information are extracted.
4. The method for speech enhancement based on AI echo cancellation according to claim 3, characterized in that: The specific process of building an echo cancellation model is as follows: Input the extracted time-frequency features and timing information into the echo cancellation model as input data of the model; The input data is fused with the echo type label to form a composite input feature that includes speech signal features and echo features; Based on the fused features, local features are extracted through the convolutional layer, and the time-frequency features and timing information are integrated through the fully connected layer to construct echo cancellation models for different echo types.
5. The method for speech enhancement based on AI echo cancellation according to claim 4, characterized in that: The specific process of extracting the target echo signal characteristics based on the echo type label and echo cancellation model is as follows: Based on the echo type label, determine the echo characteristic category of the target signal, including the echo delay, frequency response characteristics, and echo intensity; The convolutional layer and long short-term memory network in the echo cancellation model extract the time-frequency features and timing information of the input signal, and combine them with the echo type label to form a multi-dimensional feature vector; Based on the combination of echo type label and time-frequency features, the characteristics of the target echo are identified, including the time delay, intensity and frequency response of the echo; The extracted target echo characteristics can be used as model output.
6. The method for speech enhancement based on AI echo cancellation according to claim 5, characterized in that: The specific process of generating a signal estimate with echo removal in combination with an adaptive filter is as follows: Based on the extracted echo characteristics and the input speech signal, the signal is processed by an adaptive filtering algorithm to generate a preliminary estimated signal with the echo removed; The filter coefficients are adjusted by an adaptive filter to minimize the difference between the input signal and the output signal and remove the echo signal; The filter output is evaluated and the resulting denoised signal is used as the denoised signal estimate.
7. The method for speech enhancement based on AI echo cancellation according to claim 6, characterized in that: The specific process of calculating the echo cancellation effect index by comprehensively analyzing the signal quality and echo residual characteristics after echo removal is as follows: Calculate the signal quality after echo removal, including signal-to-noise ratio and speech distortion, and obtain the speech quality assessment index; Analyze the residual echo characteristics, measure the amplitude, delay, and spectrum distribution of the residual echo, and compare them with the initial echo characteristics; The echo cancellation effect index is calculated by integrating the evaluation results of signal quality and echo residual characteristics.
8. The method for speech enhancement based on AI echo cancellation according to claim 7, characterized in that: The specific process of dynamically adjusting the filter parameters through the adaptive algorithm according to the echo cancellation effect index is as follows: According to the echo cancellation effect index, the adjustment direction and amplitude of the current filter parameters are determined; According to the difference between signal quality and echo residual, the target signal quality index is set; Calculate the error and adjust the filter coefficients through the adaptive filtering algorithm, and adjust the learning rate and filter step size according to the size of the error; The filter coefficients are dynamically updated to optimize the echo cancellation effect through multiple iterations until the expected echo cancellation effect index threshold is reached.
9. The method for speech enhancement based on AI echo cancellation according to claim 8, characterized in that: The specific process of further performing speech enhancement on the signal after echo removal through a deep convolutional autoencoder is as follows: The signal after removing the echo is input into the deep convolutional autoencoder network as the input signal; In the encoder part of the autoencoder, the spatial features of the signal are extracted through the convolutional layer to further compress the representation dimension of the signal; In the decoder part, the spatial information of the signal is gradually restored through the convolution layer, and the denoised signal is generated through the deconvolution layer; The network parameters are optimized through the reconstruction error of the autoencoder, and the convolution kernel and filter coefficients are adjusted during the training process to enhance the quality of the denoised speech and output the enhanced speech signal as the final speech enhancement result.
Citation Information
Cited By
Model training method for removing nonlinear echo and display equipment
CN121438811A