Frequency domain dual-channel speech enhancement method of hearing aid
By employing a frequency-domain dual-channel speech enhancement method, utilizing signal preprocessing, feature extraction, and inter-channel correlation analysis, combined with short-time Fourier transform and deep learning techniques, the problem of speech quality degradation in noisy environments caused by traditional hearing aids is solved, achieving personalized and efficient noise suppression and speech enhancement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZUODIAN IND (HUBEI) CO LTD
- Filing Date
- 2024-10-21
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional hearing aids use a single-channel processing method, which cannot effectively process sounds of different frequencies separately, resulting in impaired speech quality. Furthermore, they have weak noise suppression capabilities in noisy environments, making it difficult to meet individual differences in needs.
A frequency-domain dual-channel speech enhancement method is adopted, which combines signal acquisition, frequency domain transformation, frequency analysis, signal processing and output control with signal preprocessing, feature extraction and inter-channel correlation analysis. Short-time Fourier transform, spectral subtraction and deep learning methods are used for noise suppression and speech enhancement.
It effectively identifies and suppresses background noise, maintains the naturalness and clarity of speech, responds quickly to environmental changes, provides stable enhancement effects, and meets individual needs.
Abstract
Description
Technical Field
[0001] This invention relates to the field of hearing aid technology, specifically a frequency domain dual-channel speech enhancement method for hearing aids. Background Technology
[0002] Hearing aids are electronic devices that amplify sound signals to help people with hearing loss receive and process sound better. Currently, hearing aids are widely used in my country and have become an important product in the field of hearing assistance. There are many types of hearing aids, including analog hearing aids, digital hearing aids, and wireless hearing aids. These hearing aids have their own characteristics in terms of performance, functions, and comfort, meeting the needs of different people with hearing loss.
[0003] Although hearing aids have achieved remarkable results in the field of hearing assistance, traditional hearing aid speech enhancement methods still have certain limitations. Traditional hearing aids usually use a single-channel processing method, which cannot effectively process sounds of different frequencies separately, resulting in impaired speech quality. In noisy environments, traditional hearing aids have weak noise suppression capabilities, which can easily lead to speech distortion. Since the degree of hearing loss and frequency characteristics of each person are different, traditional hearing aids cannot meet the needs of individual differences. Therefore, it is necessary to propose a frequency domain dual-channel speech enhancement method for hearing aids. Summary of the Invention
[0004] To address the problems in the prior art, this invention provides a frequency domain dual-channel speech enhancement method for hearing aids.
[0005] The technical solution adopted by this invention to solve its technical problem is: a frequency domain dual-channel speech enhancement method for hearing aids, comprising: signal acquisition: acquiring speech signals from the left and right ears through the microphone of the hearing aid; frequency domain conversion: converting the acquired speech signals into a frequency domain representation, typically achieved using Fast Fourier Transform; frequency analysis: analyzing the frequency components of the speech signal in the frequency domain to identify speech and noise; signal processing: adjusting the frequency components of the speech signal according to the user's hearing loss characteristics, including amplification, compression, and filtering; signal synthesis: resynthesizing the processed frequency components into a time domain signal; output control: outputting the processed signal through the speaker of the hearing aid, while adjusting the output parameters according to user feedback to achieve the best hearing effect.
[0006] It also includes signal preprocessing, feature extraction, and inter-channel correlation analysis; The signal preprocessing is the first step in the speech enhancement process, and its purpose is to improve signal quality and lay a solid foundation for subsequent feature extraction and enhancement processing. The signal preprocessing steps mainly include: Denoising: Removing background noise from speech signals using filters to improve the signal-to-noise ratio; Normalization: Normalize the signal to ensure that the signal strength is within a reasonable range, which facilitates subsequent processing; Framing: Dividing a continuous speech signal into short time frames, each containing a certain length of data, facilitates frequency domain analysis; Windowing: Apply a window function to each frame of signal to reduce spectral leakage and improve the accuracy of spectral analysis.
[0007] Feature extraction is a core step in speech enhancement methods. Its purpose is to extract information useful for speech enhancement from the preprocessed signal. The feature extraction steps include: Frequency domain transformation: The preprocessed time-domain signal is converted into a frequency-domain signal using the Fast Fourier Transform. Spectrum analysis: Analyzing the spectral characteristics of a frequency domain signal, including the amplitude, phase, and energy of the spectrum; Feature calculation: A series of feature parameters, such as spectral center, spectral entropy, and frequency distribution, are calculated based on spectral characteristics. These features will be used in subsequent speech enhancement processing.
[0008] The inter-channel correlation analysis is a unique step in dual-channel speech enhancement, aiming to improve speech enhancement by utilizing the relationship between the left and right ear signals. The channel correlation analysis step includes: Correlation calculation: Calculate the correlation between the left and right ear signals in the frequency domain, which can be achieved through methods such as correlation coefficient or mutual information; Correlation mapping: Based on the correlation calculation results, a correlation mapping between channels is constructed to guide subsequent signal fusion and enhancement; Adaptive adjustment: Based on the results of inter-channel correlation analysis, the dual-channel enhancement strategy is adaptively adjusted to optimize speech output.
[0009] In frequency domain dual-channel speech enhancement methods, feature extraction and processing are key steps that directly affect the speech enhancement effect. These steps include the application of short-time Fourier transform, noise suppression based on spectral subtraction, and feature extraction based on deep learning. Short-time Fourier transform (STFT) is a commonly used time-frequency analysis technique that can decompose a speech signal into a series of short-time spectral frames, thereby capturing the time-frequency characteristics of the speech signal. STFT obtains the spectrum that changes over time by dividing the signal into frames and applying Fourier transform to each frame. This method allows us to observe the changes in the spectrum of the speech signal over time and is very suitable for analyzing non-stationary signals. In speech enhancement, STFT is used to extract the time-frequency features of the signal, which can be used to identify speech and noise, thereby guiding subsequent enhancement processing.
[0010] The noise suppression based on spectral subtraction is a classic noise suppression technique that recovers clean speech by subtracting the estimated noise spectrum from the spectrum of noisy speech. Spectral subtraction assumes that noise and speech are independent in the frequency domain, so the speech spectrum can be recovered by subtracting the noise spectrum. Estimate the spectrum of the noise; subtract the noise spectrum from the spectrum of the noisy speech to obtain the enhanced speech spectrum; convert the spectrum back to the time domain using inverse Fourier transform.
[0011] The deep learning-based feature extraction has been widely used in the field of speech enhancement. Deep learning models: Deep learning models, such as convolutional neural networks, recurrent neural networks, or long short-term memory networks, are used to extract complex features from speech signals; Training process: The model is trained with a large amount of noisy and clean speech data, enabling it to learn how to extract useful speech features from noisy signals; deep learning-based feature extraction methods can capture more complex signal features and improve the performance of speech enhancement, especially when dealing with non-stationary noise and complex speech scenarios.
[0012] A method for using a frequency-domain dual-channel speech enhancement method for hearing aids includes the following steps: Step 1: Before using the new speech enhancement method, the following preparations need to be made: Ensure that you have suitable hardware devices, such as microphones and speakers that support frequency domain processing, and processors with sufficient computing power; install the necessary software, including implementation libraries of speech enhancement algorithms and deep learning frameworks; and collect speech datasets for training and testing, including noisy speech and clean speech samples. The second step is to collect noisy speech signals through a microphone, preprocess the collected speech signals, including denoising, normalization, framing and windowing, and extract the frequency domain features of the speech signals using short-time Fourier transform or other frequency domain analysis techniques. Step 3: Using a frequency domain filtering-based enhancement algorithm, design and apply a frequency domain filter to reduce noise; using a spectral mapping-based enhancement algorithm, calculate the spectral mapping relationship between the noisy speech and the reference speech, and adjust the spectrum of the noisy speech; using a deep learning-based enhancement algorithm, input the extracted speech features into a trained deep learning model to obtain the enhanced speech features. Step 4: Convert the enhanced speech features back to the time domain through inverse transform to obtain the enhanced speech signal; perform post-processing on the enhanced speech signal, such as smoothing filtering, to further improve the speech quality.
[0013] The beneficial effects of this invention are: The present invention discloses a frequency domain dual-channel speech enhancement method for hearing aids. The novel method can more effectively identify and suppress background noise, especially in non-stationary noise environments. In noisy environments, the novel method can better maintain the naturalness and clarity of speech, reduce noise-introduced artifacts, and adaptively adjust the enhancement strategy according to changes in noise, providing a more stable enhancement effect. The present invention discloses a frequency domain dual-channel speech enhancement method for hearing aids. This novel method can quickly respond to environmental changes and adjust the enhancement strategy in a timely manner. The deep learning algorithm can utilize the parallel processing capabilities of modern hardware, such as GPU acceleration, to achieve fast computation and real-time enhancement. Through model compression and optimization, the novel method can reduce the amount of computation and meet the needs of real-time applications. Detailed Implementation
[0014] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0015] The present invention discloses a frequency-domain dual-channel speech enhancement method for hearing aids, comprising: signal acquisition: acquiring speech signals from the left and right ears through the microphone of the hearing aid; frequency domain conversion: converting the acquired speech signals into a frequency domain representation, typically achieved using Fast Fourier Transform; frequency analysis: analyzing the frequency components of the speech signals in the frequency domain to identify speech and noise; signal processing: adjusting the frequency components of the speech signals according to the user's hearing loss characteristics, including amplification, compression, and filtering; signal synthesis: resynthesizing the processed frequency components into a time-domain signal; and output control: outputting the processed signal through the hearing aid's speaker, while simultaneously adjusting the output parameters based on user feedback to achieve optimal hearing results.
[0016] It also includes signal preprocessing, feature extraction, and inter-channel correlation analysis; the signal preprocessing is the first step in the speech enhancement process, and its purpose is to improve signal quality and lay a solid foundation for subsequent feature extraction and enhancement processing; the signal preprocessing steps mainly include: Denoising: Removing background noise from speech signals using filters to improve the signal-to-noise ratio; Normalization: Normalize the signal to ensure that the signal strength is within a reasonable range, which facilitates subsequent processing; Framing: Dividing a continuous speech signal into short time frames, each containing a certain length of data, facilitates frequency domain analysis; Windowing: Apply a window function to each frame of signal to reduce spectral leakage and improve the accuracy of spectral analysis.
[0017] Feature extraction is a core step in speech enhancement methods. Its purpose is to extract information useful for speech enhancement from the preprocessed signal. The feature extraction steps include: Frequency domain transformation: The preprocessed time-domain signal is converted into a frequency-domain signal using the Fast Fourier Transform. Spectrum analysis: Analyzing the spectral characteristics of a frequency domain signal, including the amplitude, phase, and energy of the spectrum; Feature calculation: A series of feature parameters, such as spectral center, spectral entropy, and frequency distribution, are calculated based on spectral characteristics. These features will be used in subsequent speech enhancement processing.
[0018] The inter-channel correlation analysis is a unique step in dual-channel speech enhancement, aiming to improve speech enhancement by utilizing the relationship between the left and right ear signals. The channel correlation analysis step includes: Correlation calculation: Calculate the correlation between the left and right ear signals in the frequency domain, which can be achieved through methods such as correlation coefficient or mutual information; Correlation mapping: Based on the correlation calculation results, a correlation mapping between channels is constructed to guide subsequent signal fusion and enhancement; Adaptive adjustment: Based on the results of inter-channel correlation analysis, the dual-channel enhancement strategy is adaptively adjusted to optimize speech output.
[0019] Methods for calculating inter-channel correlation Calculating inter-channel correlation is the process of analyzing the similarity between two signals. Here are some commonly used calculation methods: Pearson correlation coefficient: This is a commonly used correlation measure that measures the linear correlation between two signals at the same point in time. Spearman rank correlation coefficient: When a signal does not follow a normal distribution or has outliers, the Spearman rank correlation coefficient can be used to measure the nonlinear correlation of the signal. Mutual information: Mutual information is a method to measure the degree of information sharing between two signals. It considers not only linear relationships but also nonlinear relationships. Coherence coefficient: The coherence coefficient is used to measure the correlation between two signals in the frequency domain. It takes into account the correlation between the signals at different frequencies.
[0020] Relationship between correlation coefficient and speech enhancement Correlation coefficients play an important role in speech enhancement; the relationships between them are as follows: Correlation measurement: By calculating the correlation coefficient between the left and right ear signals, we can quantify the similarity between two signals, which is crucial for identifying and utilizing the binaural effect; Signal fusion: In dual-channel speech enhancement, the correlation coefficient can be used to guide the signal fusion strategy. For example, when reconstructing the signal, channels with high correlation can be given higher weights. Noise suppression: The correlation coefficient can also be used to identify and suppress noise, because the correlation between the left and right ear signals is usually reduced in noisy environments.
[0021] Optimization strategy for correlation coefficient To improve the efficiency of correlation coefficient utilization, the following optimization strategies are proposed: Time alignment: Time alignment technology ensures that the signals from the left and right ears are aligned in time, thereby improving the accuracy of correlation coefficient calculation; Frequency selection: In the frequency domain, based on the correlation between different frequency components, the frequency components with high correlation can be selectively enhanced to improve speech quality; Adaptive adjustment: Based on the correlation coefficient calculated in real time, the enhancement strategy is adaptively adjusted to adapt to the constantly changing speech and noise environment; Post-processing: Add post-processing steps, such as smoothing, to the enhanced signal to reduce artifacts caused by correlation calculation errors.
[0022] Through these computational methods and optimization strategies, inter-channel correlation analysis can provide effective information for dual-channel speech enhancement, thereby improving speech quality while also enhancing spatial and stereoscopic effects. In frequency domain dual-channel speech enhancement methods, feature extraction and processing are crucial steps that directly affect the enhancement effect. These steps involve the application of Short-Time Fourier Transform (STFT), noise suppression based on spectral subtraction, and feature extraction based on deep learning. STFT is a commonly used time-frequency analysis technique that decomposes the speech signal into a series of short-time spectral frames, thereby capturing the time-frequency characteristics of the speech signal. STFT obtains the time-varying spectrum by framing the signal and applying Fourier transform to each frame. This method allows us to observe the changes in the speech signal's spectrum over time and is well-suited for analyzing non-stationary signals. In speech enhancement, STFT is used to extract the time-frequency features of the signal. These features can be used to identify speech and noise, thus guiding subsequent enhancement processing.
[0023] The noise suppression based on spectral subtraction is a classic noise suppression technique that recovers clean speech by subtracting the estimated noise spectrum from the spectrum of noisy speech. Spectral subtraction assumes that noise and speech are independent in the frequency domain, so the speech spectrum can be recovered by subtracting the noise spectrum. The process involves estimating the noise spectrum, subtracting the noise spectrum from the spectrum of noisy speech to obtain the enhanced speech spectrum, and then converting the spectrum back to the time domain using an inverse Fourier transform.
[0024] The deep learning-based feature extraction has been widely used in the field of speech enhancement; deep learning models: using deep learning models such as convolutional neural networks, recurrent neural networks or long short-term memory networks to extract complex features of speech signals; Training process: The model is trained with a large amount of noisy and clean speech data, enabling it to learn how to extract useful speech features from noisy signals; deep learning-based feature extraction methods can capture more complex signal features and improve the performance of speech enhancement, especially when dealing with non-stationary noise and complex speech scenarios.
[0025] Speech enhancement algorithm implementation In the field of speech enhancement, algorithm implementation is a key step in technology deployment. The following is a detailed discussion of the implementation of three common speech enhancement algorithms. Enhancement Algorithm Based on Frequency Domain Filtering Frequency domain filtering is a traditional speech enhancement method that reduces noise and enhances speech by manipulating the signal in the frequency domain. Frequency domain filtering-based enhancement algorithms first convert the speech signal to the frequency domain through short-time Fourier transform, and then apply filters in the frequency domain to reduce noise. Steps: Perform STFT on the input noisy speech to obtain its spectral representation; Design and apply frequency domain filters, such as noise threshold filters and Wiener filters, to reduce noise components; The enhanced speech signal is obtained by performing inverse STFT on the filtered spectrum. This method is simple to implement, has relatively low computational complexity, and is suitable for real-time applications.
[0026] Spectral mapping-based enhancement algorithms Spectrum mapping-based enhancement algorithms improve speech quality by manipulating the spectrum of the speech signal in the frequency domain; spectrum mapping algorithms reduce the impact of noise by mapping the spectrum of noisy speech onto the spectrum of clean speech. Steps: Perform STFT on both the noisy speech and the reference speech; Calculate the mapping relationship between the noisy speech spectrum and the reference speech spectrum, such as proportional mapping and minimum mean square error mapping; Apply mapping relationships to adjust the spectrum of noisy speech; The enhanced speech is obtained by performing inverse STFT on the adjusted spectrum. The spectrum mapping algorithm can maintain the naturalness of the speech well while reducing noise.
[0027] Deep learning-based augmentation algorithms With the development of deep learning technology, deep learning-based speech enhancement algorithms are receiving increasing attention; deep learning-based enhancement algorithms use deep neural networks to learn the mapping relationship between noisy speech and clean speech; Steps: Build a deep learning model, such as CNN, RNN, or LSTM; The model is trained using a large amount of noisy and clean speech data to learn the mapping relationship; Noisy speech is input into a trained model to obtain enhanced speech; Deep learning algorithms can handle complex nonlinear relationships and provide high-quality speech enhancement, but they usually require significant computing resources and training time.
[0028] In practical application, the following preparations are necessary before using the new speech enhancement method: Ensure you have suitable hardware, such as microphones and speakers that support frequency domain processing, and a processor with sufficient computing power; install necessary software, including implementation libraries for speech enhancement algorithms and deep learning frameworks; collect speech datasets for training and testing, including noisy and clean speech samples; acquire noisy speech signals through the microphone, preprocess the acquired speech signals including denoising, normalization, framing, and windowing, and extract frequency domain features of the speech signals using short-time Fourier transform or other frequency domain analysis techniques; design and apply frequency domain filters to reduce noise using a frequency domain filtering-based enhancement algorithm; calculate the spectral mapping relationship between noisy speech and reference speech using a spectral mapping-based enhancement algorithm, and adjust the spectrum of the noisy speech; input the extracted speech features into a trained deep learning model using a deep learning-based enhancement algorithm to obtain enhanced speech features; convert the enhanced speech features back to the time domain using an inverse transform to obtain the enhanced speech signal; perform post-processing on the enhanced speech signal, such as smoothing filtering, to further improve speech quality.
[0029] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of protection claimed by the present invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A frequency-domain dual-channel speech enhancement method for hearing aids, characterized in that, include: S1. Signal Acquisition: Acquire speech signals from both ears through the microphone of the hearing aid; S2. Frequency domain conversion: Converting the acquired speech signal into a frequency domain representation, usually achieved using the Fast Fourier Transform; S3. Frequency Analysis: Analyze the frequency components of a speech signal in the frequency domain to identify speech and noise; S4. Signal Processing: Adjusting the frequency components of the speech signal based on the user's hearing loss characteristics, including amplification, compression, and filtering. S5. Signal Synthesis: Resynthesize the processed frequency components into a time-domain signal; S6 Output Control: The processed signal is output through the hearing aid's speaker, and the output parameters are adjusted according to the user's feedback to achieve the best hearing effect.
2. The frequency domain dual-channel speech enhancement method for a hearing aid according to claim 1, characterized in that: It also includes signal preprocessing, feature extraction, and inter-channel correlation analysis; The signal preprocessing is the first step in the speech enhancement process, and its purpose is to improve signal quality and lay a solid foundation for subsequent feature extraction and enhancement processing. The signal preprocessing steps mainly include: Denoising: Removing background noise from speech signals using filters to improve the signal-to-noise ratio; Normalization: Normalize the signal to ensure that the signal strength is within a reasonable range, which facilitates subsequent processing; Framing: Dividing a continuous speech signal into short time frames, each containing a certain length of data, facilitates frequency domain analysis; Windowing: Apply a window function to each frame of signal to reduce spectral leakage and improve the accuracy of spectral analysis.
3. The frequency domain dual-channel speech enhancement method for a hearing aid according to claim 2, characterized in that: Feature extraction is a core step in speech enhancement methods. Its purpose is to extract information useful for speech enhancement from the preprocessed signal. The feature extraction steps include: Frequency domain transformation: The preprocessed time-domain signal is converted into a frequency-domain signal using the Fast Fourier Transform. Spectrum analysis: Analyzing the spectral characteristics of a frequency domain signal, including the amplitude, phase, and energy of the spectrum; Feature calculation: A series of feature parameters, such as spectral center, spectral entropy, and frequency distribution, are calculated based on spectral characteristics. These features will be used in subsequent speech enhancement processing.
4. The frequency domain dual-channel speech enhancement method for a hearing aid according to claim 2, characterized in that: The inter-channel correlation analysis is a unique step in dual-channel speech enhancement, aiming to improve speech enhancement by utilizing the relationship between the left and right ear signals. The channel correlation analysis step includes: Correlation calculation: Calculate the correlation between the left and right ear signals in the frequency domain, which can be achieved through methods such as correlation coefficient or mutual information; Correlation mapping: Based on the correlation calculation results, a correlation mapping between channels is constructed to guide subsequent signal fusion and enhancement; Adaptive adjustment: Based on the results of inter-channel correlation analysis, the dual-channel enhancement strategy is adaptively adjusted to optimize speech output.
5. The frequency domain dual-channel speech enhancement method for a hearing aid according to claim 1, characterized in that: In frequency domain dual-channel speech enhancement methods, feature extraction and processing are key steps that directly affect the speech enhancement effect. These steps include the application of short-time Fourier transform, noise suppression based on spectral subtraction, and feature extraction based on deep learning. Short-time Fourier transform (STFT) is a commonly used time-frequency analysis technique that can decompose a speech signal into a series of short-time spectral frames, thereby capturing the time-frequency characteristics of the speech signal. STFT obtains the spectrum that changes over time by dividing the signal into frames and applying Fourier transform to each frame. This method allows us to observe the changes in the spectrum of the speech signal over time and is very suitable for analyzing non-stationary signals. In speech enhancement, STFT is used to extract the time-frequency features of the signal, which can be used to identify speech and noise, thereby guiding subsequent enhancement processing.
6. The frequency domain dual-channel speech enhancement method for a hearing aid according to claim 5, characterized in that: The noise suppression based on spectral subtraction is a classic noise suppression technique that recovers clean speech by subtracting the estimated noise spectrum from the spectrum of noisy speech. Spectral subtraction assumes that noise and speech are independent in the frequency domain, so the speech spectrum can be recovered by subtracting the noise spectrum. Estimate the spectrum of the noise; subtract the noise spectrum from the spectrum of the noisy speech to obtain the enhanced speech spectrum; convert the spectrum back to the time domain using inverse Fourier transform.
7. The frequency domain dual-channel speech enhancement method for a hearing aid according to claim 5, characterized in that: The deep learning-based feature extraction has been widely used in the field of speech enhancement. Deep learning models: Deep learning models, such as convolutional neural networks, recurrent neural networks, or long short-term memory networks, are used to extract complex features from speech signals; Training process: The model is trained with a large amount of noisy and clean speech data, enabling the model to learn how to extract useful speech features from noisy signals; Deep learning-based feature extraction methods can capture more complex signal features and improve speech enhancement performance, especially when dealing with non-stationary noise and complex speech scenarios.
8. A method of using a smart chip fixture with multi-directional adjustment function as described in any one of claims 1-7, characterized in that: Includes the following steps: Step 1: Before using the new speech enhancement method, the following preparations need to be made: Ensure that you have suitable hardware devices, such as microphones and speakers that support frequency domain processing, and processors with sufficient computing power; install the necessary software, including implementation libraries of speech enhancement algorithms and deep learning frameworks; and collect speech datasets for training and testing, including noisy speech and clean speech samples. The second step is to collect noisy speech signals through a microphone, preprocess the collected speech signals, including denoising, normalization, framing and windowing, and extract the frequency domain features of the speech signals using short-time Fourier transform or other frequency domain analysis techniques. Step 3: Use a frequency domain filtering-based enhancement algorithm to design and apply a frequency domain filter to reduce noise; An enhancement algorithm based on spectral mapping is used to calculate the spectral mapping relationship between noisy speech and reference speech, and the spectrum of the noisy speech is adjusted. An enhancement algorithm based on deep learning is used to input the extracted speech features into a trained deep learning model to obtain the enhanced speech features. Step 4: Convert the enhanced speech features back to the time domain using an inverse transform to obtain the enhanced speech signal; Post-processing techniques, such as smoothing filtering, are applied to the enhanced speech signal to further improve speech quality.