Voice signal enhancement method and system

By employing an improved least mean square algorithm and spectral subtraction, the problem of poor convergence efficiency and filtering effect of speech signal enhancement methods in different scenarios is solved, thereby improving the clarity and adaptability of speech signals and making them suitable for applications such as speech recognition, communication, and intelligent interaction.

CN121011201AActive Publication Date: 2025-11-25广东公信智能会议股份有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511444472.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-25
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing speech signal enhancement methods struggle to achieve ideal convergence efficiency and filtering effects in various scenarios, especially in complex environments such as conference rooms, where the speech signal collected by the microphone is subject to multipath interference such as wall reflections, resulting in noise. Furthermore, traditional methods are insufficient in denoising performance in low signal-to-noise ratio environments.

Method used

An improved least mean square algorithm is adopted, which dynamically adjusts the step size by combining adaptive initial step size, step size adjustment degree and difference degree, and optimizes the filtering process of speech signal by combining spectral subtraction processing, thereby enhancing the quality of speech signal.

Benefits of technology

It improves the clarity and quality of speech signals, enhances the adaptability and robustness of speech signal enhancement systems, and is applicable to fields such as speech recognition, voice communication, and intelligent voice interaction. In particular, it significantly reduces background noise interference in complex noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121011201A_ABST
    Figure CN121011201A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, in particular to a voice signal enhancement method and system, and the method comprises the steps: obtaining a real-time voice signal and a historical voice signal of a target region, and carrying out the filtering of the real-time voice signal of the target region through an improved least mean square algorithm, so as to enhance the voice quality; a concept of self-adaptive initial step length is introduced, and the step length is dynamically adjusted according to the maximum characteristic value of an autocorrelation matrix of a real-time signal and the step length adjustment degree so as to adapt to the change of a voice signal; the step length adjustment reflects the convergence trend of the voice signal filtering process; the difference degree is used for quantifying the difference between the real-time voice signal and the historical voice signal, the characteristics of the voice signal can be effectively evaluated, and the filtering process is dynamically adjusted. According to the invention, the problem that the existing algorithm is difficult to realize ideal convergence efficiency and filtering effect in different scenes is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing. More particularly, the present application relates to a speech signal enhancement method and system. BACKGROUND

[0002] With the wide application of speech communication and speech recognition technology, the quality of speech signals is increasingly important in various systems. However, in actual environments, speech signals are often affected by noise, echo, interference and other factors, resulting in a decrease in speech quality, which may seriously affect the normal operation of the system and user experience. Especially in low signal-to-noise ratio environments, noise and speech signals are mixed, and traditional speech enhancement methods face problems such as insufficient de-noising performance and speech distortion. Therefore, how to effectively extract clear speech information from noise-polluted speech signals has become a key problem to be solved in the field of speech processing.

[0003] At present, speech signal enhancement methods mainly include time domain and frequency domain methods. Time domain methods mainly process signal waveforms, but have certain limitations in noise suppression; frequency domain methods analyze signals in the frequency domain, which can improve the quality of speech to some extent. However, in complex environments, the performance of these methods is still restricted by various factors. Therefore, a new speech signal enhancement method is urgently needed to improve the intelligibility and intelligibility of speech through more accurate signal processing means, thereby improving the performance of speech communication and speech recognition systems.

[0004] As a core technology of adaptive filtering, the least mean square algorithm performs well in speech signal processing, and can effectively improve the quality of speech in various scenarios due to its dynamic adaptability, low computational complexity and hardware friendliness. However, in practical applications, especially in conference room recording scenarios, the speech signals collected by microphones are often disturbed by multi-path interference such as wall reflection, resulting in the existence of noise. When using the least mean square algorithm for speech enhancement, the step size is crucial. Too large step size may increase the steady-state error, thereby affecting the filtering effect; while too small step size will cause the algorithm to react slowly in a dynamic environment, and cannot adapt to the changes of speech signals in time, thereby slowing down the convergence speed. Making it difficult for existing algorithms to achieve ideal convergence efficiency and filtering effect in different scenarios. SUMMARY

[0005] To solve the problem of the existing algorithm in the background art that it is difficult to achieve ideal convergence efficiency and filtering effect in different scenarios, the present application provides solutions in the following aspects.

[0006] In a first aspect, the present invention provides a speech signal enhancement method, comprising: acquiring a real-time speech signal and a historical speech signal of a target region; filtering the real-time speech signal using an improved least mean square algorithm to enhance the real-time speech signal, and obtaining the enhanced real-time speech signal; wherein the improved least mean square algorithm includes an adaptive initial step size, the adaptive initial step size being positively correlated with the degree of difference between the real-time speech signal and the historical speech signal, the step size adjustment degree, and the initial step size; the step size adjustment degree characterizes the filtering convergence trend of the speech signal.

[0007] The above technical solution improves the step size of the least mean square algorithm by dynamically adjusting the step size according to the degree of difference and convergence trend of the speech signal, so that the filtering process can maintain stability and efficiency in different speech environments. This solves the problem that existing algorithms are difficult to achieve ideal convergence efficiency and filtering effect in different scenarios.

[0008] Furthermore, the adaptive initial step size for: In the formula, To indicate the degree of difference, For step size adjustment degree, The largest eigenvalue of the autocorrelation matrix of the real-time speech signal. This is the initial step size.

[0009] Furthermore, the degree of difference for, In the formula, The first in real-time voice signal Each signal amplitude, It is the average value of the amplitudes of all signals in the historical speech signal. The standard deviation of the amplitude of all signals in the historical speech signal is given. This represents the total number of signal amplitudes in the real-time speech signal. Z These are the zero-crossing counts of the historical voice signal and the real-time voice signal, respectively. This represents the total number of historical speech signals.

[0010] Furthermore, the step size adjustment degree , In the formula, , These are the first in the historical speech signal. The error signal value, the first One error signal value, This represents the maximum value of the error signal in the real-time speech signal. This represents the total number of error signal values ​​in the historical speech signal. These are the preset hyperparameters.

[0011] The above technical solution introduces a step size adjustment factor to adaptively optimize the step size of the least mean square algorithm, enabling the filtering process to more accurately adapt to the dynamic changes in the speech signal. The calculation method of the step size adjustment factor comprehensively considers the distribution of error signals in historical speech signals, weights the accumulated errors using a normalized approach, and incorporates a smoothing factor to ensure that the step size adjustment effectively reflects the changing trend of signal errors. This method, while maintaining filtering stability, improves the algorithm's adaptability to noisy environments, allowing the step size to maintain sufficient adjustment space when the signal changes drastically, while avoiding over-adjustment when the signal tends to stabilize, thereby improving the convergence speed and robustness of the filtering.

[0012] Furthermore, the step size adjustment degree , In the formula, , These are the first in the historical speech signal. The error signal value, the first One error signal value, This represents the maximum value of the error signal in the real-time speech signal. This represents the total number of error signal values ​​in the historical speech signal. These are the preset hyperparameters.

[0013] The above technical solution improves filtering performance by optimizing the calculation method of step size adjustment, enabling the least mean square algorithm to adapt more accurately to error changes in different speech environments. It utilizes the distribution characteristics of error signal values ​​in historical speech signals, introduces normalization processing, combines weighted adjustment of error accumulation, and uses a logarithmic function to smooth the step size changes, ensuring that the step size adjustment reflects error dynamics while avoiding instability caused by excessive numerical fluctuations.

[0014] Furthermore, the initial step size , , It represents the largest eigenvalue of the autocorrelation matrix of the real-time speech signal.

[0015] The above technical solution sets the initial step size to the reciprocal of the largest eigenvalue of the autocorrelation matrix of the real-time speech signal, enabling the least mean square algorithm to adaptively adjust the step size under different signal environments, thus ensuring the stability and efficiency of the filtering process. By utilizing the eigenvalue characteristics of the autocorrelation matrix, the step size can adapt to the statistical characteristics of the signal, achieving dynamic adjustment and avoiding instability caused by an excessively large step size or a decrease in convergence speed caused by an excessively small step size.

[0016] Further, the real-time voice signal and the historical voice signal of the conference room are acquired, specifically, real-time voice data and historical voice data are collected by arranging microphones in the conference room, and the collected real-time voice data and historical voice data are equally sampled and discretized to obtain the real-time voice signal and the historical voice signal.

[0017] Further, the error signal value is obtained by using a long short-term memory network to obtain a predicted voice signal of the historical voice signal, and a difference value between the historical voice signal and the predicted voice signal is taken as the error signal value.

[0018] Further, the enhanced real-time voice signal is subjected to spectral subtraction processing.

[0019] The above technical solution introduces spectral subtraction processing after enhancing the real-time voice signal, and the spectral subtraction further optimizes the signal-to-noise ratio of the signal, so that the voice feature is more prominent. Especially in a complex or low signal-to-noise ratio environment, the interference caused by background noise can be significantly reduced, so that voice recognition, voice interaction and other applications can obtain more accurate input data.

[0020] In a second aspect, the present application provides a voice signal enhancement system comprising a memory and a processor, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a voice signal enhancement method according to any one of the above aspects is implemented.

[0021] The present application has the following beneficial effects: The present application can dynamically optimize the filtering processing of the voice signal and improve the clarity and quality of the voice signal by combining the improved least mean square algorithm. By combining the adaptive initial step size, the step size adjustment degree and the difference degree, the present application can accurately filter according to the characteristics of the real-time voice signal and the historical voice signal, so that the voice signal can quickly converge and effectively suppress noise under different environmental noise conditions. Not only the quality of the voice signal is improved, but also the adaptability and robustness of the voice signal enhancement system are enhanced, which is widely applicable to the fields of voice recognition, voice communication and intelligent voice interaction. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a flowchart schematically showing a voice signal enhancement method according to an embodiment of the present application; Figure 2 is a structural block diagram schematically showing a voice signal enhancement system according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] A voice signal enhancement method embodiment.

[0024] As Figure 1As shown, a flowchart of a voice signal enhancement method according to an embodiment of the present application comprises the following steps: S1: obtaining real-time voice signals and historical voice signals in a target area.

[0025] In one embodiment, real-time voice signals and historical voice signals in a target area are obtained to support subsequent functions such as voice analysis, conference recording generation, or voice recognition. Specifically, taking a conference room as an example, the following steps are included: First, a plurality of high-sensitivity microphone arrays are arranged in the conference room to cover the effective pickup range of the entire conference room, ensuring that voice data of different positions and different speakers during the conference can be fully collected. The microphone array can capture real-time voice data and historical voice data stored in the past period of time to form a complete voice signal database.

[0026] Second, the collected voice data is subjected to equal sampling and discretization processing to obtain time-stable voice signals. The sampling frequency can be set to 10 Hz by default, thereby ensuring sufficient time resolution at a lower computational cost and ensuring that the key features of the voice signal are retained. Of course, the sampling frequency can be dynamically adjusted according to actual application requirements, for example, the sampling frequency can be appropriately increased in a high-precision voice analysis scenario, while the sampling frequency can be reduced in a resource-limited or low-time resolution requirement scenario to reduce data storage and computational overhead.

[0027] In order to improve the quality of the voice signal and reduce the interference of environmental noise, after obtaining the real-time voice signal, it will be subjected to high-pass filter processing first. The high-pass filter can effectively remove low-frequency noise such as background noise caused by air conditioners, projection equipment, power lines, etc., thereby improving the clarity and distinguishability of the voice signal. In addition, the high-pass filter processing can also reduce the influence of low-frequency resonance effect on the voice features, so that the subsequent voice recognition, speaker separation or voice enhancement algorithms can more accurately extract the voice content and speaker features.

[0028] S2: filtering the real-time voice signal using an improved least mean square algorithm to enhance the real-time voice signal.

[0029] In one embodiment, the improved least mean square algorithm includes an adaptive initial step size, which is for: , wherein is the difference degree, is the step size adjustment degree, is the maximum eigenvalue of the real-time voice signal autocorrelation matrix, wherein the calculation of the maximum eigenvalue of the real-time voice signal autocorrelation matrix is obtained by the prior art, and the present scheme will not be described in detail, is an initial step length; By introducing an improved least mean square algorithm with adaptive initial step length, the step length can be intelligently adjusted according to the dynamic changes of the speech signal, thereby improving the convergence speed and stability of the filtering. Considering the difference degree of the speech signal, the step length adjustment trend and the signal autocorrelation characteristics, the step length can be quickly adjusted to an appropriate range in the case of large changes, so as to speed up the adaptability of the filtering process, and the step length range is limited when the change is small, to prevent instability caused by excessive adjustment. Through this adaptive optimization, the algorithm can maintain efficient noise suppression ability in different noise environments, while preserving the key features of the speech signal to the greatest extent, so that the enhanced speech signal is clearer and more accurate, providing high-quality input data for speech recognition, intelligent speech processing and other applications.

[0030] It should be noted that the initial step length is adjusted according to the difference degree and the step length adjustment degree. The greater the difference degree, the greater the gap between the latest speech signal and the historical speech signal; the greater the step length adjustment degree, the poorer the filtering effect of the initial step length, and the greater step length is needed to speed up the steady-state convergence speed of the filtering.

[0031] The difference degree is, , wherein, is the first signal amplitude in the real-time speech signal, is the average value of all signal amplitudes in the historical speech signal, is the standard deviation of all signal amplitudes in the historical speech signal, is the total number of signal amplitudes in the real-time speech signal, , Z are the zero-crossing times of the historical speech signal and the real-time speech signal, respectively, is the total number of historical speech signals.

[0032] By introducing a difference degree calculation method based on statistical characteristics, the filtering process can more accurately reflect the change trend between the real-time speech signal and the historical speech signal. The dispersion degree of the signal amplitude, the overall mean deviation and the change of the zero-crossing times are considered to measure the fluctuation characteristics of the real-time speech signal compared with the historical speech signal. Through this calculation method, the algorithm can enhance the step length adjustment amplitude when the signal changes greatly, and slow down the step length when the signal is stable, to avoid signal distortion caused by excessive correction.

[0033] The step length adjustment degree , , wherein, , are the first and second signal amplitudes in the historical speech signal, respectively, ​the first error signal value in the historical speech signal, the second error signal value in the historical speech signal, the first error signal value in the historical speech signal, the second error signal value in the historical speech signal, the maximum value of the error signal values in the real-time speech signal, the total number of error signal values in the historical speech signal, a preset hyperparameter.

[0034] Through the optimization calculation mode of the step adjustment degree, the filtering algorithm can more accurately adapt to the error distribution characteristics of the historical speech signal, and dynamic adjustment of the step is realized to improve the convergence efficiency and robustness. The cumulative influence of the historical error signal values is comprehensively considered, and normalization processing is used, so that the step adjustment can flexibly reflect the error change trend, while avoiding the excessive influence of extreme errors on the adjustment process. By introducing a smoothing factor, it is ensured that the step adjustment can respond to error fluctuations and maintain overall stability, so that the best filtering effect can be achieved in different noise environments. This optimization strategy improves the noise resistance and adaptability of the algorithm, making the speech signal enhancement more accurate and effectively reducing noise interference.

[0035] the initial step size , , the maximum eigenvalue of the autocorrelation matrix of the real-time speech signal. By associating the initial step size with the maximum eigenvalue of the autocorrelation matrix of the real-time speech signal, the startup process of the filtering algorithm is optimized, so that the step size can more accurately match the characteristics of the signal at the initial stage of signal processing. This setting allows the step size to automatically adjust according to the autocorrelation of the speech signal, ensuring that the filtering algorithm can quickly adapt to the dynamic changes of the signal at the initial stage, avoiding the problems of instability or slow convergence caused by excessively large or small step sizes.

[0036] The error signal value is specifically obtained by using a long short-term memory network to obtain a predicted speech signal of the historical speech signal, and the difference between the historical speech signal and the predicted speech signal is taken as the error signal value.

[0037] In another embodiment, the step adjustment degree , , wherein , the first error signal value in the historical speech signal, the second error signal value in the historical speech signal, the first error signal value in the historical speech signal, the second error signal value in the historical speech signal, the maximum value of the error signal values in the real-time speech signal, the total number of error signal values in the historical speech signal, the total number of error signal values in the historical speech signal, a preset hyperparameter.

[0038] By optimizing the step size adjustment, combined with the error signal characteristics and dynamic changes of historical speech signals, the step size in the filtering process can be more accurately adjusted to adapt to the volatility of the signal and the noise characteristics. This method considers the cumulative effect of error values and makes the step size adjustment more smooth and adaptive through weighted calculation of historical errors, thereby improving the stability and convergence efficiency of the filtering algorithm. The introduction of logarithmic smoothing processing of error values further enhances the adaptability to different types of signals, avoids excessive adjustment caused by excessive step size, and ensures that moderate adjustment can be maintained when the error is small.

[0039] S3: Obtain the enhanced real-time speech signal.

[0040] In one embodiment, after obtaining the enhanced real-time speech signal, further spectral subtraction processing is performed to further improve the quality of the speech signal and reduce noise interference. As a common noise suppression technique, spectral subtraction processing mainly removes the estimated noise component from the frequency spectrum of the speech signal to effectively suppress the influence of background noise. This processing step can significantly improve the intelligibility and audibility of the speech signal, especially in noisy environments. Specifically, in the enhanced real-time speech signal, the estimated value of the noise component is extracted by analyzing its spectral characteristics, and is subtracted from the signal spectrum to effectively suppress noise.

[0041] Through spectral subtraction processing, residual noise in the speech signal can be further eliminated, especially in high-noise environments such as conference rooms, streets, or industrial areas where noise is strong. This improves the purity of the speech signal. This not only helps speech signal recognition and processing, but also improves the accuracy and response speed of the system in intelligent speech recognition, voice interaction and other applications. The application of spectral subtraction processing can also avoid speech distortion or recognition errors caused by excessive noise to some extent, ensuring that the enhanced speech signal is clearer and more natural, and improving user experience.

[0042] This invention, by combining an improved least mean square algorithm, effectively improves the quality of real-time audio signals in conference rooms. First, filtering enhances the real-time audio signal, and an adaptive step-size adjustment method ensures stable and efficient convergence of the filtering process, thereby optimizing the clarity and audibility of the audio signal. Furthermore, the calculation of the degree of difference adaptively adjusts the step size based on the dynamic characteristics of the audio signal, avoiding the limitations of static parameters and making the algorithm more adaptable and flexible. In addition, spectral subtraction further reduces background noise interference, ensuring the purity of the audio signal, especially in noisy environments, enhancing the accuracy and stability of speech recognition. Overall, this invention significantly improves the quality of audio signals and is suitable for various real-time audio processing scenarios such as conferences and speech recognition, maintaining high-quality audio output, especially in complex noisy environments.

[0043] An embodiment of a speech signal enhancement system: like Figure 2 As shown, a block diagram of a speech signal enhancement system according to an embodiment of the present invention includes a processor and a memory.

[0044] This invention also provides a speech signal enhancement system. For example... Figure 2 As shown, the system includes a processor and a memory, the memory storing computer program instructions, which, when executed by the processor, implement the speech signal enhancement method according to the present invention.

[0045] The aforementioned voice signal enhancement system also includes other components well known to those skilled in the art, such as communication interfaces. Their settings and functions are known in the art and will not be described in detail here.

[0046] In this description, the terms "communication" and "communicate" are used broadly. For example, a device can communicate information to another device, even though the information need not be received explicitly by the other device. In other words, one device can communicate information to another device by placing the information in a location where the other device is able to retrieve the information, even though one device does not know exactly where or when another device will retrieve the information. The term "communication" can include one or both of these actions, and also can include other actions associated with these actions. For example, the process of placing information in a location where another device is able to retrieve the information can include the actions of encoding the information on a physical medium, transmitting encoded information on a physical medium, or other actions associated with these actions. Similarly, the process of retrieving information can include the actions of receiving the information on a physical medium, decoding encoded information on a physical medium, or other actions associated with these actions. Further, one device can communicate information to another device by causing another device to communicate the information. These provisions are merely examples of what is meant to be a "communication" or "communicate," and these provisions are not intended to limit the scope of the present application to these examples. Additionally, the present application contemplates that devices can communicate information in the form of signals, messages, data, or other information.

[0047] In the description of the specification, the meaning of "a plurality of" or "several" is at least two, for example, two, three, or more, unless specifically defined otherwise.

[0048] While the present application has been illustrated and described in detail in the drawings and foregoing description, such illustration and description is to be considered illustrative or exemplary and not restrictive; the present application is not limited to the disclosed embodiments. Various modifications and changes can be made thereto without departing from the spirit and scope of the present application as set forth in the claims below. It is to be understood that the replications of any of the elements of the disclosed embodiments can be made in the practice of the present application.

Claims

1. A method for enhancing speech signals, characterized in that, include: Acquire real-time and historical voice signals of the target area; The real-time speech signal is filtered using an improved least mean square algorithm to enhance it, and the enhanced real-time speech signal is obtained. Among them, the improved least mean square algorithm includes an adaptive initial step size, which is positively correlated with the degree of difference between the real-time speech signal and the historical speech signal, the degree of step size adjustment, and the initial step size. The step size adjustment degree characterizes the filtering convergence trend of the speech signal.

2. The speech signal enhancement method according to claim 1, characterized in that, The adaptive initial step size for: In the formula, To indicate the degree of difference, For step size adjustment degree, The largest eigenvalue of the autocorrelation matrix of the real-time speech signal. This is the initial step size.

3. The speech signal enhancement method according to claim 1, characterized in that, The degree of difference for, In the formula, The first in real-time voice signal Each signal amplitude, It is the average value of the amplitudes of all signals in the historical speech signal. The standard deviation of the amplitude of all signals in the historical speech signal is given. This represents the total number of signal amplitudes in the real-time speech signal. Z These are the zero-crossing counts of the historical voice signal and the real-time voice signal, respectively. This represents the total number of historical speech signals.

4. The speech signal enhancement method according to claim 1, characterized in that, The step size adjustment , In the formula, , These are the first in the historical speech signal. The error signal value, the first One error signal value, This represents the maximum value of the error signal in the real-time speech signal. This represents the total number of error signal values ​​in the historical speech signal. These are the preset hyperparameters.

5. The speech signal enhancement method according to claim 1, characterized in that, The step size adjustment , In the formula, , These are the first in the historical speech signal. The error signal value, the first One error signal value, This represents the maximum value of the error signal in the real-time speech signal. This represents the total number of error signal values ​​in the historical speech signal. These are the preset hyperparameters.

6. The speech signal enhancement method according to claim 1, characterized in that, The initial step size , , It represents the largest eigenvalue of the autocorrelation matrix of the real-time speech signal.

7. The speech signal enhancement method according to claim 1, characterized in that, The acquisition of real-time and historical voice signals from the conference room specifically involves: Real-time and historical voice data are collected by deploying microphones in the conference room, and the collected real-time and historical voice data are discretized by equal sampling to obtain real-time and historical voice signals.

8. The speech signal enhancement method according to claim 1, characterized in that, The error signal value is specifically defined as follows: the predicted speech signal obtained from the historical speech signal using a long short-term memory network, and the difference between the historical speech signal and the predicted speech signal is used as the error signal value.

9. A speech signal enhancement method according to claim 1, characterized in that, It also includes performing spectral subtraction processing on the enhanced real-time voice signal.

10. A speech signal enhancement system, characterized in that, It includes a memory and a processor, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a speech signal enhancement method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Wavelet transform and variable-step least mean square algorithm-based voice denoising method

    CN101894561A

  • Optical fiber monitoring speech enhancement technology based on backward Rayleigh scattering

    CN108696312A

  • Sound processing method and system based on variable step size LMS algorithm

    CN112116914A

  • Voice noise reduction method and device, and storage medium

    CN117153181A

  • Multiple input multiple output (MIMO) audio signal processing for speech de-reverberation

    US20180182411A1