A method for estimating the signal-to-noise ratio of transmitted speech
By performing frame-based and Fourier analysis on noisy voice signals, combined with principal component analysis, separating speech and noise, the problem of signal-to-noise ratio estimation in noise environments is solved, and the performance and noise removal effect of speech enhancement algorithms are improved.
Patent Information
- Application Number
- CN202310329327.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-03-29
AI Technical Summary
The prior art has affected speech recognition and communication quality in noisy environments, and lacks fast and effective signal-to-noise ratio estimation methods, resulting in poor performance of speech enhancement algorithms.
By performing frame-based and window-based preprocessing of the noisy voice signal, short-time Fourier analysis is performed, initial segmentation threshold point α is set, speech and noise signals are separated, pure speech structure is extracted using principal component analysis, and signal-to-noise ratio is calculated.
It realizes accurate estimation of the signal-to-noise ratio without pure speech prior knowledge, improving the performance and noise removal effect of the speech enhancement algorithm, and the results are consistent with the human ear experience.
Smart Images

Figure CN116486836B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech signal processing, and more particularly to a method for estimating a signal-to-noise ratio of transmitted speech. Background Art
[0002] In nature, sound is a crucial medium for humans to communicate with each other and a crucial component of the human auditory perception system. Humans possess the ability to selectively block out surrounding background noise in complex environments while simultaneously capturing information from multiple sound sources. With technological advancements, the application of technologies such as voice communication and speech recognition is becoming increasingly widespread. While many smart devices can effectively identify speech content and speaker information, in real life, speech is easily eroded in noisy environments, and noise interference can affect recognition accuracy and communication quality. Therefore, speech enhancement technology has become a hot research topic for researchers both domestically and internationally.
[0003] Speech enhancement is a preprocessing technique that processes contaminated, noisy signals to produce clearer speech signals. Its primary goal is to improve speech clarity and intelligibility, while minimizing distortion and maximizing the signal-to-noise ratio (SNR) to reduce noise. Classic speech enhancement algorithms include spectral subtraction, subspace-based algorithms, and Kalman filtering. These algorithms all require a method to accurately calculate the SNR to evaluate their performance.
[0004] Currently, speech quality assessment algorithms for speech enhancement are primarily categorized into two main categories: subjective and objective. Subjective methods, such as MOS, CMOS, and ABX tests, primarily rely on human scoring of speech, requiring significant human and material resources. Therefore, there is an urgent need for objective evaluation methods that can conveniently and quickly assess speech quality, such as signal-to-noise ratio (SNR), log-likelihood ratio (LR), log-spectral distance (LSD), and short-term objective intelligibility. Objective evaluation methods offer the advantages of saving both time and effort. SNR has long been a classic method for measuring the performance of speech enhancement algorithms, providing an intuitive understanding of the strength of speech and signal. However, its calculation relies on prior knowledge of pure speech signals. In real life, only contaminated, noisy signals are available. The SNR estimate of this noisy signal largely determines the performance of speech enhancement algorithms and the effectiveness of noise removal. Summary of the Invention
[0005] Based on the limitations of existing evaluation methods that require knowledge of the original speech signal, the present invention provides a signal-to-noise ratio estimation method that does not require prior knowledge of a pure speech signal. It can more accurately separate the speech signal and the noise signal from the power spectrum of the noisy signal, and estimate the signal-to-noise ratio of the noisy signal based on the separated speech and noise signals. The estimated signal-to-noise ratio of the present invention has good monotonicity and applicability, and the evaluation result is consistent with the subjective auditory perception of the human ear.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A method for estimating the signal-to-noise ratio of transmitted speech, comprising the following steps:
[0008] (1) Preprocessing the speech signal, performing frame division and windowing preprocessing on the noisy speech signal to obtain a short-time stationary speech signal x(t);
[0009] (2) Performing short-time Fourier analysis related to the time sequence on the preprocessed speech signal x(t) to obtain a three-dimensional dynamic spectrogram X(k) representing the change of the speech signal spectrum over time;
[0010] (3) Setting an initial segmentation threshold point α for separating the speech signal and the noise signal;
[0011] (4) Respectively performing approximate region segmentation on the spectrogram X(k) of the noisy speech signal in the time direction and the frequency direction according to the initial segmentation threshold point α, and dividing the spectrogram X(k) of the noisy speech signal into a global speech part and an edge noise part;
[0012] (5) Performing secondary signal-noise separation on the spectrogram of the global speech part to obtain a local voiceprint part and a local noise part respectively;
[0013] (6) Calculating the power of the local voiceprint part, and taking this power as the total power P of the estimated clean speech signal signal ;
[0014] (7) Calculating the power of the edge noise part and the power of the local noise part respectively, and accumulating the two to obtain an estimated value P of the total power of the noise part noise ;
[0015] (8) Calculating the signal-to-noise ratio according to P signal and P noise as:
[0016]
[0017] Further, the initial segmentation threshold point α in the step (3) is determined by the sampling frequency of the noisy speech signal and the number of points of the Fourier transform, and the calculation steps are as follows:
[0018] s1. Determining the maximum value in the frequency direction of the spectrogram from the sampling frequency f s of the noisy speech signal Then the estimated range of the noise frequency f is: f ∈ [ξ × f max , f max , where ξ is the noise ratio, which is determined by the sampling frequency of the noisy speech and the power spectrum;
[0019] s2. Calculate the threshold point α for the first separation of the low-frequency speech part and the high-frequency noise part according to the estimated noise frequency range above:
[0020]
[0021] where N is the number of sampling points for short-time Fourier analysis.
[0022] Furthermore, in step (4), the starting value and the ending value in the time direction are determined by endpoint detection, and the initial segmentation threshold point α is used for judgment in the frequency direction: the part less than the initial segmentation threshold point α is judged as the global speech, and the part higher than the initial segmentation threshold point α is the edge noise part, thereby determining the effective region of the speech signal in the frequency direction.
[0023] Furthermore, in step (5), the principal component analysis method is used for the secondary signal noise separation to extract the main speech structure, i.e., the voiceprint information, of the global speech part, calculate the proportion of the local voiceprint part in the global speech according to the extracted principal components, and the remaining part is the local noise.
[0024] Furthermore, in step (6), the total power P of the clean speech signal signal = β × P audio , P audio is the power of the local voiceprint part, and β ∈ [0, +∞] is an adjustable parameter for weighting the power of the local voiceprint part.
[0025] Beneficial effects:
[0026] The present invention first performs the first separation of the signal noise to obtain the global speech and the edge noise respectively; on this basis, the principal component analysis is performed on the global speech part for the secondary signal noise separation to extract the part containing the clean speech structure; removing the clean speech part to obtain the spectrum sequence of only the local noise signal; thereby obtaining that the power of the noise signal is the sum of the cumulative powers of the edge noise and the local noise, and then calculating the signal-to-noise ratio for the clean speech part and the noise signal part, improving the accuracy of the signal-to-noise ratio calculation, and the whole calculation process is simple and easy to implement.
[0027] In the present invention, the signal-to-noise ratio is estimated for the power spectrum of the noisy signal, and the result can more intuitively reflect the noise strength of the power spectrum, which is convenient for using the signal-to-noise ratio to evaluate the effectiveness of various speech noise reduction and speech separation algorithms, as well as to measure problems such as the sound quality and the channel transmission quality in communication. Description of the Drawings
[0028] Figure 1 is the overall structural block diagram of the method algorithm of the present invention;
[0029] Figure 2It is a flowchart of the primary signal-noise separation module;
[0030] Figure 3 It is the spectrogram of the clean speech signal;
[0031] Figure 4 It is the spectrogram of the noisy speech signal;
[0032] Figure 5 It is the schematic diagram of the primary signal-noise separation;
[0033] Figure 6 They are the separated global speech and edge noise parts. Detailed implementation manners
[0034] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0035] As Figure 1 shown, a method for estimating the signal-to-noise ratio of transmitted speech proposed by the present invention includes the following steps.
[0036] Taking Figure 3 the clean speech shown as the test object to demonstrate the specific process of the algorithm of the present invention. In order to obtain the signal contaminated by noise, additive noise is superimposed on Figure 3 it for mixing to obtain the spectrogram of the noisy signal, as Figure 4 shown.
[0037] Since the speech signal is a non-stationary signal, some preprocessing operations need to be performed before analyzing the noisy speech signal, including framing and windowing the speech signal to obtain a short-time stationary speech signal.
[0038] The processed signal is subjected to short-time Fourier analysis to obtain the spectrogram of the signal. After the operation of the primary signal-noise separation module, the result of the primary separation of the speech signal and the noise signal is obtained, that is, two parts: the global speech and the edge noise.
[0039] The global speech part is input into the secondary signal-noise separation module, and the final clean speech signal and the local noise part are obtained through principal component analysis.
[0040] The edge noise and the local noise are accumulated to obtain the noise signal and the clean speech signal, and they are calculated according to the signal-to-noise ratio calculation formula:
[0041]
[0042] where P signal and P noise are the power of the clean speech signal and the power of the noise signal respectively.
[0043] The specific implementation process of the primary signal-noise separation module is asFigure 2 As shown, the effective estimation range of the noise is determined by the sampling frequency f of the speech signal s and the noise ratio ξ.
[0044] Specifically, in this embodiment, the effective estimation range of the noise is:
[0045] f ∈ [ξ × f max , f max .
[0046] Among them, the value of the noise ratio ξ is:
[0047] And, in this embodiment, f max is the maximum value in the frequency direction of the spectrogram, and its value is half of the sampling frequency:
[0048] Calculate the primary separation threshold point of the speech part and the noise part according to the estimated effective range of the noise:
[0049]
[0050] Among them, N is the number of sampling points for short-time Fourier analysis, and the sampling frequency of the speech signal and the number of points for Fourier transform determine the threshold point for signal noise separation.
[0051] The criterion for the primary separation of the signal and the noise is to judge the part greater than the threshold point as the edge noise part and the part less than the threshold point as the global speech, thereby determining the effective region in the frequency direction of the global speech signal.
[0052] Specifically, in this embodiment, the initial value and the final value in the time direction of the global speech signal are determined by endpoint detection based on short-time energy and zero-crossing rate.
[0053] The process of regionally segmenting the spectrogram in combination with the effective frequency range of the speech signal is as Figure 5 shown. The spectrogram will be segmented from the frequency direction, and the separated edge noise part and the global speech part are as Figure 6 shown. In the first sub-graph, there is basically no energy of the low-frequency speech part in the edge noise part, and the global speech part is as Figure 6 shown in the second sub-graph, which contains the overall low-frequency structure of the speech.
[0054] For the secondary signal noise separation, the principal component analysis method is used to extract the main speech structure, i.e., the voiceprint information, from the global speech part.
[0055] Specifically, the number of principal components in the principal component analysis is set to 8.
[0056] In the secondary signal noise separation method, the proportion of the local voiceprint part in the global speech is calculated based on the extracted principal components, and the remaining part is the local noise.
[0057] The total power calculation formula of the clean speech signal is:
[0058] P signal = β × P audio
[0059] where P audio is the power of the local voiceprint part, and β ∈ [0, +∞] is an adjustable parameter for weighting the power of the local voiceprint part.
[0060] The specific embodiments of the present invention described above do not constitute a limitation on the protection scope of the present invention. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. A method for estimating the signal-to-noise ratio of transmitted speech, characterized in that, The steps are as follows: (1) Preprocess the speech signal. Frame and window the noisy speech signal to obtain a short-time stationary speech signal x(t). (2) Perform short-time Fourier analysis related to time sequence on the preprocessed speech signal x(t) to obtain a three-dimensional dynamic spectrogram X(k) representing the variation of the speech signal spectrum over time. (3) Set an initial segmentation threshold point α for separating the speech signal and the noise signal. (4) According to the initial segmentation threshold point α, approximately segment the spectrogram X(k) of the noisy speech signal in the time direction and the frequency direction, and segment the spectrogram X(k) of the noisy speech signal into a global speech part and an edge noise part. (5) Perform secondary signal-noise separation on the global speech part spectrogram to obtain a local voiceprint part and a local noise part respectively. (6) Calculate the power of the local voiceprint part, and use this power as the total power P of the estimated clean speech signal signal ; (7) Calculate the power of the edge noise part and the power of the local noise part respectively, and accumulate the two to obtain the estimated value P of the total power of the noise part noise ; (8) According to P signal and P noise Calculate the signal-to-noise ratio as:
2. The SNR estimation method for transmitted speech according to claim 1, characterized in that The calculation steps of the initial segmentation threshold point α in step (3) are as follows: s1. Determine the maximum value in the frequency direction of the spectrogram from the sampling frequency f of the noisy speech signal s Then the estimated range of the noise frequency f is: f ∈ [ξ × f max , f max , where ξ is the noise ratio; s2. According to the estimated noise frequency range mentioned above, calculate the threshold point α for the first separation of the low-frequency speech part and the high-frequency noise part: where N is the number of sampling points for short-time Fourier analysis.
3. The method for estimating the signal-to-noise ratio of transmitted speech according to claim 2, wherein The value of the noise ratio ξ is determined by the sampling frequency and power spectrum of the noisy speech.
4. The method for estimating the signal-to-noise ratio of transmitted voice according to claim 1, characterized in that, In step (4), the starting value and the ending value in the time direction are determined by endpoint detection, and the initial segmentation threshold point α is used for judgment in the frequency direction: the part less than the initial segmentation threshold point α is judged as the global speech, and the part higher than the initial segmentation threshold point α is the edge noise part, thereby determining the effective region of the speech signal in the frequency direction.
5. The method for estimating the signal-to-noise ratio for transmitted speech according to claim 1, characterized in that, In step (5), the secondary signal-noise separation adopts the principal component analysis method to extract the main speech structure, i.e., the voiceprint information, of the global speech part, calculate the proportion of the local voiceprint part in the global speech according to the extracted principal components, and the remaining part is the local noise.
6. A method for estimating the signal-to-noise ratio of transmitted speech according to claim 1, characterized in that, In the step (6), the total power P of the pure voice signal signal = β × P audio , where P audio is the partial power of the local voiceprint, and β ∈ [0, +∞] is an adjustable parameter for weighting the partial power of the local voiceprint.
Citation Information
Patent Citations
Voice signal fundamental frequency estimation method and device
CN114822577A
Method (variants) of filtering the noisy speech signal in complex jamming environment
RU2580796C1