A method, device and electronic device for speech enhancement
Through frequency domain processing and window function technology, the steady-state and non-stable components in the speech signal are accurately estimated and the signal gain is performed, which solves the problem of difficult to identify and suppress non-stable noise and improves the clarity of speech.
Patent Information
- Application Number
- CN202311117201.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-08-31
AI Technical Summary
The prior art is difficult to effectively recognize and suppress non-steady state noise, resulting in a great impact on speech intelligibility and recognition rate.
By converting the input signal from the time domain to the frequency domain, time recursive averaging is performed to obtain a smooth periodic graph, the power spectral density estimates of each component are determined using the causal window and the non-causal window, and the signal gain is performed based on these estimates to suppress non-stable noise.
Spectral estimation of different components is realized, effectively suppressing non-steady state noise and improving speech clarity.
Smart Images

Figure CN116913306B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech signal technology, and in particular to a speech enhancement method, device and electronic device. Background Art
[0002] During driving, non-stationary noise occurs very frequently, such as the sound of horns, brakes, and coughs. Non-stationary noise has a great impact on speech intelligibility and recognition rate, so it is very necessary to study non-stationary noise algorithms. Non-stationary noise is an unexpected and sudden interference to speech input. Compared with speech phonemes, non-stationary noise has a short duration and a large diffusion range in the frequency domain. Traditional speech enhancement technology assumes that the noise is quasi-static, making it possible to estimate the noise power spectral density (PSD) through time smoothing. For non-stationary interference, this assumption no longer holds due to its rapidly changing characteristics.
[0003] At the beginning, a method based on non-local filtering was proposed for speech enhancement of non-stationary noise: first, the non-stationary noise component is enhanced through a modified speech estimator; second, the diffusion map is used to obtain the geometric structure of the non-stationary component, and then the obtained geometric structure is used to estimate the PSD of the non-stationary noise through non-local diffusion filtering. Finally, OM-LSA (optimally modified-log spectral amplitude) is used to suppress non-stationary noise and enhance speech.
[0004] However, a disadvantage of the non-local diffusion filter is that the same pattern of non-stationary noise must appear multiple times, and non-stationary noise that appears only once cannot be identified and therefore cannot be suppressed.
[0005] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve. Summary of the invention
[0006] In view of this, the embodiments of the present application provide a speech enhancement method, device and electronic device to solve the problem in the prior art that non-stable noise is difficult to identify and suppress.
[0007] According to a first aspect of an embodiment of the present application, a speech enhancement method is provided, comprising:
[0008] Convert the input signal from the time domain to the frequency domain to obtain a first spectrum signal;
[0009] Performing time recursive averaging calculation on the first spectrum signal to obtain a smoothed periodogram, wherein the smoothed periodogram includes energy spectrum values of each frequency point in each frame;
[0010] Using a first causal window following each energy spectrum value on the smooth periodogram, taking the minimum value of all energy spectrum values within the first causal window as a first power spectrum level value corresponding to the energy spectrum value being followed;
[0011] Calculating a power spectrum density estimation value of a steady-state component in the first spectrum signal using each first power spectrum level value; the steady-state component includes a speech component and quasi-static noise;
[0012] Using the second causal window and the non-causal window that follow each energy spectrum value on the smooth periodogram at the same time, the larger value is taken as the second power spectrum level value of the energy spectrum value being followed between the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window;
[0013] Calculating a power spectrum density estimation value of a non-stationary component in the first spectrum signal using each second power spectrum level value;
[0014] Performing signal gain on the first spectrum signal according to the power spectrum density estimation value of the steady-state component and the power spectrum density estimation value of the non-steady-state component;
[0015] The first spectrum signal after signal gain is converted from the frequency domain to the time domain to obtain a speech output signal.
[0016] A second aspect of an embodiment of the present application provides a speech enhancement device, comprising:
[0017] A preprocessing module, used for converting an input signal from a time domain to a frequency domain to obtain a first spectrum signal;
[0018] An energy calculation module, used for performing time recursive averaging calculation on the first spectrum signal to obtain a smooth periodogram, wherein the smooth periodogram includes energy spectrum values of each frame and each frequency point;
[0019] A first component calculation module is used to use a first causal window following each energy spectrum value on a smooth periodogram, take the minimum value of all energy spectrum values in the first causal window as a first power spectrum level value corresponding to the energy spectrum value being followed, and also to use each first power spectrum level value to calculate a power spectrum density estimation value of a steady-state component in a first spectrum signal; the steady-state component includes a speech component and a quasi-static noise;
[0020] A second component calculation module is used to use a second causal window and a non-causal window that simultaneously follow each energy spectrum value on the smooth periodogram, and take a larger value between the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window as a second power spectrum level value of the energy spectrum value being followed, and is also used to calculate a power spectrum density estimate of a non-steady-state component in the first spectrum signal using each second power spectrum level value;
[0021] The gain output module is used to perform signal gain on the first spectrum signal according to the power spectrum density estimation value of the steady-state component and the power spectrum density estimation value of the non-steady-state component, and is also used to convert the first spectrum signal after signal gain from the frequency domain to the time domain to obtain a speech output signal.
[0022] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0023] Compared with the prior art, the embodiments of the present application have at least the following beneficial effects: the embodiments of the present application determine the power spectrum density estimation value of each component by utilizing causal windows and non-causal windows for energy spectrum values, thereby more accurately realizing spectrum estimation of different components, thereby performing signal gain, effectively suppressing non-steady-state noise, and improving speech clarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 It is a schematic diagram of an application scenario of an embodiment of the present application;
[0026] Figure 2 It is a flowchart of a speech enhancement method provided in an embodiment of the present application;
[0027] Figure 3 It is a sub-process diagram of a speech enhancement method provided in an embodiment of the present application;
[0028] Figure 4 is a time domain waveform diagram of an input signal provided in an embodiment of the present application;
[0029] Figure 5 is a time domain waveform diagram for estimating transient noise provided by an embodiment of the present application;
[0030] Figure 6 is a time domain waveform diagram of a speech output signal provided in an embodiment of the present application;
[0031] Figure 7 is a spectrogram of an input signal provided in an embodiment of the present application;
[0032] Figure 8 is a spectrogram of a speech output signal provided in an embodiment of the present application;
[0033] Fig. 9 is a structural schematic diagram of a speech enhancement device provided in an embodiment of the present application;
[0034] Fig.10 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0036] A speech enhancement method, device and electronic device according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0037] Figure 1 1 is a schematic diagram of an application scenario of an embodiment of the present application. The application scenario may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a server 104 and a network 105.
[0038] The first terminal device 101 can be hardware or software. When the first terminal device 101 is hardware, it can be various electronic devices with a display screen and supporting communication with the server 104, including but not limited to vehicle systems, smart phones, tablet computers, laptop computers, and desktop computers; when the first terminal device 101 is software, it can be installed in the electronic devices described above. The first terminal device 101 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present application are not limited to this. Furthermore, various applications can be installed on the first terminal device 101, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.
[0039] The second terminal device 102 can be hardware or software. When the second terminal device 102 is hardware, it can be various electronic devices with a display screen and supporting communication with the server 104, including but not limited to a car system, a smart phone, a tablet computer, a laptop computer, and a desktop computer; when the second terminal device 102 is software, it can be installed in the electronic device described above. The second terminal device 102 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiment of the present application does not limit this. Furthermore, various applications can be installed on the second terminal device 102, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.
[0040] The third terminal device 103 can be hardware or software. When the third terminal device 103 is hardware, it can be various electronic devices with a display screen and supporting communication with the server 104, including but not limited to vehicle systems, smart phones, tablet computers, laptop portable computers and desktop computers, etc.; when the third terminal device 103 is software, it can be installed in the electronic devices described above. The third terminal device 103 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present application are not limited to this. Furthermore, various applications can be installed on the third terminal device 103, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.
[0041] The server 104 may be a server that provides various services, for example, a background server that receives a request sent by a terminal device that establishes a communication connection with the server, and the background server may receive and analyze the request sent by the terminal device, and generate a processing result. The server 104 may be a single server, or a server cluster composed of several servers, or a cloud computing service center, which is not limited in the present embodiment of the application.
[0042] It should be noted that the server 104 can be hardware or software. When the server 104 is hardware, it can be various electronic devices that provide various services for the first terminal device 101, the second terminal device 102, and the third terminal device 103. When the server 104 is software, it can be multiple software or software modules that provide various services for the first terminal device 101, the second terminal device 102, and the third terminal device 103, or it can be a single software or software module that provides various services for the first terminal device 101, the second terminal device 102, and the third terminal device 103, and the embodiment of the present application does not limit this.
[0043] The network 105 can be a wired network connected by coaxial cable, twisted pair and optical fiber, or it can be a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), infrared, etc., which is not limited in the embodiments of the present application.
[0044] It should be noted that the specific types, quantities and combinations of the first terminal device 101, the second terminal device 102, the third terminal device 103, the server 104 and the network 105 can be adjusted according to the actual needs of the application scenario, and the embodiments of the present application do not limit this.
[0045] Figure 2 It is a flow chart of a speech enhancement method provided in an embodiment of the present application. Figure 2 The speech enhancement method can be Figure 1 The first terminal device, the second terminal device, the third terminal device or the server is executed, and it can be especially applied to the vehicle-mounted voice analysis scenario of the smart cockpit. Figure 2 As shown, the speech enhancement method comprises:
[0046] S201: Convert an input signal from the time domain to the frequency domain to obtain a first spectrum signal;
[0047] S202: performing time recursive averaging calculation on the first spectrum signal to obtain a smoothed periodogram, where the smoothed periodogram includes energy spectrum values of each frequency point in each frame;
[0048] S203: using a first causal window following each energy spectrum value on the smooth periodogram, taking the minimum value of all energy spectrum values within the first causal window as a first power spectrum level value corresponding to the energy spectrum value being followed;
[0049] S204: Calculate a power spectrum density estimation value of a steady-state component in the first spectrum signal using each first power spectrum level value; the steady-state component includes a speech component and quasi-static noise;
[0050] S205: using the second causal window and the non-causal window that simultaneously follow each energy spectrum value on the smooth periodogram, taking the larger value between the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window as the second power spectrum level value of the energy spectrum value being followed;
[0051] S206: Calculate a power spectrum density estimation value of a non-stationary component in the first spectrum signal using each second power spectrum level value;
[0052] S207: performing signal gain on the first spectrum signal according to the power spectrum density estimation value of the steady-state component and the power spectrum density estimation value of the non-steady-state component;
[0053] S208: Convert the first spectrum signal after signal gain from the frequency domain to the time domain to obtain a speech output signal.
[0054] It can be understood that the input signal includes speech components, quasi-static noise, and transient noise, among which the speech components and quasi-static noise are relatively stable, and the transient noise changes relatively faster. Therefore, when tracking the rapid changes in the PSD of transient noise, the speech components and quasi-static noise can be regarded as pseudo-stable steady-state components, and the transient noise is a non-steady-state component different from the steady-state component.
[0055] The input signal initially received in step S201 is originally in the time domain, and the speech output signal to be output in the final step S208 is also in the time domain. In this embodiment, the speech enhancement analysis and processing of the signal in the entire method are performed in the frequency domain. Therefore, it is necessary to convert the input signal from the time domain to the frequency domain in step S201, and convert the first spectrum signal after signal gain from the frequency domain to the time domain in step S208.
[0056] Specifically, step S201 converts the input signal from the time domain to the frequency domain, including windowing the input signal and performing STFT (Short-Time Fourier Transform) transformation. The input signal y(n) in the time domain includes an additive speech signal x(n), quasi-static noise d(n) and transient noise t(n), which can be expressed as:
[0057] y(n)=x(n)+d(n)+t(n);
[0058] The process of windowing and STFT transforming the input signal y(n) to obtain the first spectrum signal Y(k,l) can be expressed as:
[0059]
[0060] Where k is the frequency index, l is the frame index, h(n) is the N-point analysis window, which can be a Hamming window, M is the number of overlapping points between adjacent frames, and Y(k,l) is the signal amplitude of the kth frequency point in the lth frame of the first spectrum signal. It can be understood that the use of short-time frames will reduce the large changes in speech data between adjacent frames, but at the same time, short-time frames will also reduce frequency resolution. After experimental verification, when the sampling rate is 16kHz, the number of points per frame is 64, which will achieve better results. At this time, the duration of each frame is 4ms.
[0061] At this time, the first spectrum signal obtained after conversion to the frequency domain can be expressed as: Y(k,l)=X(k,l)+T(k,l)+D(k,l), where X(k,l), T(k,l) and D(k,l) are the STFT transform components of the speech signal x(n), quasi-static noise d(n) and transient noise t(n), respectively.
[0062] It can be understood that in step S208, the action of converting the first spectrum signal after signal gain from the frequency domain to the time domain is exactly the opposite of step S201. The first spectrum signal after signal gain is subjected to IFFT (Inverse Fast Fourier Transform) transformation, and then passed through a synthesis window to finally output a speech output signal in the time domain.
[0063] Furthermore, step S202 performs time recursive averaging calculation on the first spectrum signal to obtain a smoothed periodogram, including:
[0064] The first spectrum signal is subjected to time recursive averaging calculation according to a time recursive averaging formula to obtain a smoothed periodogram. The time recursive averaging formula includes:
[0065] S(k,l)=α s S(k,l-1)+(1-α s )|Y(k,l)| 2 ;
[0066] Among them, Y(k,l) is the signal amplitude of the kth frequency point in the lth frame of the first spectrum signal, S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, α s is the preset smoothing parameter.
[0067] In the time recursive average formula, the smaller the preset smoothing parameter is, the greater the weight of the current frame is, so that the rapid changes of the power spectrum density estimation value of each component can be captured. Therefore, in this embodiment, a smaller value can be assigned to the preset smoothing parameter, so the value of the preset smoothing parameter in this embodiment is less than 0.9. The energy spectrum value is used to determine the power spectrum density estimation value of the steady-state component in steps S203-S204, and is also used to determine the power spectrum density estimation value of the non-steady-state component in steps S205-S206. Therefore, different preset smoothing parameters can be set for different components to calculate their corresponding energy spectrum values. For example, for the smooth periodogram used in step S203, its preset smoothing parameter can be set to 0.7, and for the smooth periodogram used in step S205, its preset smoothing parameter can be set to 0.85. Of course, in addition to these two specific values, different preset smoothing parameters can be selected according to actual needs and actual signal performance, such as 0.75, 0.8, etc., which are not limited here.
[0068] Further, step S203 uses the first causal window following each energy spectrum value on the smooth periodogram, and takes the minimum value of all energy spectrum values in the first causal window as the first power spectrum level value corresponding to the energy spectrum value being followed. Assuming that the length of the first causal window is L0, the first power spectrum level value can be expressed as:
[0069]
[0070] Where S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, L0 is the total number of frames or the length of the first causal window, is the minimum value of all energy spectrum values in the first causal window. At this time, the number of sub-windows can be 2, and the number of frames contained in each sub-window can be 5.
[0071] Further, the process of calculating the power spectrum density estimation value of the steady-state component in the first spectrum signal by using each first power spectrum level value in step S204 is as follows: Figure 3 As shown, including:
[0072] S301: Determine whether the ratio of each energy spectrum value to the corresponding first power spectrum level value is greater than a first preset threshold; if so, determine the first indication value to be 1; if not, determine the first indication value to be 0;
[0073] S302: Obtaining the existence probability of the non-steady-state component through smoothing recursive calculation according to the first indicator value and a preset probability smoothing factor;
[0074] S303: determining an estimated smoothing factor according to the existence probability and a preset fixed smoothing factor;
[0075] S304: Obtain a power spectrum density estimation value of a steady-state component in the first spectrum signal through smoothing recursive calculation according to the first spectrum signal and the estimated smoothing factor.
[0076] Specifically, the ratio of each energy spectrum value to the corresponding first power spectrum level value determined in step S301 can be expressed as S r (k,l), specifically:
[0077]
[0078] Further, the first preset threshold can be expressed as δ, and the value of the first indicator value I(k,l) is as follows:
[0079]
[0080] The first preset threshold value may be 1.67, which is not limited here, and the specific value may be adjusted according to the scenario.
[0081] Step S302 is a process of obtaining the existence probability of the non-steady-state component by smoothing recursively calculating according to the first indicator value and the preset probability smoothing factor, including:
[0082] According to the first indicator value and the preset probability smoothing factor, the existence probability of the non-steady-state component is calculated according to the first smoothing recursive formula; the first smoothing recursive formula includes:
[0083] p(k,l)=α p p(k,l-1)+(1-α p )I(k,l);
[0084] Where I(k,l) is the first indicator value of the kth frequency point in the lth frame, α p is the preset probability smoothing factor, p(k,l) is the existence probability of the non-steady-state component of the kth frequency point in the lth frame, and its initial value is the initial value of I(k,l).
[0085] The further step S303 is a process of determining the estimated smoothing factor according to the existence probability and the preset fixed smoothing factor, including:
[0086] The estimated smoothing factor is determined by calculating the estimated factor calculation formula according to the existence probability and the preset fixed smoothing factor. The estimated factor calculation formula includes:
[0087]
[0088] in, is the estimated smoothing factor of the kth frequency point in the lth frame; α is the preset fixed smoothing factor.
[0089] The last step S304 obtains the power spectrum density estimation value of the steady-state component in the first spectrum signal by smoothing recursively calculating the first spectrum signal and the estimated smoothing factor, including:
[0090] According to the first spectrum signal and the estimated smoothing factor, the power spectrum density estimation value of the steady-state component in the first spectrum signal is obtained according to the second smoothing recursive formula; the second smoothing recursive formula includes:
[0091]
[0092] Wherein, Y(k,l) is the signal amplitude of the kth frequency point in the lth frame of the first spectrum signal, is the power spectral density estimate of the steady-state component of the kth frequency point in the lth frame.
[0093] It can be understood that the above preset probability smoothing factor and preset fixed smoothing factor have values between 0 and 1, wherein the preset probability smoothing factor can be 0.7 and the preset fixed smoothing factor can be 0.5, which can be adjusted according to actual conditions.
[0094] It can be understood that most of the speech components and quasi-static noise can be captured by steps S203-S204. However, the beginning of the speech factor will be mistakenly determined as a sudden non-steady-state component, which cannot be captured by the tiled recursive smoothing of steps S203-S204 and therefore cannot be suppressed. For this reason, it is necessary to use a non-causal window that considers future time frames to distinguish the beginning of the speech factor and the non-steady-state component. The power of the non-steady-state component will drop sharply after a very short period of time, while the power of the speech factor is stable after the beginning. Therefore, it is necessary to add a non-causal window to complete the distinction between the two, namely steps S205-S206.
[0095] Specifically, step S205 uses the second causal window and the non-causal window that follow each energy spectrum value on the smooth periodogram at the same time, and takes the larger value of the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window as the second power spectrum level value of the energy spectrum value to be followed, including:
[0096] The minimum value of all energy spectrum values within the second causal window following each energy spectrum value is determined using the causal window formula, where the causal window formula is:
[0097]
[0098] Where S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, L is the total number of frames in the second causal window, is the minimum value of all energy spectrum values within the second causal window;
[0099] The non-causal window formula is used to determine the minimum value of all energy spectrum values within the non-causal window following each energy spectrum value, where the non-causal window formula is:
[0100]
[0101] Where T is the total number of frames in the non-causal window, is the minimum value of all energy spectrum values within the non-causal window;
[0102] The larger value among the minimum value of all energy spectrum values within the second causal window and the minimum value of all energy spectrum values within the non-causal window is taken as the second power spectrum level value of the followed energy spectrum value.
[0103] At this time, the second power spectrum level value can be expressed as This maximum value is used as the power spectrum level value of the non-steady-state component, thereby ensuring that the spectrum at the beginning of the speech phoneme is compared with the following speech phoneme instead of being compared with the previous environmental noise. Using the new decision criterion, the beginning part of the speech phoneme can be effectively suppressed.
[0104] It can be understood that the above determination of the power spectral density estimate mainly uses the MCRA (MinimaControlled Recursive Averaging) control idea, but the window used is either a causal window or a causal window and a non-causal window according to the type of component to be calculated.
[0105] It can be understood that the calculation of the power spectrum density estimation value of the non-steady-state component in step S206 is similar to that in step S204. It should be noted that the power spectrum level values used in these two steps are different.
[0106] Step S207 is to calculate the signal gain of the first spectrum signal, and specifically implement the steps of calculating and filtering the spectrum gain value through a filter after obtaining the power spectrum density estimation values of the steady-state component and the non-steady-state component, and superimposing the spectrum gain value with the signal of the first spectrum signal.
[0107] Finally, in step S208, the first spectrum signal after signal gain is converted from the frequency domain to the time domain, and the non-steady-state component in the obtained speech output signal is suppressed, thereby obtaining a clearer speech effect.
[0108] by Figure 4 For example, the input signal Figure 4 is the time domain waveform of the input signal. Figure 4 The input signal is used to perform the speech enhancement method in this embodiment, and the obtained Figure 5 and Figure 6 ,in Figure 5 The time domain waveform diagram of the transient noise estimated by the method of this embodiment has a relatively accurate estimation effect. Figure 6 The time domain waveform diagram of the voice output signal output by the method of this embodiment has a relatively clear and accurate voice effect. Figure 4 The spectrogram of the input signal is as follows Figure 7 As shown, Figure 6 The spectrogram corresponding to the speech output signal is as follows Figure 8 As shown, comparison Figure 7 and Figure 8 It can also be found from the spectrogram that the method of this embodiment has a good suppression effect on transient noise and quasi-static noise, while causing little damage to the speech and well preserving the characteristics of the speech.
[0109] The method of the embodiment of the present application determines the power spectral density estimation value of each component by utilizing causal windows and non-causal windows for energy spectrum values, thereby realizing more accurate spectrum estimation of different components, thereby performing signal gain, effectively suppressing non-steady-state noise, and improving speech clarity.
[0110] All the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here. It should be understood that the size of the sequence number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0111] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0112] Fig. 9 Schematic diagram of a speech enhancement device provided in an embodiment of the present application. Fig. 9 As shown, the speech enhancement device comprises:
[0113] The preprocessing module 901 is used to convert the input signal from the time domain to the frequency domain to obtain a first spectrum signal;
[0114] An energy calculation module 902 is used to perform time recursive averaging calculation on the first spectrum signal to obtain a smooth periodogram, where the smooth periodogram includes energy spectrum values of each frame and each frequency point;
[0115] The first component calculation module 903 is used to use the first causal window following each energy spectrum value on the smooth periodogram, take the minimum value of all energy spectrum values in the first causal window as the first power spectrum level value corresponding to the energy spectrum value being followed, and also to calculate the power spectrum density estimation value of the steady-state component in the first spectrum signal using each first power spectrum level value; the steady-state component includes the speech component and the quasi-static noise;
[0116] A second component calculation module 904 is used to use the second causal window and the non-causal window that follow each energy spectrum value on the smooth periodogram at the same time, and take the larger value between the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window as the second power spectrum level value of the energy spectrum value being followed, and is also used to calculate the power spectrum density estimation value of the non-steady-state component in the first spectrum signal using each second power spectrum level value;
[0117] The gain output module 905 is used to perform signal gain on the first spectrum signal according to the power spectrum density estimation value of the steady-state component and the power spectrum density estimation value of the non-steady-state component, and is also used to convert the first spectrum signal after signal gain from the frequency domain to the time domain to obtain a speech output signal.
[0118] The device of the embodiment of the present application determines the power spectral density estimation value of each component by utilizing causal windows and non-causal windows for the energy spectrum value, thereby realizing more accurate spectrum estimation of different components, thereby performing signal gain, effectively suppressing non-steady-state noise, and improving speech clarity.
[0119] In an exemplary embodiment, the process of performing time recursive averaging calculation on the first spectrum signal to obtain a smoothed periodogram includes:
[0120] The first spectrum signal is subjected to time recursive averaging calculation according to a time recursive averaging formula to obtain a smoothed periodogram. The time recursive averaging formula includes:
[0121] S(k,l)=α s S(k,l-1)+(1-α s )|Y(k,l)| 2 ;
[0122] Among them, Y(k,l) is the signal amplitude of the kth frequency point in the lth frame of the first spectrum signal, S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, α s is the preset smoothing parameter.
[0123] In an exemplary embodiment, the value of the preset smoothing parameter is less than 0.9.
[0124] In an exemplary embodiment, the process of calculating the power spectrum density estimation value of the steady-state component in the first spectrum signal using each first power spectrum level value includes:
[0125] Determine whether the ratio of each energy spectrum value to the corresponding first power spectrum level value is greater than a first preset threshold;
[0126] If yes, determine the first indication value to be 1, if no, determine the first indication value to be 0;
[0127] According to the first indicator value and the preset probability smoothing factor, the existence probability of the non-steady-state component is obtained by smoothing recursively calculating;
[0128] Determine an estimated smoothing factor according to the existence probability and a preset fixed smoothing factor;
[0129] According to the first spectrum signal and the estimated smoothing factor, a power spectrum density estimation value of the steady-state component in the first spectrum signal is obtained through smoothing recursive calculation.
[0130] In an exemplary embodiment, the process of obtaining the existence probability of the non-steady-state component by smoothing recursively calculating according to the first indicator value and the preset probability smoothing factor includes:
[0131] According to the first indicator value and the preset probability smoothing factor, the existence probability of the non-steady-state component is calculated according to the first smoothing recursive formula; the first smoothing recursive formula includes:
[0132] p(k,l)=α p p(k,l-1)+(1-α p )I(k,l);
[0133] Where I(k,l) is the first indicator value of the kth frequency point in the lth frame, α p is the preset probability smoothing factor, and p(k,l) is the existence probability of the non-stationary component of the kth frequency point in the lth frame.
[0134] In an exemplary embodiment, the process of determining the estimated smoothing factor according to the existence probability and the preset fixed smoothing factor includes:
[0135] The estimated smoothing factor is determined by calculating the estimated factor calculation formula according to the existence probability and the preset fixed smoothing factor. The estimated factor calculation formula includes:
[0136]
[0137] in, is the estimated smoothing factor of the kth frequency point in the lth frame; α is the preset fixed smoothing factor.
[0138] In an exemplary embodiment, obtaining the power spectrum density estimate of the steady-state component in the first spectrum signal by smoothing recursive calculation according to the first spectrum signal and the estimated smoothing factor includes:
[0139] According to the first spectrum signal and the estimated smoothing factor, the power spectrum density estimation value of the steady-state component in the first spectrum signal is obtained according to the second smoothing recursive formula; the second smoothing recursive formula includes:
[0140]
[0141] Wherein, Y(k,l) is the signal amplitude of the kth frequency point in the lth frame of the first spectrum signal, is the power spectral density estimate of the steady-state component of the kth frequency point in the lth frame.
[0142] In an exemplary embodiment, the process of using a second causal window and a non-causal window that simultaneously follow each energy spectrum value on a smooth periodogram, and taking a larger value between the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window as the second power spectrum level value of the energy spectrum value being followed, includes:
[0143] The minimum value of all energy spectrum values within the second causal window following each energy spectrum value is determined using the causal window formula, where the causal window formula is:
[0144]
[0145] Where S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, L is the total number of frames in the second causal window, is the minimum value of all energy spectrum values within the second causal window;
[0146] The non-causal window formula is used to determine the minimum value of all energy spectrum values within the non-causal window following each energy spectrum value, where the non-causal window formula is:
[0147]
[0148] Where T is the total number of frames in the non-causal window, is the minimum value of all energy spectrum values within the non-causal window;
[0149] The larger value among the minimum value of all energy spectrum values in the second causal window and the minimum value of all energy spectrum values in the non-causal window is taken as the second power spectrum level value of the energy spectrum value to be followed.
[0150] Fig.10 Schematic diagram of an electronic device 10 provided in an embodiment of the present application. Fig.10 As shown, the electronic device 10 of this embodiment includes: a processor 1001, a memory 1002, and a computer program 1003 stored in the memory 1002 and executable on the processor 1001. When the processor 1001 executes the computer program 1003, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 1001 executes the computer program 1003, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0151] The electronic device 10 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 10 may include, but is not limited to, a processor 1001 and a memory 1002. Those skilled in the art will appreciate that Fig.10 The electronic device 10 is merely an example and does not limit the electronic device 10 . The electronic device 10 may include more or fewer components than shown in the figure, or different components.
[0152] Processor 1001 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0153] The memory 1002 may be an internal storage unit of the electronic device 10, for example, a hard disk or memory of the electronic device 10. The memory 1002 may also be an external storage device of the electronic device 10, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card (FlashCard), etc. equipped on the electronic device 10. The memory 1002 may also include both an internal storage unit of the electronic device 10 and an external storage device. The memory 1002 is used to store computer programs and other programs and data required by the electronic device.
[0154] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0155] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium, such as a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The readable storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0156] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A speech enhancement method, characterized in that: include: Convert the input signal from the time domain to the frequency domain to obtain a first spectrum signal; Performing time recursive averaging calculation on the first spectrum signal to obtain a smoothed periodogram, wherein the smoothed periodogram includes energy spectrum values of each frequency point in each frame; Using a first causal window following each of the energy spectrum values on the smooth periodogram, taking the minimum value of all the energy spectrum values within the first causal window as a first power spectrum level value corresponding to the energy spectrum value being followed; Calculating a power spectrum density estimation value of a steady-state component in the first spectrum signal using each of the first power spectrum level values; the steady-state component includes a speech component and quasi-static noise; Utilizing a second causal window and a non-causal window that simultaneously follow each of the energy spectrum values on the smooth periodogram, a larger value is taken between the minimum value of all the energy spectrum values in the second causal window and the minimum value of all the energy spectrum values in the non-causal window as a second power spectrum level value of the energy spectrum value being followed; Calculating a power spectrum density estimation value of a non-stationary component in the first spectrum signal using each of the second power spectrum level values; Performing signal gain on the first spectrum signal according to the power spectrum density estimation value of the steady-state component and the power spectrum density estimation value of the non-steady-state component; The first spectrum signal after signal gain is converted from the frequency domain to the time domain to obtain a speech output signal.
2. The method according to claim 1, characterized in that The process of performing time recursive averaging calculation on the first spectrum signal to obtain a smoothed periodogram includes: The first spectrum signal is subjected to time recursive averaging calculation according to a time recursive averaging formula to obtain a smoothed periodogram, wherein the time recursive averaging formula includes: S(k,l)=a s S(k,l-1)+(1-a s )|Y(k,l)| 2 ; Wherein, Y(k,l) is the signal amplitude of the kth frequency point in the lth frame of the first spectrum signal, S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, α s is the preset smoothing parameter.
3. The method according to claim 2, characterized in that The value of the preset smoothing parameter is less than 0.
9.
4. The method according to claim 1, characterized in that The process of calculating the power spectrum density estimation value of the steady-state component in the first spectrum signal by using each of the first power spectrum level values includes: Determine whether the ratio of each of the energy spectrum values to the corresponding first power spectrum level value is greater than a first preset threshold; If yes, determine the first indication value to be 1, if no, determine the first indication value to be 0; According to the first indicator value and a preset probability smoothing factor, obtaining the existence probability of the non-steady-state component through smoothing recursive calculation; Determining an estimated smoothing factor according to the existence probability and a preset fixed smoothing factor; According to the first spectrum signal and the estimated smoothing factor, a power spectrum density estimation value of the steady-state component in the first spectrum signal is obtained through smoothing recursive calculation.
5. The method according to claim 4, characterized in that The process of obtaining the existence probability of the non-steady-state component by smoothing recursively calculating according to the first indicator value and the preset probability smoothing factor includes: According to the first indicator value and the preset probability smoothing factor, the existence probability of the non-steady-state component is calculated according to the first smoothing recursive formula; the first smoothing recursive formula includes: p(k,l)=a p p(k,l-1)+(1-a p )I(k,l); Where I(k,l) is the first indication value of the kth frequency point in the lth frame, α p is the preset probability smoothing factor, and p(k,l) is the existence probability of the non-steady-state component of the kth frequency point in the lth frame.
6. The method according to claim 5, characterized in that The process of determining the estimated smoothing factor according to the existence probability and the preset fixed smoothing factor includes: The estimated smoothing factor is determined by calculating the estimated factor calculation formula according to the existence probability and the preset fixed smoothing factor, and the estimated factor calculation formula includes: in, is the estimated smoothing factor of the kth frequency point in the lth frame; α is the preset fixed smoothing factor.
7. The method according to claim 6, characterized in that Obtaining a power spectrum density estimate of a steady-state component in the first spectrum signal by smoothing recursive calculation according to the first spectrum signal and the estimated smoothing factor includes: According to the first spectrum signal and the estimated smoothing factor, a power spectrum density estimation value of the steady-state component in the first spectrum signal is obtained by calculating according to a second smoothing recursive formula; the second smoothing recursive formula includes: Wherein, Y(k,l) is the signal amplitude of the kth frequency point in the lth frame in the first spectrum signal, is the power spectral density estimate of the steady-state component at the kth frequency point in the lth frame.
8. The method according to any one of claims 1 to 7, characterized in that The process of using the second causal window and the non-causal window that simultaneously follow each of the energy spectrum values on the smooth periodogram, and taking the larger value between the minimum value of all the energy spectrum values in the second causal window and the minimum value of all the energy spectrum values in the non-causal window as the second power spectrum level value of the energy spectrum value to be followed, comprises: The minimum value of all the energy spectrum values within a second causal window following each of the energy spectrum values is determined using a causal window formula, wherein the causal window formula is: Where S(k,l) is the energy spectrum value of the kth frequency point in the lth frame, L is the total number of frames in the second causal window, is the minimum value of all the energy spectrum values within the second causal window; The minimum value of all the energy spectrum values within the non-causal window following each of the energy spectrum values is determined using a non-causal window formula, wherein the non-causal window formula is: Where T is the total number of frames of the non-causal window, is the minimum value of all the energy spectrum values within the non-causal window; A larger value is taken between the minimum value of all the energy spectrum values within the second causal window and the minimum value of all the energy spectrum values within the non-causal window as the second power spectrum level value of the energy spectrum value to be followed.
9. A speech enhancement device, characterized in that: include: A preprocessing module, used for converting an input signal from a time domain to a frequency domain to obtain a first spectrum signal; An energy calculation module, used for performing time recursive averaging calculation on the first spectrum signal to obtain a smooth periodogram, wherein the smooth periodogram includes energy spectrum values of each frame and each frequency point; A first component calculation module is used to use a first causal window following each of the energy spectrum values on the smooth periodogram, take the minimum value of all the energy spectrum values in the first causal window as a first power spectrum level value corresponding to the energy spectrum value being followed, and is also used to calculate a power spectrum density estimate of a steady-state component in the first spectrum signal using each of the first power spectrum level values; the steady-state component includes a speech component and a quasi-static noise; A second component calculation module is used to use a second causal window and a non-causal window that simultaneously follow each of the energy spectrum values on the smooth periodogram, and take a larger value among the minimum value of all the energy spectrum values in the second causal window and the minimum value of all the energy spectrum values in the non-causal window as a second power spectrum level value of the energy spectrum value being followed, and is also used to calculate a power spectrum density estimate of a non-steady-state component in the first frequency spectrum signal using each of the second power spectrum level values; The gain output module is used to perform signal gain on the first spectrum signal according to the power spectrum density estimation value of the steady-state component and the power spectrum density estimation value of the non-steady-state component, and is also used to convert the first spectrum signal after signal gain from the frequency domain to the time domain to obtain a speech output signal.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Transient noise suppression method based on spectrum estimation
CN103456310A
Transient noise suppression-oriented real-time speech enhancement method
CN110739005A