A speech noise reduction method, apparatus, electronic device and storage medium
The speech denoising method using frequency band segmentation and dynamic threshold adjustment solves the problems of real-time performance and resource consumption on microprocessors, achieving low-latency and efficient speech denoising.
Patent Information
- Application Number
- CN202511223293.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing technologies struggle to balance real-time performance with low resource consumption for voice noise reduction on microprocessors, especially in single-channel architectures where it is difficult to distinguish between human voices and noise, resulting in poor noise reduction performance or excessive resource consumption.
By segmenting the speech signal into frequency bands and using a dual second-order filter to divide the signal into human voice and noise sub-bands in the time domain, noise reduction is performed by combining intensity and bandwidth parameters, and the threshold and noise reduction coefficient are dynamically adjusted to achieve targeted suppression of noise.
It effectively distinguishes human voice from noise under low latency conditions, reduces computational complexity and resource consumption, and improves the accuracy and efficiency of noise reduction. It is suitable for single-channel structures.
Smart Images

Figure CN120708640B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech signal processing technology, specifically to a speech noise reduction method, apparatus, electronic device, and storage medium. Background Technology
[0002] Noise reduction methods in related technologies include: time-domain based noise reduction, such as adaptive filters, subspace decomposition, Kalman filtering, and time-domain Wiener filtering; and frequency-domain based noise reduction, such as spectral subtraction, frequency-domain Wiener filtering, and optimal improved logarithmic spectral amplitude estimation + minimum tracking method. Adaptive filter noise reduction requires reference noise and is not suitable for single-channel structures. Subspace decomposition algorithms are computationally complex and resource-intensive. Time-domain Wiener filtering can only reduce high-frequency noise and has limited suppression of in-band noise. Frequency-domain filtering requires Fourier transform, which is affected by window length and results in high latency. There are technical challenges in balancing real-time performance with low resource consumption on microprocessors. Summary of the Invention
[0003] This application provides a speech noise reduction method, apparatus, electronic device, and storage medium, which can solve the technical problem of balancing real-time performance and low resource consumption on microprocessors.
[0004] Firstly, this embodiment provides a speech noise reduction method, including:
[0005] Acquire the first time-domain signal of the speech to be denoised, which includes the target sampling point;
[0006] The first time-domain signal is subjected to frequency band segmentation to obtain multiple sub-band time-domain signals; each sub-band time-domain signal includes the signal corresponding to the target sampling point.
[0007] Based on the intensity parameters corresponding to each of the sub-band time-domain signals, determine the first sub-band time-domain signal of human voice and the second sub-band time-domain signal of noise from among the multiple sub-band time-domain signals;
[0008] The second sub-band time-domain signal is denoised to obtain the third sub-band time-domain signal.
[0009] The second time-domain signal after denoising of the speech to be denoised is determined based on the first sub-band time-domain signal and the third sub-band time-domain signal.
[0010] In some embodiments, determining a second sub-band time-domain signal that is noise among the plurality of sub-band time-domain signals based on the intensity parameter corresponding to each sub-band time-domain signal includes:
[0011] Determine whether the value of the strength parameter is less than a preset strength threshold to obtain a first determination result;
[0012] If the first determination result indicates that the value of the intensity parameter of the sub-band time domain signal is less than the intensity threshold, then the sub-band time domain signal is determined as the second sub-band time domain signal.
[0013] In some embodiments, the noise reduction processing of the second sub-band time-domain signal to obtain the third sub-band time-domain signal includes:
[0014] Determine whether the value of the bandwidth parameter of the second sub-band time domain signal is greater than a preset bandwidth threshold to obtain a second determination result;
[0015] If the second judgment result indicates that the value of the bandwidth parameter of the second sub-band time domain signal is greater than the bandwidth threshold, the second sub-band time domain signal is denoised according to the preset first denoising coefficient to obtain the third sub-band time domain signal.
[0016] In some embodiments, the method further includes:
[0017] Obtain the initial time-domain signal and initial energy parameters;
[0018] The average intensity parameter is determined based on the initial time-domain signal and the initial energy parameter;
[0019] The intensity threshold is determined based on the average intensity parameter and the preset intensity coefficient.
[0020] In some embodiments, after determining the second sub-band time-domain signal that is noise among the plurality of sub-band time-domain signals based on the intensity parameter corresponding to each sub-band time-domain signal, the method further includes:
[0021] The intensity threshold is updated based on the intensity parameters of the second sub-band time-domain signal to obtain the updated intensity threshold.
[0022] In some embodiments, determining the denoised second time-domain signal based on the first sub-band time-domain signal and the third sub-band time-domain signal includes:
[0023] The first sub-band time domain signal is denoised according to the preset second denoising coefficient to obtain the fourth sub-band time domain signal;
[0024] The third sub-band time-domain signal and the fourth sub-band time-domain signal are merged to obtain the second time-domain signal.
[0025] In some embodiments, the merging process of the third sub-band time-domain signal and the fourth sub-band time-domain signal to obtain the second time-domain signal includes:
[0026] Determine the overlap interval between the third sub-band time-domain signal and the fourth sub-band time-domain signal;
[0027] The first signal value of the time-domain signal of the third sub-band in the overlapping interval is weighted based on a preset first gain parameter to obtain a first weighted value; and the second signal value of the time-domain signal of the fourth sub-band in the overlapping interval is weighted based on a preset second gain parameter to obtain a second weighted value; the sum of the values of the first gain parameter and the second gain parameter is a preset threshold.
[0028] The target signal value of the second time-domain signal is determined based on the first weighted value and the second weighted value.
[0029] Secondly, this embodiment also provides a voice noise reduction device, including:
[0030] The first acquisition module is used to acquire the first time-domain signal in the speech to be denoised, which includes the target sampling point;
[0031] The segmentation module is used to perform frequency band segmentation processing on the first time domain signal to obtain multiple sub-band time domain signals; each sub-band time domain signal includes the signal corresponding to the target sampling point;
[0032] The first determining module is used to determine, based on the intensity parameter corresponding to each of the sub-band time-domain signals, a first sub-band time-domain signal of human voice and a second sub-band time-domain signal of noise among the multiple sub-band time-domain signals;
[0033] The first noise reduction module is used to perform noise reduction processing on the second sub-band time domain signal to obtain the third sub-band time domain signal;
[0034] The second determining module is used to determine the second time-domain signal of the speech to be denoised after denoising based on the first sub-band time-domain signal and the third sub-band time-domain signal.
[0035] Thirdly, this embodiment also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the method described above.
[0036] Fourthly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in the method described.
[0037] In embodiments of the present invention, a first time-domain signal containing target sampling points is subjected to frequency band segmentation to obtain multiple sub-band time-domain signals. Based on the intensity parameters corresponding to each sub-band time-domain signal, a first sub-band time-domain signal representing human voice and a second sub-band time-domain signal representing noise are distinguished. After denoising the second sub-band time-domain signal, the denoised second time-domain signal is determined by combining it with the first sub-band time-domain signal. This allows for application to a single-channel structure without the need for reference noise, reducing computational complexity and resource consumption. At the same time, targeted processing of noise in different frequency bands improves the technical problem in related technologies where it is difficult to balance real-time performance with low microprocessor resource requirements. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 A schematic flowchart of a speech noise reduction method provided in an embodiment of this application;
[0040] Figure 2 This is another schematic flowchart illustrating a speech noise reduction method provided in an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of the structure of a speech noise reduction device provided in an embodiment of this application;
[0042] Figure 4 This is a schematic diagram of the hardware structure of a voice noise reduction device provided in an embodiment of this application. Detailed Implementation
[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0044] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0045] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0046] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not preclude applicability to or configuration to devices performing additional tasks or steps. Furthermore, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more conditions or values may in practice be based on additional conditions or values beyond those conditions.
[0047] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0048] Low-latency noise reduction is a function of a recorder or audio acquisition device used to reduce background noise in audio to improve audio clarity and ensure that the audio data is transmitted without too much delay due to the introduction of algorithms, thus ensuring the real-time nature of the audio data.
[0049] Considering real-time performance and the need for low computational resource consumption when running on microprocessors, this application proposes a time-domain low-latency speech denoising algorithm. This algorithm can meet the low-latency audio denoising requirements of live broadcasts of battlefields or sporting events, live stage performances, and film crew recordings, while also running on processors with limited resources, thus reducing costs.
[0050] Figure 1 This is a flowchart illustrating a speech noise reduction method provided in an embodiment of this application, as shown below. Figure 1 As shown, the audio file tagging method provided in this application embodiment may include, but is not limited to, the following steps and combinations thereof.
[0051] Step 101: Obtain the first time-domain signal of the speech to be denoised, including the target sampling point.
[0052] In this embodiment, the speech to be denoised can be the original speech containing noise, which can include both effective speech such as human voices that need to be preserved, and interfering speech such as environmental noise and electronic noise. The speech to be denoised can be determined according to the actual situation and is not limited here. As an example, the speech to be denoised can be from an audio acquisition device such as a microphone or recording equipment. The target sampling point can be a specific sampling point selected as the processing object in the time-domain waveform of the speech to be denoised. The target sampling point can be any one or more sampling points in the speech to be denoised. As an example, the target sampling point can be a target number of continuous sampling points, where the target number can be 10. The first time-domain signal can be a segment of time-domain signal extracted from the speech to be denoised and containing the above-mentioned target sampling points. The time-domain signal can be a waveform signal with time as the horizontal axis and signal amplitude as the vertical axis, and the time-domain signal can directly reflect the change characteristics of the signal over time.
[0053] In this embodiment, obtaining the first time-domain signal including the target sampling points in the speech to be denoised can be achieved by performing analog-to-digital conversion on the analog signal of the speech to be denoised, obtaining a discrete sampling point sequence arranged in chronological order; starting from the beginning of the sampling point sequence, the sampling points are sequentially stored in a buffer. When the number of sampling points in the buffer reaches the target number, an initial window is formed; the sampling points corresponding to the initial window are determined as the first time-domain signal. It should be noted that when a new sampling point is input, a sliding window mechanism is used to remove the earliest stored sampling point in the buffer, and a new sampling point is stored at the same time, so that the buffer always maintains the target number of sampling points, forming an updated window, and the sampling points corresponding to the updated window are determined as the first time-domain signal.
[0054] Step 102: Perform frequency band segmentation on the first time domain signal to obtain multiple sub-band time domain signals; each sub-band time domain signal includes the signal corresponding to the target sampling point.
[0055] In this embodiment, the sub-band time domain signal can be a time domain signal obtained by frequency band division of the first time domain signal. The sub-band time domain signal can correspond to a time domain signal of a specific frequency range (band). The sub-band time domain signal can still retain the characteristics of the time domain signal with time as the horizontal axis and amplitude as the vertical axis, but only contains the signal of a certain frequency band in the original first time domain signal.
[0056] In this embodiment, the frequency band segmentation of the first time-domain signal can be achieved by using multiple dual second-order filters with different passband ranges to filter the first time-domain signal, resulting in multiple sub-band time-domain signals. Since the first time-domain signal contains the target sampling point, after filtering by multiple filters, each sub-band time-domain signal output by each filter contains the signal of the target sampling point in the corresponding frequency band; that is, each sub-band time-domain signal includes the signal corresponding to the target sampling point. The dual second-order filters can be digital filters based on second-order difference equations, featuring simple structure, low computational complexity, and ease of design and implementation. By configuring coefficients, low-pass, band-pass, or high-pass filtering functions can be achieved, making them suitable for extracting signals from specific frequency bands.
[0057] This application uses a dual second-order filter to perform frequency band segmentation on the first time-domain signal, obtaining multiple sub-band time-domain signals containing the signals corresponding to the target sampling points. Compared with the quadrature mirror (QMF) filter bank in related technologies, it significantly reduces the amount of computation, is easier to implement on a microprocessor, and is also easier to accurately divide the frequency bands with more noise distribution and the frequency bands where human voice is mainly distributed.
[0058] Step 103: Determine the first sub-band time domain signal of human voice and the second sub-band time domain signal of noise from the multiple sub-band time domain signals based on the intensity parameters corresponding to each sub-band time domain signal.
[0059] In this embodiment, the intensity parameter can be a physical quantity used to describe the strength of the sub-band time-domain signal. As an example, the intensity parameter can be the energy, power, or root mean square (RMS) level of the sub-band time-domain signal. The value of the intensity parameter can directly reflect the cumulative energy or average intensity of the signal in the sub-band time-domain signal. By analyzing the intensity parameter of each sub-band time-domain signal, it is possible to distinguish between sub-band time-domain signals that mainly include human voice and sub-band time-domain signals that mainly include noise. The first sub-band time-domain signal can be the human voice sub-band time-domain signal that mainly includes human voice, as determined by the intensity parameter; the second sub-band time-domain signal can be the noise sub-band time-domain signal that mainly includes noise, as determined by the intensity parameter.
[0060] In some embodiments, determining a second sub-band time-domain signal that is noise among multiple sub-band time-domain signals based on the intensity parameter corresponding to each sub-band time-domain signal includes:
[0061] Determine whether the value of the strength parameter is less than the preset strength threshold to obtain the first determination result;
[0062] If the value of the intensity parameter of the sub-band time domain signal, as determined by the first judgment result, is less than the intensity threshold, then the sub-band time domain signal is identified as the second sub-band time domain signal.
[0063] In this embodiment, the preset intensity threshold can be a critical value used to distinguish whether a sub-band time-domain signal is a human voice sub-band time-domain signal or a noise sub-band time-domain signal. As an example, the intensity threshold can be a pre-set critical value based on the analysis of a large amount of sample data, and can be dynamically adjusted according to the actual scenario. Before determining whether the value of the intensity parameter is less than the preset intensity threshold, the method further includes: calculating the intensity parameter of each sub-band time-domain signal. The process of calculating the intensity parameter can be determined according to the actual situation and is not limited here.
[0064] In this embodiment, the values of the intensity parameters of each sub-band time-domain signal can be compared with a preset intensity threshold to obtain a first judgment result; if the value of the intensity parameter of a certain sub-band time-domain signal is less than the intensity threshold, then the intensity of the sub-band time-domain signal is low, which meets the characteristics of a noise signal, and the sub-band time-domain signal is determined as the second sub-band time-domain signal.
[0065] In embodiments of the present invention, the second sub-band time-domain signal, which is identified as noise, is determined by comparing the intensity parameters of the sub-band time-domain signal with a preset intensity threshold. The difference in intensity characteristics between human voice and noise enables accurate identification of the noise sub-band time-domain signal, allowing for targeted localization of the noise sub-band time-domain signal. This improves the technical problem in related noise reduction methods where it is difficult to accurately distinguish between noise and human voice, resulting in insufficient noise reduction targeting or damage to effective speech, and enhances the accuracy of noise reduction processing.
[0066] In some embodiments, the method further includes:
[0067] Obtain the initial time-domain signal and initial energy parameters;
[0068] The average intensity parameters are determined based on the initial time-domain signal and initial energy parameters;
[0069] The strength threshold is determined based on the average strength parameter and the preset strength coefficient.
[0070] In this embodiment, the initial time-domain signal can be the time-domain signal acquired in the initial stage of the speech processing to be denoised. The initial time-domain signal can be a time-domain signal used for noise estimation and calculating the intensity threshold. Typically, the initial time-domain signal can be the first few frames of time-domain signal acquired when the audio acquisition device is first powered on. Since the noise signal contained in the first few frames of time-domain signal is relatively stable and does not contain human voice signals, it can be used as the basis for noise estimation. It should be noted that the initial time-domain signal can include the signal corresponding to the initial sampling points. The initial sampling points can be the target number of sampling points.
[0071] In this embodiment, obtaining the initial time-domain signal and initial energy parameters can be achieved by acquiring the initial time-domain signal of the first m frames of the speech to be denoised, where m is greater than or equal to 2; determining the initial energy parameters based on the sampled values of each initial sampling point, wherein the process of determining the initial energy parameters based on the initial sampled values can be determined according to the actual situation and is not limited here. Determining the average intensity parameter based on the initial time-domain signal and initial energy parameters can be achieved by calculating the arithmetic mean of the initial energy parameter values of the first m frames of the initial time-domain signal to obtain the average intensity parameter value, thereby realizing noise estimation.
[0072] In this embodiment, the preset intensity coefficient can be a coefficient pre-set based on the energy characteristics difference between human voice and noise. The preset intensity coefficient can be used to convert the average intensity parameter into an intensity threshold that can distinguish between human voice and noise. As an example, the preset intensity coefficient can be 2 or 3. Determining the intensity threshold based on the average intensity parameter and the preset intensity coefficient can be achieved by multiplying the average intensity parameter and the preset intensity coefficient as the intensity threshold.
[0073] In the embodiments of the present invention, the average intensity parameter is calculated by using the initial time domain signal and the initial energy parameter, and the intensity threshold is determined by combining the preset intensity coefficient. The intensity threshold of the current environment is determined based on the average intensity parameter of the initial time domain signal of the previous few frames, so that the intensity threshold is more in line with the actual noise environment, thereby improving the technical problem of low accuracy in distinguishing human voice from noise in different noise environments caused by the fixed intensity threshold in related technologies.
[0074] In some embodiments, after determining the second sub-band time-domain signal that is noise among the multiple sub-band time-domain signals based on the intensity parameter corresponding to each sub-band time-domain signal, the method further includes:
[0075] The intensity threshold is updated based on the intensity parameters of the second sub-band time-domain signal to obtain the updated intensity threshold.
[0076] In this embodiment, after determining the second sub-band time-domain signal, the arithmetic mean of the intensity parameter value and the intensity threshold of the second sub-band time-domain signal can be calculated to obtain the updated average intensity parameter. The intensity threshold is then re-determined based on the average intensity parameter and a preset intensity coefficient. Alternatively, the intensity parameter value and the intensity threshold of the second sub-band time-domain signal can be weighted to obtain the updated average intensity parameter. The intensity threshold is then re-determined based on the average intensity parameter and the preset intensity coefficient. It should be noted that when environmental noise increases, the intensity parameter value of the second sub-band time-domain signal increases, the updated average intensity parameter value increases, and consequently, the intensity threshold increases. Conversely, when environmental noise decreases, the intensity parameter value of the second sub-band time-domain signal decreases, the updated average intensity parameter value decreases, and the intensity threshold decreases.
[0077] In the embodiments of the present invention, by updating the intensity threshold based on the intensity parameter of the second sub-band time domain signal and dynamically adjusting the intensity threshold in combination with the current noise energy, the intensity threshold can be adapted to the changing environmental noise in real time. This can improve the technical problem in related technologies where the accuracy of distinguishing human voice from noise decreases when the noise environment changes due to the fixed intensity threshold.
[0078] Step 104: Perform noise reduction processing on the second sub-band time domain signal to obtain the third sub-band time domain signal.
[0079] In this embodiment, an appropriate noise reduction strategy can be adopted based on the noise characteristics of the second sub-band time-domain signal (e.g., noise intensity, bandwidth, frequency distribution, etc.). For example, different noise reduction coefficients can be selected based on the bandwidth parameter of the sub-band time-domain signal. A stronger attenuation coefficient can be used for wideband noise sub-band time-domain signals, while a weaker attenuation coefficient can be used for narrowband noise sub-band time-domain signals. Furthermore, the noise reduction level can be dynamically adjusted based on the noise intensity parameter, where a stronger noise reduction is applied when the noise intensity is high. Additionally, noise signals in the sub-band time-domain signal can be further filtered out using methods such as filtering. Through these filtering processes, the noise energy in the second sub-band time-domain signal can be suppressed, ultimately resulting in a noise-reduced third sub-band time-domain signal.
[0080] In some embodiments, the second sub-band time-domain signal is denoised to obtain the third sub-band time-domain signal, including:
[0081] Determine whether the value of the bandwidth parameter of the second sub-band time domain signal is greater than a preset bandwidth threshold to obtain a second determination result;
[0082] If the value of the bandwidth parameter of the second sub-band time domain signal, as indicated by the second judgment result, is greater than the bandwidth threshold, the second sub-band time domain signal is denoised according to the preset first denoising coefficient to obtain the third sub-band time domain signal.
[0083] In this embodiment, the bandwidth parameter can be a parameter used to describe the width of the frequency range covered by the second sub-band time-domain signal. The value of the bandwidth parameter can be the difference between the highest and lowest frequencies of the second sub-band time-domain signal. The bandwidth parameter can directly reflect the size of the frequency coverage range of the sub-band. The preset bandwidth threshold can be a critical value used to distinguish whether the second sub-band time-domain signal is a wideband sub-band time-domain signal or a narrowband sub-band time-domain signal.
[0084] In this embodiment, before determining whether the bandwidth parameter of the second sub-band time-domain signal is greater than a preset bandwidth threshold, the method further includes: calculating the difference between the highest and lowest frequencies of the second sub-band time-domain signal to obtain the value of the bandwidth parameter of the second sub-band time-domain signal. The first noise reduction coefficient can be a preset noise reduction attenuation coefficient for broadband noisy sub-band time-domain signals with bandwidth parameters greater than the bandwidth threshold. The value of the first noise reduction coefficient can be set according to the broadband noise suppression requirements and is not limited here. It should be noted that the first noise reduction coefficient can be different for different second sub-band time-domain signals.
[0085] In this embodiment, the bandwidth parameter values of each second sub-band time-domain signal can be compared with a preset bandwidth threshold to obtain a second judgment result. If the bandwidth parameter value of a certain second sub-band time-domain signal is greater than the bandwidth threshold, it indicates that the second sub-band time-domain signal is a broadband noisy sub-band time-domain signal, and the second sub-band time-domain signal is a sub-band time-domain signal with more noise distribution. The second sub-band time-domain signal is denoised using a preset first denoising coefficient. For example, the first denoising coefficient is multiplied by the amplitude parameter value of the second sub-band time-domain signal to obtain the amplitude parameter value of the third sub-band time-domain signal whose noise is effectively suppressed.
[0086] In the embodiments of the present invention, by determining the relationship between the bandwidth parameter of the second sub-band time domain signal and the preset bandwidth threshold, the first noise reduction coefficient is used to perform noise reduction processing on the noise sub-band time domain signal with a larger bandwidth parameter value. This can achieve targeted suppression of broadband sub-bands with more noise distribution, and can improve the technical problems in related technologies where noise reduction processing lacks differentiation in noise suppression of different bandwidth sub-bands, resulting in insufficient broadband noise suppression or excessive attenuation of narrowband voice signals, thereby improving the accuracy and effectiveness of noise reduction.
[0087] Step 105: Determine the second time-domain signal of the speech to be denoised after denoising based on the first sub-band time-domain signal and the third sub-band time-domain signal.
[0088] In this embodiment, the first sub-band time-domain signal and the denoised third sub-band time-domain signal can be integrated to form a complete signal covering the entire frequency range of the original speech to be denoised. Integration methods include, but are not limited to: performing in-band noise suppression on the first sub-band time-domain signal to obtain the denoised first sub-band time-domain signal; and merging the denoised first sub-band time-domain signal and the third sub-band time-domain signal according to their time-domain order. In some embodiments, for overlapping sections of the first and third sub-band time-domain signals, gain weighting can be used for smooth transition to avoid signal abrupt changes; for non-overlapping sections, the signal characteristics of each processed first and third sub-band time-domain signal can be directly retained, and finally, a second time-domain signal with continuous time and complete frequency is formed through time-domain superposition.
[0089] In some embodiments, determining the denoised second time-domain signal based on the first sub-band time-domain signal and the third sub-band time-domain signal includes:
[0090] The first sub-band time domain signal is denoised according to the preset second denoising coefficient to obtain the fourth sub-band time domain signal;
[0091] The time-domain signals of the third and fourth sub-bands are merged to obtain the second time-domain signal.
[0092] In this embodiment, the preset second noise reduction coefficient can be a noise reduction attenuation coefficient preset for the first sub-band time-domain signal. The second noise reduction coefficient can be the same as or different from the first noise reduction coefficient. The attenuation strength of the second noise reduction coefficient can be weaker than that of the first noise reduction coefficient. The function of the second noise reduction coefficient is to suppress noise in the first sub-band time-domain signal without damaging the main human voice. The first noise reduction coefficient and the second noise reduction coefficient can be the noise reduction depth of the sub-band time-domain signal.
[0093] In this embodiment, the first sub-band time-domain signal is denoised according to a preset second denoising coefficient to obtain the fourth sub-band time-domain signal. The second denoising coefficient is multiplied by the amplitude parameter of the first sub-band time-domain signal to reduce in-band noise while preserving human voice energy, resulting in the value of the amplitude parameter of the fourth sub-band time-domain signal. The third and fourth sub-band time-domain signals are then merged to obtain the second time-domain signal. The amplitude parameter values of the sampling points corresponding to the same time in the third and fourth sub-band time-domain signals are superimposed, causing signals from different frequency bands to recombine and form a second time-domain signal covering the full frequency range of the original signal. The second time-domain signal can be the denoised complete speech signal.
[0094] In an embodiment of the present invention, a fourth sub-band time-domain signal is obtained by suppressing in-band noise of the first sub-band time-domain signal according to the second noise reduction coefficient, and the fourth sub-band time-domain signal is combined with the third sub-band time-domain signal to obtain a second time-domain signal. This can effectively suppress noise in the human voice sub-band time-domain signal while retaining the main human voice signal. This can improve the technical problem in related technologies that only focus on noise sub-band noise reduction while ignoring noise in the human voice sub-band, resulting in the final signal still containing in-band noise, and improve the purity of the speech signal after noise reduction.
[0095] In some embodiments, the third sub-band time-domain signal and the fourth sub-band time-domain signal are combined to obtain a second time-domain signal, including:
[0096] Determine the overlap region between the time-domain signals of the third sub-band and the fourth sub-band;
[0097] The first signal value of the time-domain signal of the third sub-band in the overlapping interval is weighted based on the preset first gain parameter to obtain the first weighted value; and the second signal value of the time-domain signal of the fourth sub-band in the overlapping interval is weighted based on the preset second gain parameter to obtain the second weighted value; the sum of the values of the first gain parameter and the second gain parameter is a preset threshold.
[0098] The target signal value of the second time-domain signal is determined based on the first weighting value and the second weighting value.
[0099] In this embodiment, the overlapping interval can be the range of sampling points where the third sub-band time-domain signal and the fourth sub-band time-domain signal overlap in the time domain. The overlapping interval can be a region where both sub-band signals have valid signal values in the same time interval. The first gain parameter can be a preset weighting coefficient for the third sub-band time-domain signal within the overlapping interval. The value of the first gain parameter can dynamically change with the position of the sampling point within the overlapping interval. For example, the value of the first gain parameter can decrease as the position of the sampling point changes, and the value of the first gain parameter can be 1-0.2.
[0100] In this embodiment, the second gain parameter can be a preset weighting coefficient for the fourth sub-band time-domain signal within the overlapping interval. The value of the second gain parameter can dynamically change with the sampling point position. For example, the value of the second gain parameter can increase as the sampling point position changes. The value of the second gain parameter can be 0.2-1, and the sum of the second gain parameter and the first gain parameter is always equal to a preset threshold. The preset threshold can be determined according to actual conditions and is not limited here. As an example, the preset threshold can be 1. The first signal value can be the amplitude parameter value of the third sub-band time-domain signal within the overlapping interval at the corresponding sampling point; the second signal value can be the amplitude parameter value of the fourth sub-band time-domain signal within the overlapping interval at the corresponding sampling point.
[0101] In this embodiment, determining the overlap interval between the third sub-band time-domain signal and the fourth sub-band time-domain signal can be achieved by analyzing the time-domain distribution of the third and fourth sub-band time-domain signals to obtain their overlap interval on the time axis. A first gain parameter can be assigned to the sampling points of the third sub-band time-domain signal within the overlap interval, and a second gain parameter can be assigned to the sampling points of the fourth sub-band time-domain signal. A first weighted value is obtained by multiplying the first signal value of the third sub-band time-domain signal within the overlap interval by the corresponding first gain parameter, and a second weighted value is obtained by multiplying the second signal value of the fourth sub-band time-domain signal by the corresponding second gain parameter. The first weighted value and the second weighted value of each sampling point are added together to obtain the target signal value of that sampling point. The signal values of the corresponding sub-band time-domain signals in the non-overlapping intervals can be directly retained to form the second time-domain signal.
[0102] In the embodiments of the present invention, by determining the overlapping interval of the third sub-band time domain signal and the fourth sub-band time domain signal, and using complementary gain parameters for weighted merging, the signal values in the overlapping interval can be smoothly transitioned, and the sub-band time domain signals can be seamlessly merged. This improves the technical problem in related technologies where direct merging of sub-band signals leads to signal abrupt changes in the transition region, the introduction of ticking sounds or distortion, and enhances the auditory quality of the noise-reduced speech signal.
[0103] The following describes the speech noise reduction method provided in the embodiments of this application.
[0104] The purpose of this application is to provide a low-latency speech denoising method. This method involves noise estimation, sub-band segmentation, level detection, threshold determination, followed by denoising processing and noise updating. Using 10 sampling points and 48kHz audio, the latency is less than 1ms. This method achieves both low latency and can run on low-resource platforms, making it applicable to various audio scenarios. Figure 2 This is another schematic flowchart illustrating a speech noise reduction method provided in an embodiment of this application, as shown below. Figure 2 As shown, speech noise reduction methods include:
[0105] Start the real-time audio data stream. Initialize the noise estimate and calculate the initial threshold T. For example, the noise estimate is calculated using the average energy parameter of the previous few frames. The energy parameter for each frame is E = (1 / n) * sqrt(x1*x1 + x2*x2... + xn*xn), where x is the sample value within a frame; n is the frame length. The initial threshold T can be two or three times the noise estimate. Buffer 10 sample points as one frame of time-domain signal. For example, a sliding window technique can be used, with every 10 sample points as one frame of time-domain signal. For each frame of time-domain signal, data sample calculations are performed, for example, extracting statistics reflecting the characteristics of the current audio segment.
[0106] Subbanding is performed on the time-domain signal to obtain multiple subband time-domain signals. For example, the purpose of subbanding is to use digital filters to divide the time-domain signal into frequency bands, suppressing frequency bands with high noise levels and thus preserving the human voice frequency band. Common subbanding methods use low-pass, band-pass, and high-pass filters. Dual second-order filters can be used to divide the time-domain signal into multiple subbands such as 0-400Hz, 400-800Hz, 800-2000Hz, 2000-4000Hz, and 4000-8000Hz. The energy values of the multiple subband time-domain signals are then statistically analyzed. For example, the signal energy value can be calculated based on the signal level, typically using the RMS method. A noise reduction depth can be set for each subband time-domain signal.
[0107] The system determines whether the energy value of the sub-band time-domain signal is less than a threshold T. For example, in most scenarios, the energy of human voice is greater than that of noise. If the energy of human voice is less than that of noise, the human voice will be masked by the noise. Based on this criterion, an initial threshold T is set through data analysis such as noise estimation, and then dynamically adjusted. The threshold T is mainly used to distinguish between noise and human voice. If it is higher than the threshold, the sub-band time-domain signal is the human voice sub-band time-domain signal; if it is lower than the threshold, the sub-band time-domain signal is the noise sub-band time-domain signal.
[0108] If so, record the noise energy value to update the noise estimate. For example, when the current frame sub-band time-domain signal is determined to be a noisy sub-band time-domain signal, changing environmental noise can be dynamically identified. The noise estimate is updated by statistically analyzing the noise energy. Noise reduction is performed on the noisy sub-band time-domain signal based on the noise reduction depth. For example, based on the threshold discrimination result, the human voice sub-band time-domain signal and the noise sub-band time-domain signal are determined. If it is a noisy sub-band time-domain signal, noise reduction is performed on the frame sub-band time-domain signal using the set noise reduction coefficient.
[0109] If not, then in-band noise suppression is performed on the human voice sub-band time-domain signal based on the noise reduction depth. For example, noise is suppressed on the noise portion of the human voice sub-band time-domain signal using a set noise reduction coefficient. Speech post-processing. For example, the speech segment is smoothed using fade-in / fade-out techniques to avoid audio data loss or the introduction of ticking sounds due to swallowed sounds or discontinuities at the end.
[0110] This application provides a low-latency noise reduction scheme architecture. Starting from a 48kHz audio data stream, the level threshold and noise estimate are initialized, and 10 sampling points (0.2ms delay) are buffered. A bandpass filter is used for sub-band segmentation, and the average electrical value within the current frequency band is calculated. Combining sliding window technology and threshold judgment, it is determined whether human voice is present. If present, the human voice time-domain signal is post-processed and filtered to suppress in-band noise; if absent, the noise level value is calculated, and the level threshold and noise estimate are updated accordingly, while noise reduction processing is performed. The signal is then subjected to speech post-processing filtering.
[0111] Figure 3 This is a schematic diagram of the structure of a speech noise reduction device provided in an embodiment of this application, as shown below. Figure 3 As shown, the voice noise reduction device 300 includes:
[0112] The first acquisition module 301 is used to acquire the first time-domain signal including the target sampling point in the speech to be denoised;
[0113] The segmentation module 302 is used to perform frequency band segmentation processing on the first time domain signal to obtain multiple sub-band time domain signals; each sub-band time domain signal includes the signal corresponding to the target sampling point.
[0114] The first determining module 303 is used to determine, based on the intensity parameters corresponding to each sub-band time domain signal, the first sub-band time domain signal of human voice and the second sub-band time domain signal of noise among multiple sub-band time domain signals;
[0115] The first noise reduction module 304 is used to perform noise reduction processing on the second sub-band time domain signal to obtain the third sub-band time domain signal;
[0116] The second determining module 305 is used to determine the second time-domain signal after denoising of the speech to be denoised based on the first sub-band time-domain signal and the third sub-band time-domain signal.
[0117] In some embodiments, the first determining module 303 is further configured to determine whether the value of the intensity parameter is less than a preset intensity threshold and obtain a first determination result; if the first determination result indicates that the value of the intensity parameter of the sub-band time domain signal is less than the intensity threshold, then the sub-band time domain signal is determined as the second sub-band time domain signal.
[0118] In some embodiments, the first noise reduction module 304 is further configured to determine whether the value of the bandwidth parameter of the second sub-band time domain signal is greater than a preset bandwidth threshold, and obtain a second determination result; if the second determination result indicates that the value of the bandwidth parameter of the second sub-band time domain signal is greater than the bandwidth threshold, the second sub-band time domain signal is denoised according to a preset first noise reduction coefficient to obtain a third sub-band time domain signal.
[0119] In some embodiments, the speech noise reduction device 300 further includes: a second acquisition module for acquiring an initial time-domain signal and an initial energy parameter; a third determination module for determining an average intensity parameter based on the initial time-domain signal and the initial energy parameter; and a fourth determination module for determining an intensity threshold based on the average intensity parameter and a preset intensity coefficient.
[0120] In some embodiments, the speech noise reduction device 300 further includes an update module for updating the intensity threshold based on the intensity parameters of the second sub-band time-domain signal to obtain an updated intensity threshold.
[0121] In some embodiments, the second determining module 305 is further configured to perform noise reduction processing on the first sub-band time domain signal according to a preset second noise reduction coefficient to obtain a fourth sub-band time domain signal; and to perform merging processing on the third sub-band time domain signal and the fourth sub-band time domain signal to obtain a second time domain signal.
[0122] In some embodiments, the second determining module 305 is further configured to: determine the overlapping interval of the third sub-band time-domain signal and the fourth sub-band time-domain signal; weight the first signal value of the third sub-band time-domain signal in the overlapping interval based on a preset first gain parameter to obtain a first weighted value; and weight the second signal value of the fourth sub-band time-domain signal in the overlapping interval based on a preset second gain parameter to obtain a second weighted value; the sum of the values of the first gain parameter and the second gain parameter is a preset threshold; and determine the target signal value of the second time-domain signal based on the first weighted value and the second weighted value.
[0123] To implement the method of the embodiments of this application, Figure 4 This is a schematic diagram of the hardware structure of a speech noise reduction device provided in an embodiment of this application, as shown below. Figure 4 As shown in the illustration, this application embodiment also provides a speech noise reduction device 40, which may include: a memory 401 for storing a computer program; and a processor 402 for implementing the method described above when executing the computer program. For example, the processor 402 may be used to implement the steps in any of the methods described above, which will not be elaborated further here.
[0124] It should be noted that the speech noise reduction device and the speech noise reduction method provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0125] Of course, in practical applications, such as Figure 4As shown, the speech noise reduction device 40 may further include at least one network interface 403. The various components in the speech noise reduction device are coupled together via a bus system 404. It is understood that the bus system 404 is used to enable communication between these components. In addition to a data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 4 Various buses are labeled as bus system 404. The number of processors 402 can be at least one. Network interface 403 is used for wired or wireless communication between the voice noise reduction device and other devices. Memory 401 in this embodiment is used to store various types of data to support the operation of the voice noise reduction device. The methods disclosed in the above embodiments can be applied to processor 402, or implemented by processor 402. Processor 402 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 402 or by instructions in software form. The processor 402 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 402 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly reflected in the combined execution of hardware and software modules in a microcontroller. The software module can reside in a storage medium located in memory 401. Processor 402 reads information from memory 401 and, in conjunction with its hardware, completes the steps of the aforementioned method. In an exemplary embodiment, the voice noise reduction device 40 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to execute the aforementioned method.
[0126] Specifically, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, such as a memory 401 storing the computer program, which can be executed by a processor 402 to complete the aforementioned method steps. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0127] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0128] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0129] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0130] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0131] The above provides a detailed description of a speech noise reduction method, apparatus, electronic device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A speech noise reduction method, characterized in that, include: Acquire the first time-domain signal of the speech to be denoised, which includes the target sampling point; The first time-domain signal is subjected to frequency band segmentation to obtain multiple sub-band time-domain signals; each sub-band time-domain signal includes the signal corresponding to the target sampling point. Based on the intensity parameter corresponding to each of the sub-band time-domain signals, determine the first sub-band time-domain signal that is human voice and the second sub-band time-domain signal that is noise from among the multiple sub-band time-domain signals; The second sub-band time-domain signal is denoised to obtain the third sub-band time-domain signal. The second time-domain signal after denoising of the speech to be denoised is determined based on the first sub-band time-domain signal and the third sub-band time-domain signal. The step of determining the second time-domain signal after denoising the speech to be denoised based on the first sub-band time-domain signal and the third sub-band time-domain signal includes: The first sub-band time domain signal is denoised according to the preset second denoising coefficient to obtain the fourth sub-band time domain signal; The third sub-band time-domain signal and the fourth sub-band time-domain signal are merged to obtain the second time-domain signal; The step of merging the third sub-band time-domain signal and the fourth sub-band time-domain signal to obtain the second time-domain signal includes: Determine the overlap interval between the third sub-band time-domain signal and the fourth sub-band time-domain signal; The first signal value of the time-domain signal of the third sub-band in the overlapping interval is weighted based on a preset first gain parameter to obtain a first weighted value; and the second signal value of the time-domain signal of the fourth sub-band in the overlapping interval is weighted based on a preset second gain parameter to obtain a second weighted value; the sum of the values of the first gain parameter and the second gain parameter is a preset threshold. The target signal value of the second time-domain signal is determined based on the first weighted value and the second weighted value.
2. The method according to claim 1, characterized in that, Determining a second sub-band time-domain signal that is noise among the multiple sub-band time-domain signals based on the intensity parameter corresponding to each sub-band time-domain signal includes: Determine whether the value of the strength parameter is less than a preset strength threshold to obtain a first determination result; If the first determination result indicates that the value of the intensity parameter of the sub-band time domain signal is less than the intensity threshold, then the sub-band time domain signal is determined as the second sub-band time domain signal.
3. The method according to claim 1, characterized in that, The step of denoising the second sub-band time-domain signal to obtain the third sub-band time-domain signal includes: Determine whether the value of the bandwidth parameter of the second sub-band time domain signal is greater than a preset bandwidth threshold to obtain a second determination result; If the second judgment result indicates that the value of the bandwidth parameter of the second sub-band time domain signal is greater than the bandwidth threshold, the second sub-band time domain signal is denoised according to the preset first denoising coefficient to obtain the third sub-band time domain signal.
4. The method according to claim 2, characterized in that, The method further includes: Obtain the initial time-domain signal and initial energy parameters; The average intensity parameter is determined based on the initial time-domain signal and the initial energy parameter; The intensity threshold is determined based on the average intensity parameter and the preset intensity coefficient.
5. The method according to claim 2, characterized in that, After determining the second sub-band time-domain signal that is noise among the multiple sub-band time-domain signals based on the intensity parameter corresponding to each sub-band time-domain signal, the method further includes: The intensity threshold is updated based on the intensity parameters of the second sub-band time-domain signal to obtain the updated intensity threshold.
6. A voice noise reduction device, characterized in that, include: The first acquisition module is used to acquire the first time-domain signal in the speech to be denoised, which includes the target sampling point; The segmentation module is used to perform frequency band segmentation processing on the first time domain signal to obtain multiple sub-band time domain signals; each sub-band time domain signal includes the signal corresponding to the target sampling point; The first determining module is used to determine, based on the intensity parameter corresponding to each of the sub-band time-domain signals, a first sub-band time-domain signal that is human voice and a second sub-band time-domain signal that is noise from among the multiple sub-band time-domain signals. The first noise reduction module is used to perform noise reduction processing on the second sub-band time domain signal to obtain the third sub-band time domain signal; The second determining module is used to determine the second time-domain signal of the speech to be denoised after denoising based on the first sub-band time-domain signal and the third sub-band time-domain signal. The second determining module is further configured to perform noise reduction processing on the first sub-band time-domain signal according to a preset second noise reduction coefficient to obtain a fourth sub-band time-domain signal; and to perform merging processing on the third sub-band time-domain signal and the fourth sub-band time-domain signal to obtain the second time-domain signal. The second determining module is further configured to determine the overlap interval between the third sub-band time-domain signal and the fourth sub-band time-domain signal; The first signal value of the time-domain signal of the third sub-band in the overlapping interval is weighted based on a preset first gain parameter to obtain a first weighted value; and the second signal value of the time-domain signal of the fourth sub-band in the overlapping interval is weighted based on a preset second gain parameter to obtain a second weighted value; the sum of the values of the first gain parameter and the second gain parameter is a preset threshold; and the target signal value of the second time-domain signal is determined according to the first weighted value and the second weighted value.
7. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Voice gain control method and computer storage medium
CN112242147A
Speech processing apparatus
JP2019086724A