Method and system for background noise estimation

By dividing the audio signal buffer and constructing a cost function, combined with confidence-adjusted noise reduction, the accuracy problem in background noise estimation is solved, and robust noise reduction for all audio signals is achieved.

CN114981888BActive Publication Date: 2026-03-24DOLBY INTERNATIONAL AB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-18
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies fail to accurately identify audio segments containing only noise in background noise estimation, causing noise reduction methods to fail, especially when the signal exists at different times and frequencies. Furthermore, existing methods may discard narrowband tone components or make assumptions about the nature of the signal, making them unsuitable for all types of audio signals.

Method used

By dividing the audio signal into multiple buffers, calculating the time-frequency samples and energy changes in each buffer, and combining the median and standard deviation, a cost function is constructed to determine the background noise. The noise reduction process is then adjusted using confidence values ​​to avoid making assumptions about the signal properties.

Benefits of technology

It achieves accurate estimation of background noise in various audio signals, avoids discarding narrowband tone components, and remains robust during signal fade-in and fade-out, applicable to all types of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114981888B_ABST
    Figure CN114981888B_ABST
Patent Text Reader

Abstract

Background noise estimation and noise reduction are disclosed, in an embodiment, a method includes: obtaining an audio signal; dividing the audio signal into a plurality of buffers; determining time-frequency samples of each buffer of the audio signal; determining, for each buffer and each frequency, a median (or mean) of energies and a measure of energy variation based on samples in the buffer and samples in adjacent buffers, the samples in the buffer and the samples in the adjacent buffers together spanning a specified time range of the audio signal; combining the median (or mean) of energies and the measure of energy variation into a cost function; for each frequency: determining a signal energy of a particular buffer of the audio signal corresponding to a minimum of the cost function; selecting the signal energy as an estimated background noise of the audio signal; and using the estimated background noise to reduce noise in the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to the following prior applications: Spanish application P202030040 (reference number: D19149ES), filed January 21, 2020; U.S. Provisional Application 63 / 000,223 (reference number: D19149USP1), filed March 26, 2020; and U.S. Provisional Application 63 / 117,313 (reference number: D19149USP2), filed November 23, 2020, which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to audio signal processing. Background Technology

[0004] Unlike professional settings, background noise is a potential problem in user-generated audio content (UGC) due to limitations of the equipment used and the uncontrolled acoustic environment in which recording takes place. Such background noise, besides being bothersome, can be amplified by processing tools that apply heavy dynamic range compression and equalization to the audio content. Therefore, noise reduction is a crucial element in the audio processing chain to reduce background noise. Noise reduction relies on successfully measuring the noise floor, which can be obtained by analyzing the power spectrum of a recording segment containing only background noise. Such segments can be manually identified by the user, found automatically, or obtained by requiring the performer / speaker to remain silent for the first few seconds of recording. However, there are still scenarios where it is impossible to obtain audio content segments containing only noise.

[0005] Existing methods based on finding quiet segments of audio (manually or automatically) fail in cases where such segments don't exist, for example, because the signal exists at different times and frequencies. Other methods are based on fitting the audio spectrum to a smooth curve that passes through a minimum. Such methods often discard narrow-band tonal components of noise, such as electrical hum. Other methods, based on calculating the level distribution at each frequency and selecting a low percentage (e.g., 10%) of the distribution as noise, are not robust to, for example, fade-in and fade-out signals. Finally, other methods rely on assumptions about the nature of the signal (e.g., assuming the signal is speech) and therefore cannot be generalized to all types of audio signals. Summary of the Invention

[0006] Implementation methods for background noise estimation and noise reduction are disclosed.

[0007] In one embodiment, a method includes: obtaining an audio signal; dividing the audio signal into a plurality of buffers; determining time-frequency samples of the audio signal for each buffer; for each buffer and each frequency, determining a measure and median of energy variation based on samples in the buffer and samples in adjacent buffers, the samples in the buffer and samples in adjacent buffers together spanning a specified time range of the audio signal; combining the median and the measure of energy variation into a cost function; for each frequency: determining a signal energy of a specific buffer of the audio signal corresponding to the minimum value of the cost function; selecting the signal energy as an estimated noise floor of the audio signal; and using the estimated noise floor to reduce noise in the audio signal.

[0008] In this embodiment, the mean is determined instead of the median.

[0009] In the embodiment, the measure and median or mean of the change are scaled to between 0.0 and 1.0.

[0010] In an embodiment, the combination of the change with the mean or median is the sum of their values ​​plus the reciprocal of the sum of their product and 1.

[0011] In an embodiment, the combination of the change and the median or mean is the sum of their squared values.

[0012] In an embodiment, the combination of the change with the median or mean is the sum of the square of the median or mean and the sigmoid of the energy variance.

[0013] In an embodiment, the combination of the change with the median or mean is the sum of the median or mean and the sigmoid of the variance.

[0014] In this embodiment, the change is replaced by the difference between the maximum energy value and the minimum energy value in the buffer spanning a specified time range.

[0015] In an embodiment, the buffer having a variance and median or mean calculated for blocks of the audio signal includes at least one buffer where the overall signal energy is below a predefined threshold, and the at least one buffer is not used to estimate the noise floor of the audio signal.

[0016] In this embodiment, the predefined threshold is determined relative to the maximum level of the audio signal.

[0017] In this embodiment, the predefined threshold is determined relative to the average level of the audio signal.

[0018] In an embodiment, the method further includes: analyzing the distribution of blocks of the audio signal using one or more processors, estimating the noise floor at each frequency based on the distribution; selecting a block k and a frequency f; and replacing the estimated noise at frequency f with the value calculated from block k if the increased cost is less than a second predefined threshold.

[0019] In an embodiment, the method further includes determining a confidence value based on the value of the energy change at the selected buffer.

[0020] In this embodiment, the confidence values ​​are smoothed over frequency.

[0021] In an embodiment, reducing noise in the audio signal further includes applying gain reduction at each frequency, the gain reduction decreasing as the confidence value at the frequency decreases.

[0022] In an embodiment, the method further includes: selecting a frequency f1 using one or more processors; calculating, using one or more processors, the average of the discrete derivatives of the spectrum in segments of a predefined size for all intervals of a predetermined size above the selected frequency f1; and selecting, using one or more processors, a segment with the negative derivative as the cutoff frequency f when the maximum negative derivative is less than a predefined value. c ; and using one or more processors to replace spectral values ​​above the cutoff frequency with the average value of the spectrum in a frequency band having an upper boundary adjacent to the cutoff frequency.

[0023] In an embodiment, the cost function increases with the increase of the median or mean, and also increases with the increase of the measure of the energy change.

[0024] In this embodiment, the cost function is non-linear.

[0025] In an embodiment, the cost function is symmetric on both the measure of energy change and the mean or median.

[0026] In this embodiment, the cost function is asymmetric, and when the measure of the energy change is less than a predefined threshold, the weight of the measure of the energy change is less than the weight of the mean or median.

[0027] In one embodiment, a system includes: one or more processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the operations described in the foregoing methods.

[0028] In one embodiment, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform any of the operations described in the foregoing methods.

[0029] Other embodiments disclosed herein relate to systems, apparatuses, and computer-readable media. Details of the disclosed embodiments are set forth in the accompanying drawings and description below. Other features, objects, and advantages will be apparent from this specification, the drawings, and the claims.

[0030] The specific embodiments disclosed herein offer one or more of the following advantages. The disclosed systems and methods can be used to estimate the noise floor when a reliable estimate of the audio signal's noise floor is unavailable (e.g., in segments containing only background noise). Unlike existing solutions, the disclosed systems and methods do not discard narrowband tonal components of the audio signal (e.g., electrical hum) and are robust to, for example, fade-in and fade-out of the audio signal. Furthermore, no assumptions need to be made about the properties of the audio signal, allowing the disclosed systems and methods to be applied to all types of audio signals. Attached Figure Description

[0031] In the accompanying drawings, for ease of description, specific arrangements or orders of schematic elements are shown, such as those representing devices, units, instruction blocks, and data elements. However, those skilled in the art will understand that the specific order or arrangement of the schematic elements in the drawings does not imply a requirement for a particular processing sequence or order, or processing separation. Furthermore, the inclusion of schematic elements in the drawings does not mean that such elements are required in all embodiments, or that in some embodiments, the features represented by such elements may not be included in other elements or combined with other elements.

[0032] Furthermore, in drawings that use connecting elements such as solid or dashed lines or arrows to illustrate connections, relationships, or associations between two or more other schematic elements, the absence of any such connecting element does not imply the absence of connections, relationships, or associations. In other words, some connections, relationships, or associations between elements are not shown in the drawings to avoid obscuring this disclosure. Additionally, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, in cases where a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or more signal paths that may be required to affect the communication.

[0033] Figure 1 This is a block diagram of a system for background noise estimation and noise reduction according to an embodiment.

[0034] Figures 2A to 2CThe diagram (from top to bottom) illustrates the signal energy, median (μ), and standard deviation (σ) at a certain frequency buffer according to an embodiment.

[0035] Figure 3 The cost functions of μ and σ according to an embodiment are illustrated.

[0036] Figure 4A The illustration shows an example energy level for each buffer i at a given frequency f according to an embodiment, highlighting the buffers corresponding to the minimum cost function J(i, f).

[0037] Figure 4B The illustration shows an embodiment of the target Figure 4A The sample values ​​(μ) of buffer i and frequency f in dB.

[0038] Figure 4C The illustration shows an embodiment of the target Figure 4A Example standard deviation (σ) of buffer i and frequency f in dB.

[0039] Figure 4D The illustration shows an example minimum of the cost function J(i, f) for buffer i and frequency f according to an embodiment, and highlights the value of argmin. i The buffer corresponding to {J(i, f)}.

[0040] Figure 5A The illustration shows an example of noise level (dB) as a function of frequency f according to an embodiment.

[0041] Figure 5B The illustration shows an example standard deviation of the estimated noise according to an embodiment, where at each frequency f, the standard deviation corresponds to a buffer with the lowest cost function at a given frequency.

[0042] Figure 5C The following is illustrated based on an embodiment. Figure 5B The standard deviation σ shown in the figure Figure 5A The confidence level of the noise estimation.

[0043] Figure 6 The figure illustrates the gain curve (transfer function) for noise reduction according to an embodiment.

[0044] Figure 7A The illustration shows a significant reduction in noise floor at high frequencies according to the embodiment.

[0045] Figure 7B The illustration shows the embodiment of the present invention. Figure 7AThe noise spectrum above frequency f1 shown is divided into segments with a length of L points and predefined overlap, and the average derivative of the points in each segment is calculated and sorted in ascending order of the frequency of the corresponding segment.

[0046] Figure 7C The illustration shows the finding of a first average derivative with a value greater than a predefined negative value according to an embodiment.

[0047] Figure 7D The figure illustrates the calculated cutoff frequency f according to an embodiment. c The average noise spectrum in the previous small region is used to replace the values ​​above f. c The value of the noise spectrum.

[0048] Figure 8 This is a flowchart of a process for background noise estimation and noise reduction according to an embodiment.

[0049] Figure 9 An implementation reference according to an embodiment is shown. Figures 1 to 8 A block diagram of an example system describing the features and processes.

[0050] The same reference numerals used in all the figures indicate the same elements. Detailed Implementation

[0051] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the various embodiments described. It will be apparent to those skilled in the art that different implementations can be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Several features are described below, each of which can be used independently of each other or in any combination with other features.

[0052] Nomenclature

[0053] As used herein, the term “comprising” and variations thereof should be understood as open-ended terms meaning “including but not limited to”. Unless the context explicitly states otherwise, the term “or” should be understood as “and / or”. The term “based on” should be understood as “at least partially based on”. The terms “one example implementation” and “example implementation” should be understood as “at least one example implementation”. The term “another implementation” should be understood as “at least one other implementation”. The term “determined, determines, or determining” should be understood as obtaining, receiving, calculating, estimating, predicting, or acquiring. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0054] System Overview

[0055] The disclosed embodiments identify segments of an audio recording at each frequency of an audio signal (e.g., an audio file or audio stream) where the energy is lower than that of other segments of the audio recording, and where the variance of the energy is relatively small. The energy of such a segment at a frequency of interest is considered the stable noise level at that frequency. At each frequency, selecting a suitable segment is considered a minimization problem, where a preference is given to segments with low energy and low variance, thus finding the optimal trade-off between the two independent variables. If, at a certain frequency, the level identified as noise floor corresponds to a relatively high variance, then the confidence associated with such frequency is low. The confidence value is used to inform subsequent noise reduction units, so the gain attenuation applied to suppress noise is reduced according to the confidence value, thereby allowing a conservative approach where potentially inaccurate noise estimates do not negatively impact the output quality of the noise reduction. In cases where the noise floor drops significantly at higher frequencies (e.g., typically due to bandwidth limitations in a lossy codec), the estimated noise value before reduction is maintained until the end of the spectrum to avoid a reduction in attenuation gain due to the smoothing of the attenuation gain at frequencies around the reduction region.

[0056] Figure 1 This is a block diagram of a system 100 for background noise estimation and noise reduction according to an embodiment. The system 100 includes a spectrum generation unit 101, a buffer 102, a root mean square (RMS) calculator 103, a statistical analysis unit 104 (“STATS”), a cost function unit 105, an optional smoothing unit 106, a noise reduction unit 107, and a partitioning unit 108.

[0057] In an embodiment, the input audio signal x(t) (e.g., an audio file or audio stream) is divided by partitioning unit 108 into a plurality of buffers 102, each buffer comprising N samples (e.g., 4096 samples) at a Z kHz sampling rate (e.g., 48 kHz) and overlapping with adjacent buffers by a percentage Y (e.g., 50% overlap). Spectrum generation unit 101 applies a frequency transform to the contents of the plurality of buffers 102 to obtain a time-frequency representation X(n,f), which comprises buffers with M frequency intervals (e.g., 4096 samples) at a Z kHz sampling rate (e.g., 48 kHz). For example, 4096 samples, 50% overlap, and a sampling rate of 48 kHz result in a frequency resolution of approximately 12 Hz for each buffer. In some embodiments, the frequency transform is a short-time Fourier transform (STFT), which outputs time-frequency data (e.g., a time-frequency slice).

[0058] For each buffer i, the RMS calculator 103 calculates the buffer's RMS level in the time domain and defines a mute threshold relative to the maximum RMS (e.g., -80 dB below the maximum RMS). The mute threshold is calculated by analyzing the entire audio signal and is therefore limited to "offline" use cases. Alternatively, the mute threshold can be defined as a fixed number (e.g., -100 dBFS) or a fixed number depending on the bit depth of the input audio file / stream (e.g., -90 dBFS for a 16-bit signal and -140 dBFS for a 24-bit signal). Mute buffers are those buffers with an RMS level below the mute threshold.

[0059] For each frequency f and each buffer i, the statistical analysis unit 104 calculates the median and a measure of the variation of energy in the samples of the j buffers (e.g., standard deviation, variance, range (maximum-minimum), interquartile range), where the j buffers belong to a block (e.g., a 1-second audio segment) of the audio signal x(t) centered at buffer i. Equations [1] and [2] describe the operation of the statistical analysis unit 104 using the median μ and standard deviation σ of the energy in the samples of the j buffers, as follows:

[0060] μ(i, f) = median(20 * Log(|X) i (j, f)|)), [1]

[0061] σ(i, f) = std(20 * Log(|X) i (j, f)|)). [2]

[0062] Audio signal blocks containing one or more silence buffers (determined by a silence threshold) are not used to calculate the median and standard deviation. In some embodiments, the median can be replaced with the mean to reduce computational cost.

[0063] Figures 2A to 2C The diagram (from top to bottom) illustrates the signal energy, median μ, and standard deviation σ at a certain frequency buffer according to an embodiment. The goal is to find the audio signal block at each frequency that best represents the background noise of the audio signal, i.e., the block with a small median / mean μ and standard deviation σ. Instead of introducing a threshold, cost function unit 105 calculates the numerical joint minimization of the cost function J(μ(i,f), σ(i,f)) after rescaling μ and σ to be in the interval [0.0, 1.0], i.e., normalizing them.

[0064]

[0065] Once the relationship with argmin is determined i The buffer k(f) corresponding to {J(i, f)} makes the background noise of the audio file / stream equal to the median / mean of buffer k:

[0066] noise (f) =μ(k(f), f) [4]

[0067] The audio block corresponding to buffer k includes some adjacent buffers of buffer k, referred to as the selected block at frequency f. Figure 3 The cost functions of μ and σ according to equation [3] are illustrated.

[0068] Note that posterior rescaling of μ and σ requires obtaining their values ​​over the entire audio file. If noise estimation is to be performed online, and the file is being recorded or processed, a fixed range of [μ] for these two variables can be introduced based on prior empirical observations. max μ min ] and [σ max , σ min This will rescale the variable, making it the same as before.

[0069] μ(i, f) = 0, if μ(i, f) ≤ μ min [5]

[0070] μ(i,f)=(μ(i,f)-μ min ) / (μ max -μ min If μ min <μ(i,f)<μ max [6]

[0071] μ(i, f) = 1, if μ(i, f) ≥ μ min [7]

[0072] σ can be rescaled in a similar manner using equations [5] through [7] and by replacing μ with σ.

[0073] In some embodiments, the following changes to the cost function are considered (while still assuming that μ and σ are a posteriori or online based on their maximum and minimum values, or on guessed maximum and minimum values). The cost function can be expressed in quadratic terms:

[0074] J(i, f) = μ 2 (i, f) + σ 2 (i, f). [8]

[0075] The roles and importance of μ and σ can be altered, thereby breaking the symmetry of the cost function. One approach is to transform σ such that it gives a lower cost below a certain threshold, a higher cost above that threshold, and a smooth transition between the two. This formula minimizes J(i, f) for smaller values ​​of σ. One possible implementation is to use the sigmoid function shown in equation [9]:

[0076]

[0077] Here, α = 10 is a good example of a scaling factor for the sigmoid function.

[0078] In some embodiments, the quadratic term μ 2 (i, f) can be replaced by the linear term μ(i, f) to give less weight to blocks with lower levels, thereby avoiding potential underestimation.

[0079] Advantageously, there is a preference for noise estimation based on adjacent frequencies selected from the same audio block to avoid accidental underestimation of outliers in the noise curve, which is otherwise very smooth. One embodiment of this is achieved by examining the frequency distribution of a selected block k(f), for example, by visualizing a histogram of the selected block's location in the audio file. If in a certain block If large clusters are found and few random outliers are observed, it can be assumed that block k is mainly background noise, and this can force the estimation of outlier frequencies within the same block. For the corresponding block... The frequency can be used to calculate costs. And if the cost increase is less than a certain threshold: Replace noise(f) = μ(k, f) with The slight variation of this rule is that as long as the cost difference is less than J Th Choose to The surrounding n k The noise estimate corresponding to the minimum cost within each buffer range.

[0080] Figure 4AThe illustration shows an example of the noise level corresponding to the minimum of the cost function J(i, f) for a given buffer i and frequency f. Figure 4B The figure shows an example mid-to-mean (μ) value in dB for buffer i and frequency f. Figure 4C The figure shows an example standard deviation (σ) in dB for buffer i and frequency f. Figure 4D The diagram illustrates an example cost function for buffer i and frequency f, and the buffer argmin that reaches its minimum value. i {J(i, f}).

[0081] In an embodiment, an optional smoothing unit 106 applies smoothing to the estimated noise floor to avoid fluctuations caused by estimating adjacent intervals from different blocks of the audio signal. The smoothing unit 106 replaces each value of noise(f) with the average value of values ​​in a frequency band near f. This frequency band can be rectangular, triangular, etc. In some embodiments, a smoothing function that reaches zero at the band boundaries can be used. For perceptual reasons, the width of the frequency band is exponential and corresponds to a constant fraction of octaves. In some embodiments, the constant fraction is 1 / 100, which is a very narrow bandwidth used to maintain sufficient resolution to accurately measure the noise components.

[0082] By associating small confidence levels with frequencies having high variance, or vice versa, a confidence value c(f) representing the reliability of the estimate can be obtained based on the value of σ(k):

[0083] c(f) = 0, if σ ≥ σ H

[10]

[0084] If σ L <σ<σ H

[11]

[0085] c(f) = 1, if σ ≤ σ L

[12]

[0086] The example value determined based on experience is σ. H =14 and σ L =7.5. The confidence level can be used to inform the noise reduction unit 107 about the accuracy of the background noise estimation, thus improving noise reduction and avoiding unwanted artifacts at frequencies where the estimation is considered inaccurate.

[0087] Figure 5A The figure illustrates an example of estimated noise level (dB) as a function of frequency f. Figure 5B The diagram shows... Figure 5A The example standard deviation of the estimated noise shown is the standard deviation of the buffer where the cost function has the lowest value at a given frequency f. Figure 5C It shows the basis Figure 5B The standard deviation σ shown in the figure Figure 5A The confidence level of the noise estimate. Note that, according to equation

[12] , when σ is less than σ L When the confidence level is 1; according to equation

[11] , when σ is in σ L With σ H When the confidence level is between these values, it is given by the following formula: And according to equation

[10] , when σ is greater than σ H When the confidence level is 0.

[0088] In embodiments, the noise reduction unit 107 is a band-based or FFT-based extender. On any given frame, frequency ranges where the energy approaches the estimated noise floor are attenuated, with the attenuation gain proportional to the proximity of the energy to the noise floor. In some embodiments, the gain attenuation G(n,f) is used by L(n,f) as described below. Figure 6 The curve shown in the figure is similar to the curve in the figure.

[0089] Specifically, let N(f) be the energy level of the noise in dB, and let S(n, f) be the energy level of the audio content at frame n and frequency f. In some embodiments, a threshold Th in decibels is defined, and the level above the threshold is calculated as:

[0090] L(n,f)=10Log(S(n,f))-(N(f)+Th).

[13]

[0091] refer to Figure 6 Figure 601 (also known as the "noise reduction curve") and bypass curve 602 are shown. At a given input level (dB), the gain reduction is the difference between the input level (x-axis) and the desired output level (dB) (y-axis). The gain curve 601 has a slope of 1 above a threshold 603, and a slope below the threshold point 603 corresponding to a selected ratio (e.g., typically 5 or greater), and transitions smoothly or abruptly around the threshold point 603. When a confidence level c(f) is provided by the cost function unit 106, the noise reduction unit 107 uses said confidence level to attenuate the noise reduction effect at frequencies with lower confidence levels by scaling the gain reduction in decibels using the confidence level:

[0092] G(i,f)=c(f)G(i,f).

[14]

[0093] In some embodiments, the confidence level can also be smoothed by the smoothing unit 105 to ensure a continuous transition between full noise reduction in the high-confidence frequency band and zero noise reduction in the low-confidence frequency band.

[0094] like Figure 7A As shown, in cases where the noise floor drops significantly at high frequencies (e.g., often due to bandwidth limitations in lossy codecs), the estimated noise value before reduction will remain until the end of the spectrum. This is to avoid a reduction in attenuation gain due to the smoothing of the attenuation gain at frequencies around the reduction region.

[0095] In some embodiments, the reduced frequency is determined by: 1) selecting a first frequency f1, and estimating the cutoff frequency f at a frequency higher than the first frequency f1. c ,like Figure 7A As shown; 2) Divide the noise spectrum above f1 into segments with a length of L points and a predefined overlap (e.g., 50%), such as Figure 7B As shown; 3) and, within each segment, calculate the average derivative, sort by the increasing frequency of its corresponding block, and find the first derivative with a value less than a predefined negative value (e.g., -20dB), such as Figure 7C As shown; and 4) calculate f c The average value n of the noise spectrum in the previous small area c and will be higher than f c Replace the value of the noise spectrum with n c ,like Figure 7D As shown. Note that step (3) should be interpreted as a significant decrease in the frequency spectrum, and the frequency of the corresponding segment is considered to be the cutoff frequency f. c

[0096] Example process

[0097] Figure 8 This is a flowchart of a process 800 for background noise estimation and noise reduction according to an embodiment. Process 800 can be used as described in the reference. Figure 8 The device architecture shown in the figure is used for implementation.

[0098] Process 800 begins by acquiring an audio signal (e.g., a file, a stream) using one or more processors (801), dividing the audio signal into multiple buffers (802), and generating time-frequency samples for each buffer of the audio signal (803), as per reference. Figure 1 As described in Figure 7.

[0099] Process 800 continues: for each buffer and for each frequency, the median (or mean) and standard deviation of the energy are determined based on the energy in samples in the buffer and samples in adjacent buffers, which together span a specified time range of the audio signal (804); and the median and standard deviation are combined into a cost function (805), as referenced. Figure 1 As described in Figure 7.

[0100] Process 800 continues: for each frequency, the noise floor of the audio signal is estimated as the signal energy of a specific buffer of the audio signal corresponding to the minimum of the cost function (806), and the estimated noise floor is used to reduce noise in the audio signal (807), as referenced. Figure 1 As described in Figure 7.

[0101] Example System Architecture

[0102] Figure 9 An implementation reference according to an embodiment is shown. Figures 1 to 8 A block diagram of an example system describing the features and processes. System 900 includes any device capable of playing audio, including but not limited to: smartphones, tablets, wearable computers, in-vehicle computers, game consoles, surround sound systems, and kiosks.

[0103] As shown, system 900 includes a central processing unit (CPU) 901 capable of executing various processes based on programs stored, for example, in read-only memory (ROM) 902 or loaded from, for example, storage unit 908 into random access memory (RAM) 903. RAM 903 also stores data required by the CPU 901 for executing various processes as needed. CPU 901, ROM 902, and RAM 903 are interconnected via bus 909. Input / output (I / O) interface 905 is also connected to bus 904.

[0104] The following components are connected to I / O interface 905: input unit 906, which may include a keyboard, mouse, etc.; output unit 907, which may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 908, which may include a hard disk or another suitable storage device; and communication unit 909, which may include a network interface card such as a network card (e.g., wired or wireless).

[0105] In some implementations, the input unit 906 includes one or more microphones located at different locations (depending on the host device), which enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0106] In some embodiments, the output unit 907 includes a system having a variety of numbers of speakers. For example... Figure 9 As illustrated, the output unit 907 (depending on the capabilities of the host device) can render audio signals in various formats, such as mono, stereo, immersive, binaural, and other suitable formats.

[0107] Communication unit 909 is configured to communicate with other devices (e.g., via a network). Drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted on drive 910 so that computer programs read from it are installed into storage unit 908. Those skilled in the art will understand that although system 900 is described as including the components described above, in practice, some of these components may be added, removed, and / or replaced, and all such modifications or changes fall within the scope of this disclosure.

[0108] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include computer program products comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 909, and / or installed from removable medium 911, such as... Figure 9 As shown.

[0109] Typically, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be implemented by control circuitry (e.g., with...) Figure 9 The control circuitry executes the actions described herein, which are performed by the CPU in combination with other components. Some aspects may be implemented in hardware, while others may be implemented in firmware or software (e.g., control circuitry) that can be executed by a controller, microprocessor, or other computing device. Although various aspects of the exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, other computing devices, or some combination thereof, as non-limiting examples.

[0110] Additionally, the various blocks shown in the flowchart can be viewed as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform associated functions(s). For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0111] In the context of this disclosure, a machine-readable storage medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0112] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that, when executed by the processor of the computer or other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0113] While this document contains numerous details of specific implementations, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially stated so, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof. The logical flow depicted in the drawings does not require the specific order or ordered sequence shown to achieve the desired result. Additionally, other steps may be provided from the described flow, or steps may be deleted, and other components may be added to or removed from the described system. Therefore, other implementations are within the scope of the following claims.

Claims

1. A method for estimating the noise floor of an audio signal, the method comprising: Audio signals are obtained using one or more processors; The audio signal is divided into multiple buffers using one or more processors; The one or more processors are used to determine time-frequency samples for each buffer of the audio signal; For each buffer and each frequency, the one or more processors determine a measure and median or mean of the amount of energy variation based on time-frequency samples in the buffer and time-frequency samples in adjacent buffers, which together span a specified time range of the audio signal; The one or more processors are used to combine the measure of the change and the median or mean into a cost function; For each frequency: The one or more processors are used to determine the signal energy of a specific buffer of the audio signal corresponding to the minimum value of the cost function; The signal energy is selected using the one or more processors as the estimated noise floor of the audio signal; as well as The noise in the audio signal is reduced using the one or more processors and the estimated noise floor.

2. The method as described in claim 1, wherein, The measure and median or mean of the energy change are scaled to between 0.0 and 1.

0.

3. The method as described in claim 1 or 2, wherein, The cost function increases with the increase of the median or mean, and also increases with the increase of the measure of the energy change.

4. The method as described in claim 1 or 2, wherein, The cost function is non-linear.

5. The method as described in claim 1 or 2, wherein, The cost function is symmetric about the measure and mean or median of the change.

6. The method as described in claim 1 or 2, wherein, The cost function is asymmetric, and when the measure of the energy change is less than a predefined threshold, the weight of the measure of the energy change is less than the weight of the mean or median.

7. The method as described in claim 1 or 2, wherein, The measure of the energy change is: Standard deviation; or The difference between the maximum energy value and the minimum energy value in the buffer spanning the specified time range.

8. The method of claim 7, wherein, The combination of the measure of the change with the mean or median is the sum of its values ​​plus the reciprocal of the sum of its product and 1.

9. The method of claim 7, wherein, The combination of the measure of the change and the median or mean is the sum of their squared values.

10. The method of claim 7, wherein, The combination of the measure of the energy change with the median or mean is the square of the median or mean and the sigmoid of the measure of the change.

11. The method of claim 7, wherein, The combination of the measure of the change and the median or mean is the sigmoid sum of the median or mean and the measure of the change.

12. The method of claim 7, wherein, The measure of the amount of change calculated for blocks of the audio signal and the buffer of the median or mean include at least one buffer where the overall signal energy is below a predefined threshold, and the at least one buffer is not used to estimate the noise floor of the audio signal.

13. The method of claim 12, wherein, The predefined threshold is determined relative to the maximum level of the audio signal.

14. The method of claim 12, wherein, The predefined threshold is determined relative to the average level of the audio signal.

15. The method of claim 7, further comprising: The one or more processors are used to analyze the distribution of blocks of the audio signal and to estimate the noise floor at each frequency based on the distribution; Select block k and frequency f ; If the increased cost is less than the second predefined threshold, then use the block. k Calculated value replacement frequency f Estimated noise at the location.

16. The method of claim 1 or 2, further comprising: The confidence value is determined based on the value of the metric of the change at the selected buffer.

17. The method of claim 16, wherein, The confidence values ​​are smoothed over frequency.

18. The method of claim 16, wherein, Reducing noise in the audio signal further includes: Gain reduction is applied at each frequency, and the gain reduction decreases as the confidence value at that frequency decreases.

19. The method of claim 1 or 2, further comprising: Use one or more processors to select the frequency f 1; Using one or more of the processors, for frequencies higher than the selected frequency. f For all intervals of a predetermined size, calculate the average of the discrete derivatives of the spectrum in segments of a predefined size; The one or more processors select a segment with a maximum negative derivative value as the cutoff frequency when the maximum negative derivative is less than a predefined value. f c ; as well as The processors are used to replace spectral values ​​above the cutoff frequency with the average value of the spectrum in a frequency band of a predefined length having an upper boundary adjacent to the cutoff frequency.

20. A system for estimating the noise floor of an audio signal, comprising: One or more processors; as well as A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operation of the method according to any one of claims 1 to 19.

21. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operation of the method according to any one of claims 1 to 19.

22. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the operation of the method according to any one of claims 1 to 19.

Citation Information

Patent Citations

  • System and method for generating a separated signal

    CN101558397A

  • Method and apparatus for robust acoustic feedback cancellation

    CN107786925A