Adaptive noise estimation
The adaptive noise estimation method addresses the challenge of sudden noise changes by segmenting audio inputs and selecting the most likely noise spectrum from past and future estimates, ensuring accurate noise reduction during speech segments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-21
- Publication Date
- 2026-03-31
AI Technical Summary
Existing noise estimation methods struggle to adapt to sudden changes in noise levels or spectrum in audio recordings, especially when background noise changes abruptly or when audio content is edited from different locations, and they often fail to accurately estimate noise during non-speech segments.
An adaptive noise estimation method that splits audio inputs into voice and non-voice segments, estimates the time-varying noise spectrum during non-voice segments, and selects the most likely noise spectrum from past and future estimates using similarity metrics, even during voice segments.
Effectively adapts to noise changes in audio recordings, providing accurate noise estimation even during speech, thereby enhancing noise reduction performance.
Smart Images

Figure 0007837956000003 
Figure 0007837956000004 
Figure 0007837956000005
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This application claims priority to U.S. Provisional Application No. 63 / 120,253 filed on 2 December 2020, U.S. Provisional Application No. 63 / 168,998 filed on 31 March 2021, and Spanish Patent Application No. P202030960 filed on 23 September 2020, all of which are incorporated herein by reference in their entirety.
[0002] [Technical field] This disclosure relates in general to audio signal processing, and in particular to estimating the noise floor of an audio signal for use in noise reduction. [Background technology]
[0003] Noise estimation is commonly used to reduce steady-state noise in audio recordings. Typically, noise estimation is achieved by analyzing the energy of each frequency band across segments of an audio recording that contain only noise. However, in some audio recordings, steady-state noise changes smoothly and / or abruptly over time. Examples of such abrupt changes include audio recordings where background ambient noise changes abruptly over time (e.g., turning a fan on and off in a room), or audio content obtained by editing together different audio recordings, each with a different noise floor, such as a podcast containing a series of interviews recorded in different locations. Furthermore, noise changes usually do not occur during sufficiently long segments of non-speech, so noise changes may not be detected and estimated in the early stages of an audio recording.
[0004] Some existing methods use segments of audio recordings containing only noise to perform a single estimate of the noise floor. Other existing methods perform an analysis of the entire audio recording that converges to a single underlying noise floor. However, a drawback of these methods is that they cannot adapt to changes in noise levels or spectrum. Other existing methods estimate the minimum energy envelope for each frequency band and track the estimated minimum envelope over time (for example, by smoothing the estimated minimum envelope with an appropriate time constant). However, these existing methods are commonly employed in real-time online audio signal processing architectures and cannot accurately respond to sudden changes in noise within an audio recording. [Overview of the Initiative]
[0005] An implementation for adaptive noise estimation is disclosed.
[0006] In some embodiments, the adaptive noise estimation method is: The steps include: using at least one processor to split the audio input into an audio segment and a non-audio segment, Using at least one of the aforementioned processors, the steps include: estimating the time-varying noise spectrum of each non-voice segment for each frame of the non-voice segment; Using at least one of the aforementioned processors, the steps include: estimating the audio spectrum of each audio segment for each frame of the audio segment; For each frame of each audio segment, The steps include identifying one or more non-speech frequency components of the aforementioned speech spectrum, The steps include comparing one or more non-speech frequency components with one or more corresponding frequency components of a plurality of estimated noise spectra, The steps include selecting an estimated noise spectrum from the multiple estimated noise spectra based on the comparison results, This includes, in an embodiment, the method further includes the step of using the at least one processor to reduce the noise of the audio input using the selected estimated noise spectrum.
[0007] In some embodiments, the method further includes the step of obtaining a probability of speech in each frame of the audio input and identifying frames containing speech based on the probability.
[0008] In some embodiments, the time-varying noise spectrum is estimated by calculating a moving average of the power spectrum of the non-voice segment and averaging the power spectrum of the current non-voice segment with that of at least one past non-voice segment.
[0009] In some embodiments, during the non-voice segment, the time-varying estimated noise spectrum is supplied to a noise reduction unit configured to reduce noise in the audio input using the selected estimated noise spectrum.
[0010] In some embodiments, for each audio segment, the estimated noise spectrum from the past prior to the audio segment, the estimated noise spectrum from the future following the audio segment, and the current audio frame are used to determine the estimated noise spectrum that is most likely to represent the noise of the current audio segment.
[0011] In some embodiments, determining the estimated noise spectrum most likely to represent the noise of the current audio segment is: A step of obtaining the average noise spectrum from the past and future noise spectra of the past and future non-voice segments before and after the aforementioned voice segment, respectively. The steps include determining the upper frequency limits of the past and future noise spectra, The steps include determining the cutoff frequency, which is the lowest of the two upper frequency limits, and Calculating a distance metric between the frequency components of the speech spectrum and the frequency components of the noise spectrum; Selecting, as the estimated noise spectrum of the audio input, the noise spectrum having the minimum distance metric up to the cut-off frequency among the past or future noise spectra; further comprising.
[0012] In some embodiments, the distance metric is averaged over a set of speech frames within a speech segment.
[0013] In some embodiments, a speech component is estimated within the speech segment of the audio signal and then subtracted from the actual speech component to obtain a residual spectrum as the estimated non-speech frequency component.
[0014] In some embodiments, an audio processor a splitter configured to split an audio input into segments of overlapping frames; a plurality of buffers configured to store the segments of the overlapping frames; a spectrum analysis unit configured to calculate a frequency spectrum for each segment stored in each buffer; a voice activity detector (VAD) configured to detect speech segments and non-speech segments in the audio input; an averaging unit coupled to the output of the VAD and configured to calculate a speech spectrum for each speech segment identified by the VAD output and a time-varying noise spectrum for each non-speech segment identified by the VAD output; including.
[0015] In an embodiment, an audio processor a VAD configured to detect speech segments and non-speech segments in the audio input; An averaging unit configured to be coupled to the output of the VAD and obtain an audio spectrum for each voice segment identified by the VAD output and a noise spectrum for each non-voice segment identified by the VAD output; A similarity metric unit configured to calculate a similarity metric between one or more frequency components of the current audio spectrum and corresponding one or more frequency components of each noise spectrum and select one noise spectrum from the noise spectra based on the similarity metric; A noise reduction unit configured to reduce the noise of the audio input using the selected noise spectrum; including.
[0016] Other implementations disclosed herein are directed to systems, devices, and computer-readable media. Details of the disclosed implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0017] Certain implementations disclosed herein provide one or more of the following advantages. A method for adaptively estimating noise in an audio recording when voice is present is disclosed. In an embodiment, the adaptive noise estimation is performed offline on an audio recording and estimates noise changes by looking at both before and after a given frame of the audio recording. An advantage compared to conventional adaptive noise estimation methods is that the noise floor under the voice is estimated by selecting from among the best available candidate noise floor estimates calculated before and after the current voice segment.
Brief Description of the Drawings
[0018] In the figures, specific configurations or sequences of schematic elements such as devices, units, instruction blocks, and data elements are shown for the sake of simplicity of explanation. However, as should be understood by those skilled in the art, the specific order or configuration of schematic elements in the figures does not imply that a specific order or sequence of processes or separation of processes is required. Furthermore, the inclusion of schematic elements in the figures does not imply that such elements are required in all embodiments, or that features represented by such elements are not included in or combined with other elements in some implementations.
[0019] Furthermore, where connecting elements, such as solid or dashed lines or arrows, are used in a figure to illustrate connections, relationships, or associations between or within two or more other schematic elements, the absence of any such connecting element does not mean that a connection, relationship, or association does not exist. In other words, some connections, relationships, or associations between elements are not shown in the figure so as not to obscure this disclosure. Furthermore, for the sake of clarity, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of signals, data, or instructions, it should be understood by those skilled in the art that such elements may affect the communication as one or more signal paths may be involved.
[0020] [Figure 1] This is a two-dimensional (2D) plot showing an audio waveform, speech activation over time, and a threshold used to determine the non-speech segment of the audio waveform, according to several embodiments.
[0021] [Figure 2] This is a 2D plot of thresholds used to determine speech activity over time, non-speech segments of an audio waveform, and noise segments where speech activity is below a threshold, according to several embodiments.
[0022] [Figure 3]The images show the average speech spectrum corresponding to a speech segment and two noise spectra corresponding to the non-speech segments before and after the speech segment, according to several embodiments.
[0023] [Figure 4] This is a block diagram of a system for adaptive noise estimation and noise reduction according to several embodiments.
[0024] [Figure 5] This is a flowchart of the process for noise floor estimation and noise reduction according to several embodiments.
[0025] [Figure 6] This is a block diagram of a system that implements the functions and processes described with reference to Figures 1-5, according to the embodiment.
[0026] The same reference numerals used in various drawings indicate similar elements. [Modes for carrying out the invention]
[0027] In the following detailed description, many specific details are described in order to provide a complete understanding of the various described embodiments. It will be apparent to those skilled in the art that the various described implementations may be carried out without these specific details. In other examples, well-known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure the aspects of the embodiments. Several features that can be used independently of each other or in any combination with other features are described below.
[0028] <Nomenclature> When used in this specification, the term “includes” and its variations shall be interpreted as a broad term meaning “includes, but not limited to.” The term “or” shall be interpreted as “and / or” unless explicitly indicated in the context. The term “based on” shall be interpreted as “based at least in part.” The terms “one exemplary implementation” and “exemplary implementation” should be interpreted as “at least one exemplary implementation.” The term “another implementation” should be interpreted as “at least one other implementation.” The terms “determined,” “determine,” or “to determine” should be interpreted as “obtain,” “receive,” “calculate,” “calculate,” “estimate,” “predict,” or “derive.” Furthermore, in the following descriptions and claims, unless otherwise noted, all technical and scientific terms used in this specification shall have the same meaning as those generally understood by those skilled in the art to which this disclosure belongs.
[0029] <System Overview> The disclosed embodiments use a Voice Activity Detection (VAD) classifier to divide an audio input into a voice segment containing voice and a non-voice segment not containing voice. In the non-voice segment, the noise spectrum is estimated for each frame of the non-voice segment by averaging the energy per frequency in the time domain around the current frame. In the voice segment, for each frame of the voice segment, estimated noise spectra of the temporally preceding and succeeding non-voice regions are selected by identifying one or more non-voice frequency components in the voice spectrum. One or more non-voice frequency components are compared to the corresponding one or more frequency components in the estimated noise spectra of the previous and succeeding non-voice regions using a similarity metric (e.g., distance between frequency components).
[0030] Figure 1 is a two-dimensional (2D) plot showing an audio waveform, speech activity over time, and a threshold used to determine the non-speech segment of the audio waveform, according to an embodiment. For simplicity, the amplitude values of the audio waveform are not shown in Figure 1. The horizontal axis is in units of time (e.g., milliseconds). An audio input (e.g., an audio file) containing an audio recording with speech is divided into overlapping frames. In the embodiment, the probability of speech in each frame is obtained using a VAD, and the audio input is then divided into speech and non-speech segments based on a threshold of speech probability. In the example shown, the vertical axis represents the VAD value (probability of speech being present), and the exemplary VAD threshold shown by the horizontal line is approximately 0.18. Figure 2 shows a close-up of the noise segment shown in Figure 1, where the VAD value is lower than the VAD threshold.
[0031] Any suitable VAD algorithm can be used to detect speech and non-speech segments in audio recordings, including, but not limited to, VAD algorithms based on zero crossover rate and energy measurement, linear-based energy detection, adaptive linear-based energy detection, pattern recognition, and statistical measurement.
[0032] In the embodiment, the noise spectrum of the non-voice segment is estimated using adaptive voice-aware noise estimation (AVANE) to infer the most similar and robust noise estimate within the voice segment. AVANE calculates the power spectrum of noise within the non-voice frame by calculating a moving average of the power spectrum of the non-voice frame and, for each non-voice frame, averaging the power of the current non-voice frame with one or more past non-voice frames. In the embodiment, the number of past frames to average is determined by a time constant. Any suitable moving average algorithm can be used, but is not limited to, arithmetic moving average, exponential moving average, smoothed moving average, or weighted moving average.
[0033] AVANE generates a time-varying noise spectrum used in two ways. First, during non-speech segments, the time-varying estimated noise is fed to the noise reduction system (e.g., per buffer). Second, during speech segments, the last AVANE estimate before the current speech segment and the first AVANE estimate after the current speech segment are fed to the estimation component along with the current speech frame. The estimation component determines the AVANE estimate that is most likely to represent the noise in the current speech frame.
[0034] Alternatives to AVANE estimation include subband spectral minimum tracking, as described, for example, by Doblinger, G. (1995). Computationally efficient speech enhancement by subband spectral minimum tracking. Proc. EUROSPEECH'95, Madrid, pp 1513-1516, or noise power spectral density estimation based on optimal smoothing and minimum statistics, as described, for example, by Martin, R. (2001). Noise power spectral density estimation based on optimal smoothing and minimum statistics. IEEE Transactions on Speech and Audio Processing. 9(5)504-512.
[0035] Two embodiments are proposed for estimating the underlying noise spectrum of a given speech segment. In the first embodiment, after estimating the speech components, the residual spectrum is obtained as a noise estimate by subtracting it from the actual speech components. This embodiment is unrelated to and uncoupled with AVANE, as it leads to a direct estimation of background noise. Assuming that the speech is dominated by harmonic components, the pitch is first estimated and the harmonic components are identified. Based on a sine wave model and its parameter estimation, the harmonic components are subtracted from the speech signal to obtain a residual signal. This method is described, for example, in Stylianou, Y. (1996) Harmonic plus Noise Models for Speech combined with Statistical Methods for Speech and Speaker Modification, PhD Thesis, Telecom Paris. Another possibility is to identify and subtract the sine wave of a given short-time spectrum without fundamental frequency (F0) information. This method is described, for example, in Yeh, C. (2008) Multiple Fundamental Frequency Estimation of Polyphonic Recordings. Ph.D. thesis, University Paris.
[0036] In another embodiment, the harmonic components are estimated and attenuated in the cepstrum region, for example, as described in Z. Zhang, K. Honda, and J. Wei, "Retrieving Vocal-Tract Resonance and anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique," ICASSP2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 7359-7363, doi:10.1109 / ICASSP40776.2020.9054741.
[0037] The AVANE method assumes that the underlying noise spectrum is close to either the last AVANE before the speech segment or the first AVANE after the speech segment. In this embodiment, a spectral similarity index (e.g., distance index) between the speech spectrum and the AVANE is calculated by identifying the non-speech spectral segments (e.g., high frequencies) and considering only the non-speech segments of the spectrum that are primarily noise components. In this embodiment, the spectral similarity measurement is based on the distance between the speech spectrum and the AVANE. Assuming that the signal-to-noise ratio (SNR) (defined in decibels as the ratio of the energy of the speech to the energy of the noise in the speech frequency band) is positive, an additional constraint can be added: the selected AVANE can only be accepted if the average speech spectrum of the region of interest (within a duration equal to the length of the AVANE) exceeds the AVANE to be selected.
[0038] In embodiments where harmonic subtraction is used to calculate the noise spectrum estimate, the spectral similarity measurement is not limited to the non-speech frequency region of the speech spectrum, but can be extended to the entire spectrum or limited to frequencies above a specific speech frequency range where harmonic estimation is effective, such as the lowest speech frequency range. Thus, the similarity index is calculated between the residual signal after harmonic subtraction from the speech segment and the AVANE estimates before and after the speech segment.
[0039] In the embodiment, given an audio frame, the energy spectrum of the audio frame is calculated and converted to a decibel scale. If the current audio frame is a voice frame (i.e., within a voice segment), the previously calculated average noise spectrum (dB) before and after the voice segment is retrieved, for example, from storage (e.g., memory, disk). Figure 3 shows the average voice spectrum and two noise spectra corresponding to the non-voice segments before and after the voice segment, according to several embodiments.
[0040] Given these two noise spectra and the current audio spectrum, the upper frequency limit f of the noise spectrum c The calculation is performed, and the lowest of the two limits is the "cutoff" frequency f cutoff It is retained as follows. Next, in this example, a similarity metric, which is the sum of the absolute values of the differences ("distance") between the audio spectrum and two noise spectra, is calculated, for example, over the segment from the halfway point of the audio spectrum to the cutoff frequency. The noise spectrum with the minimum distance (as defined earlier) is retained as the current estimate of the noise spectrum of the audio recording. In an alternative embodiment, the distance metric can be calculated and averaged over a set of audio frames, and the noise spectrum that gives the minimum average distance is selected as the current estimate of the noise spectrum.
[0041] Assume that audioframe is a vector of audio samples within a frame, and spectrum is the frequency spectrum of the audio samples calculated using the Fast Fourier Transform (FFT) of audioframe. spectrum=fft(audioframe) [1]
[0042] spectrum can be converted to spectrum on a dB scale by the following formula dB as follows: spectrum dB =20log 10 (abs(spectrum)) [2]
[0043] If the current frame is a noise frame, its avg_spectrum dB is retained and averaged with the past spectra over a window of a given length (e.g., 5 seconds), denoted here as avg_spectrum dB If the current frame is a speech frame, its spectrum is compared with the past noise spectrum and the future noise spectrum. Hereinafter, the speech spectrum is denoted as speech_spectrum dB and the past and future noise spectra are denoted as past_spectrum dB and future_spectrum dB respectively.
[0044] In some embodiments, the upper frequency f dB of each of past_spectrum dB and future_spectrum c is determined as follows: 1) Select a first frequency. f c is estimated above the first frequency. 2) Divide the noise spectrum above the first frequency into blocks of a specified length and overlap (e.g., 50%). 3) For each block, calculate the mean derivative in descending order of frequency of the corresponding block, and find the first derivative that has a value less than the predefined negative value (e.g., -20 dB). 4)f c The average of the noise spectrum in the earlier small region is calculated, f c Replace the higher noise spectrum values with the average noise spectrum. Step (3) is interpreted as a significant attenuation of the noise spectrum, and it should be noted that the frequency of the corresponding block is considered the upper frequency limit.
[0045] The determined upper frequency f c And the lower of the frequencies exceeding the audio frequency f1 is the cutoff frequency f cutoff Given the values, the distance between the current audio spectrum and the noise spectrum is calculated as follows:
number
[0046] As shown in equation [4], f1 and f cutoff The frequency range between `distance_past` and `distance_future` defines a spectral region where audio harmonics are almost nonexistent and background noise is dominant. The minimum value between `distance_past` and `distance_future` (given by `argmin()`) gives a noise spectrum close to the current spectrum and is selected as a noise candidate. This approach can be extended to multiple candidate noise spectra.
[0047] In embodiments for estimating and removing speech harmonics using harmonic subtraction, the methods described in equations 3a, 3b, and 4 can be extended to speech frequencies by replacing the starting exponent f1 with a lower frequency exponent, such as the lowest speech frequency or the lowest frequency for which the residual estimate is considered reliable.
[0048] Given any method described here that can estimate noise in the presence of speech (e.g., the AVANE method), the distance between the estimated spectrum and two known noise spectra can be calculated by comparing the estimate obtained from AVANE for the current frame and adjacent non-speech segments and selecting either past or future noise estimates as described above.
[0049] Figure 4 is a block diagram of a system 400 for adaptive noise estimation and noise reduction according to an embodiment. The audio input (e.g., an audio file containing speech content) is divided into overlapping frame segments by the splitting unit 401, and the resulting segments are stored in a plurality of buffers 402 and converted into a spectrum 405 by, for example, the short-time Fourier transform (STFT) block 403. The Voice Activity Detection (VAD) block 404 calculates the probability that a given audio frame contains speech. The spectrum 405 and the VAD output (speech probability) are fed to the averaging unit 406, which generates the current speech spectrum and a plurality of noise spectra 407 for each speech frame. The speech spectrum and the plurality of noise spectra 407 are input to the similarity metric unit 408. The similarity metric unit 408 selects one of the noise spectra (e.g., based on the distance metric in equations [3a, 3b]) as the noise spectrum 410 that the noise reduction block 409 uses to reduce the noise of the audio input.
[0050] In some embodiments, the noise reduction unit 409 reduces the noise in the audio input using the selected noise spectrum 410 by comparing the spectrum of the audio input with the selected estimated noise spectrum 410 and applying gain reduction to the frequency band where the energy of the input signal is less than the energy of the noise spectrum plus a predefined threshold.
[0051] <Other Embodiments> The following description of further embodiments will focus on the differences between the further embodiments and the embodiments described above. Therefore, features common to both embodiments will be omitted in the following description, and unless otherwise specifically required in the following description, it should be assumed that the features of the embodiments described above can be implemented, or at least can be implemented, in the further embodiments.
[0052] In some embodiments, multiple pre-calculated noise spectra are available. noise_spectrum i Here, i=1,..,N, [5] Furthermore, the similarity index is the distance (on a dB scale) between the current audio spectrum and multiple noise spectra, and is given by the following equation:
number
[0053] The noise spectrum corresponding to smaller distances is selected as follows: noise_spectrum K Here, K = argmin(distance i ) [7]
[0054] Multiple noise spectra can be provided a priori in applications where different noise conditions observed in audio recordings are known and measured in advance, such as in a conference call with multiple endpoints. Alternatively, multiple noise spectra can be determined by a clustering algorithm applied to multiple spectra of non-voice frames. The clustering algorithm can be, for example, a k-means clustering algorithm applied to multiple non-voice spectral vectors, or any other suitable clustering algorithm.
[0055] <Online Implementation> The above embodiment for offline calculation can be extended to real-time, online, low-latency scenarios. Note that in this case, future noise spectra after the current speech frame are not available. If candidate noise spectra are provided a priori, the selection process is applied online to all speech frames using the available (stored) noise spectra. If candidate noise spectra are not provided a priori, noise spectra can be constructed online. For example, the first noise spectrum is obtained from the first non-speech frame. When additional non-speech frames are received, their noise spectra are calculated and retained as additional noise spectra if their distance from each previously retained noise spectrum is greater than a predefined threshold. When additional non-speech frames are received, their noise spectra are calculated and clustered by a clustering algorithm (e.g., k-means clustering), and the resulting clusters are used as candidate noise spectra. The clustering process is repeated and refined whenever a sufficient number of new non-speech frames are received, or whenever non-speech frames with significant dissimilarity to existing clusters are received.
[0056] <Music Recording> In this embodiment, the audio recording includes music (or another class of audio content) instead of voice content. In this embodiment, the voice classifier VAD is replaced with an appropriate music (or another class) classifier.
[0057] <Music + Voice Recording> In this embodiment, the audio recording includes both speech and music. In this embodiment, it is desirable to remove noise from the speech and music portions while preserving the music signal. In this embodiment, the speech classifier is replaced by a multi-class classifier (e.g., a music and speech classifier) or two separate classifiers for music and speech. The probabilities of speech and music output by the classifiers are compared to a predefined threshold, and if both the speech and music probabilities are less than the predefined threshold, the frame is considered noise. Next, the aforementioned method is applied to estimate a noise spectrum suitable for the speech region, and optionally for the music region as well.
[0058] <Example of processing> Figure 5 is a flowchart of the process 500 for noise floor estimation and noise reduction according to an embodiment. Process 500 can be performed using the apparatus architecture shown in Figure 6.
[0059] Process 500 begins by dividing the audio input into a voice segment and a non-voice segment (501), and for each frame of each non-voice segment, estimates the time-varying noise spectrum of the non-voice segment (503) and the voice spectrum of the voice segment (504).
[0060] The process 500 then identifies one or more non-voice frequency components of the voice spectrum for each frame of each voice segment (505), compares one or more non-voice frequency components with one or more corresponding frequency components of a plurality of estimated noise spectra (506), and selects an estimated noise spectrum from the plurality of estimated noise spectra based on the results of the comparison (507).
[0061] In some embodiments, the multiple estimated noise spectra include estimated noise spectra for past non-voice segments and estimated noise spectra for future non-voice segments. In some embodiments, the multiple estimated noise spectra can be determined by a clustering algorithm applied to multiple noise spectra of non-voice frames. The clustering algorithm can be, for example, a k-means clustering algorithm applied to multiple non-voice spectral vectors, or any other suitable clustering algorithm.
[0062] In some embodiments, the process 500 can subsequently reduce the noise of the audio input using the selected estimated noise spectrum.
[0063] <Example System Architecture> Figure 6 shows a block diagram of an exemplary system that implements the functions and processes described with reference to Figures 1-5, according to an embodiment. System 600 includes, but is not limited to, any device capable of playing audio, such as a smartphone, tablet computer, wearable computer, vehicle computer, game console, surround sound system, or kiosk.
[0064] As shown in the figure, the system 600 includes a central processing unit (CPU) 601 that can perform various operations according to a program stored in, for example, a read-only memory (ROM) 602 or a program loaded from, for example, a memory unit 608 into a random access memory (RAM) 603. The RAM 603 also stores data necessary for the CPU 601 to perform various operations as needed. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 609. An input / output (I / O) interface 605 is also connected to a bus 604.
[0065] The following components are connected to the I / O interface 605: an input unit 606 which may include a keyboard, mouse, etc.; an output unit 607 which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 608 which includes a hard disk or another suitable storage device; and a communication unit 609 which includes a network interface card such as a network card (e.g., wired or wireless).
[0066] In some implementations, the input unit 606 includes one or more microphones located at different positions (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, or other appropriate formats).
[0067] In some implementations, the output unit 607 includes a system with varying numbers of speakers. As shown in Figure 6, the output unit 607 can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, or other appropriate formats) depending on the capabilities of the host device.
[0068] The communication unit 609 is configured to communicate with other devices (for example, via a network). The drive 610 is also connected to the I / O interface 605, if necessary. A removable medium 611, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or another suitable removable medium, is mounted on the drive 610, and as a result, computer programs read from it are installed in the storage unit 608, if necessary. Those skilled in the art will understand that although the system 600 is described as including the components described above, in actual applications it is possible to add, remove, and / or replace some of these components, and that all such changes or variations are all included within the scope of this disclosure.
[0069] According to exemplary embodiments of the present disclosure, the processing described above may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium. The computer program includes program code for performing the method. In such embodiments, the computer program may be downloaded and implemented from a network via a communication unit 609 and / or installed from a removable medium 611, as shown in Figure 6.
[0070] Typically, various exemplary embodiments of this disclosure may be implemented in hardware or dedicated circuitry (e.g., control circuits), software, logic, or any combination thereof. For example, the above-described unit may be executed by a control circuit (e.g., a CPU in combination with the other components in Figure 6). Thus, the control circuit can perform the operations described in this disclosure. Some embodiments may be implemented in hardware, while others may be firmware or software implementations that may be executed by a control unit, microprocessor, or other computing device (e.g., a control circuit). Although various embodiments of the exemplary embodiments of this disclosure have been illustrated and described using block diagrams, flowcharts, or some other graphical representations, it should be understood that the blocks, devices, systems, techniques, or methods described in this specification may, in non-limiting examples, be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or control units or other computing devices, or any combination thereof.
[0071] Furthermore, the various blocks shown in the flowchart may be considered as steps of a method, and / or as operations resulting from operations in computer program code, and / or as a plurality of coupled logic circuit elements configured to perform related functions. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium. The computer program includes program code configured to perform the methods described above.
[0072] In the context of this disclosure, a machine-readable medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be intangible and may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatus, or any suitable combination thereof. More specific examples of machine-readable storage media may include one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0073] Computer program code for performing the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to the processor of a general-purpose computer, a dedicated computer, or another programmable data processing device having a control circuit. As a result, when the program code is executed by the processor of the computer or other programmable data processing device, it causes the functions / operations specified in the flowchart and / or block diagrams to be performed. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0074] <Enumerated Example Embodiments (EEE)>
[0075] Embodiments of this disclosure may relate to one of the listed embodiments (EEE) below.
[0076] (EEE1) audio processor, A splitter configured to divide an audio input into segments of overlapping frames, Multiple buffers configured to store the overlapping frame segments, A spectral analysis unit configured to calculate the frequency spectrum for each segment stored in each buffer, A voice activity detector (VAD) configured to detect voice segments and non-voice segments in the aforementioned audio input, An averaging unit coupled to the output of the VAD and configured to calculate an audio spectrum for each audio segment identified by the VAD output, and a time-varying noise spectrum for each non-audio segment identified by the VAD output; a similarity metric unit configured to calculate a similarity metric between the current audio spectrum and one or more frequency components in each noise spectrum, and to select one noise spectrum from the plurality of noise spectra based on the similarity metric; A noise reduction unit configured to reduce the noise of the audio input using the selected noise spectrum, An audio processor that includes [this component].
[0077] (EEE2) The audio processor according to EEE1, wherein the VAD is configured to acquire the probability of sound in each frame of the audio input and to identify frames containing sound based on the probability.
[0078] (EEE3) audio processor, A voice activity detector (VAD) configured to detect voice and non-voice segments in an audio input, An averaging unit coupled to the output of the VAD and configured to acquire an audio spectrum for each audio segment identified by the VAD output and a noise spectrum for each non-audio segment identified by the VAD output, A similarity metric unit configured to calculate a similarity metric between one or more frequency components of the current audio spectrum and one or more corresponding frequency components of each noise spectrum, and to select one noise spectrum from the noise spectrum based on the similarity metric, A noise reduction unit configured to reduce noise in the audio input using a selected noise spectrum, An audio processor that includes [this component].
[0079] This specification includes numerous specific implementation details, which should be considered not as limitations on the scope of what can be claimed, but rather as descriptions of features specific to a particular implementation of a particular implementation. Specific features described in this specification in the context of separate embodiments may be combined and implemented in a single implementation. Conversely, various features described in the context of a single embodiment may be implemented separately or in any suitable partial combination in multiple embodiments. Furthermore, features are described above to operate in a particular combination and may be initially claimed as such, but one or more features from the claimed combination may, in some cases, be decoupled from the combination, and the claimed combination may be directed towards a partial combination or a variation of a partial combination. The logical flow shown in the drawings does not require a specific order or sequential sequence shown to achieve the desired result. Furthermore, other steps may be provided, or steps may be removed from the described flow, and other components may be added to or removed from the described system. Therefore, other implementations are within the scope of the following claims.
Claims
1. A method for adaptive noise estimation, The steps include: using at least one processor to divide the audio input into an audio segment and a non-audio segment; Using at least one of the aforementioned processors, the steps include: estimating the time-varying noise spectrum of each non-voice segment for each frame of the non-voice segment; Using at least one of the aforementioned processors, the steps include: estimating the audio spectrum of each audio segment for each frame of the audio segment; For each frame of each audio segment, The steps include using at least one of the aforementioned processors to identify one or more non-speech frequency components of the speech spectrum, A step of using at least one processor to compare one or more non-voice frequency components with one or more corresponding frequency components of a plurality of estimated noise spectra to obtain a distance metric for each noise spectrum, The steps include using at least one of the processors to select an estimated noise spectrum having a minimum distance metric as the estimated noise spectrum of the audio segment, A method that includes this.
2. The method according to claim 1, further comprising the step of using at least one processor to reduce the noise of the audio input using the selected estimated noise spectrum.
3. The method according to claim 1 or 2, wherein the time-varying noise spectrum is estimated by calculating a moving average of the power spectrum of the non-voice segment and averaging the power spectrum of the current non-voice segment with that of at least one past non-voice segment.
4. The method according to any one of claims 1 to 3, wherein the plurality of estimated noise spectra include estimated noise spectra for past non-voice segments and estimated noise spectra for future non-voice segments.
5. The method according to any one of claims 1 to 4, further comprising the steps of obtaining the probability of sound in each frame of the audio input and identifying a frame containing sound based on the probability.
6. The steps include obtaining the past average noise spectrum from the past noise spectrum of the non-voice segment preceding the voice segment, and the future average noise spectrum from the future noise spectrum of the non-voice segment following the voice segment, The steps include determining the upper frequency limit of the past and future average noise spectrum, The steps include determining the cutoff frequency, which is the lowest of the two upper frequency limits, The steps include: calculating the distance metric between the non-speech frequency components of the speech spectrum and the frequency components of the past and future average noise spectrum up to the cutoff frequency; The steps include selecting the average noise spectrum having the minimum distance metric from the past or future average noise spectra as the estimated noise spectrum of the audio input, The method according to any one of claims 1 to 5, further comprising:
7. The method according to any one of claims 1 to 6, wherein the distance metric is averaged over a set of audio frames in an audio segment.
8. The method according to any one of claims 1 to 7, further comprising the steps of: estimating the speech frequency components within the speech segment of an audio signal; subtracting the speech frequency components from the speech spectrum to obtain a residual spectrum as the estimated non-speech frequency components.
9. A method for adaptive noise estimation, The steps include receiving multiple pre-calculated noise spectra, The steps include: using at least one processor to divide the audio input into an audio segment and a non-audio segment; Using at least one of the aforementioned processors, the steps include: estimating the audio spectrum of each audio segment for each frame of the audio segment; For each frame of each audio segment, The steps include using at least one of the aforementioned processors to identify one or more non-speech frequency components of the speech spectrum, A step of using at least one processor to compare one or more non-voice frequency components with one or more corresponding frequency components of the plurality of pre-calculated noise spectra to obtain a distance metric for each pre-calculated noise spectrum, The steps include using at least one of the processors to select a pre-calculated noise spectrum having a minimum distance metric as the estimated noise spectrum of the audio segment, A method that includes this.
10. A non-temporary computer-readable storage medium storing instructions, wherein, when the instructions are executed by one or more processors, the one or more processors cause the one or more processors to execute the method according to any one of claims 1 to 8 or 9.
11. It is an audio processor, A splitter unit configured to divide an audio input into an audio segment and a non-audio segment, An averaging unit configured to estimate the audio spectrum for each audio segment and the time-varying noise spectrum for each non-audio segment, Distance metric unit, Identify one or more non-speech frequency components of the aforementioned speech spectrum, The one or more non-speech frequency components are compared with one or more corresponding frequency components in a plurality of estimated noise spectra. The noise spectrum having the minimum distance metric is selected as the estimated noise spectrum of the audio segment. A distance metric unit configured in such a way, An audio processor that includes [this component].
12. The audio processor according to claim 11, wherein the plurality of estimated noise spectra include estimated noise spectra for past non-voice segments and estimated noise spectra for future non-voice segments.
13. The audio processor according to claim 11 or 12, further comprising a noise reduction unit configured to reduce the noise of the audio input using the selected estimated noise spectrum.
14. The aforementioned distance metric unit is The past average noise spectrum is obtained from the noise spectrum of the non-voice segment preceding the voice segment, and the future average noise spectrum is obtained from the noise spectrum of the non-voice segment following the voice segment. The upper frequency limits of the past and future noise spectra are determined, Determine the cutoff frequency, which is the lowest of the two upper frequency limits. Up to the aforementioned cutoff frequency, calculate the distance metric between the non-voice frequency components of the voice spectrum and the frequency components of the past and future average noise spectrum. As the estimated noise spectrum of the audio input, the average noise spectrum having the minimum distance metric is selected from the past or future average noise spectra. An audio processor according to any one of claims 11 to 13, configured as follows.
15. It is an audio processor, A splitter unit configured to divide an audio input into an audio segment and a non-audio segment, An averaging unit configured to estimate the audio spectrum for each audio segment and receive multiple pre-calculated noise spectra, Distance metric unit, Identify one or more non-speech frequency components of the aforementioned speech spectrum, The one or more non-speech frequency components are compared with one or more corresponding frequency components in the plurality of pre-calculated noise spectra. The pre-calculated noise spectrum having the minimum distance metric is selected as the estimated noise spectrum of the audio segment. A distance metric unit configured in such a way, An audio processor that includes [this component].
Citation Information
Patent Citations
Speech recognition device
JP2001318687A
Noise estimation using an adaptive smoothing factor based on a teager energy ratio in a multi-channel noise suppression system
US20110099007A1
Coordination of beamformers for noise estimation and noise suppression
US20180033447A1
Method of noise reduction for speech codecs
US6453289B1