Adaptive noise estimation
By using a Voice Activity Detector (VAD) to segment speech and non-speech segments in audio recordings and performing similarity measurements to select noise spectrum for noise reduction spectrum analysis, this method solves technical problems that cannot be solved in the prior art. It addresses the technical issues of speech segments in the presence of both speech and non-speech segments, as well as speech segments within speech segments, and achieves spectral changes between speech and non-speech segments, thereby improving the accuracy of noise estimation and the noise reduction effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2021-09-21
- Publication Date
- 2026-05-05
AI Technical Summary
Existing noise estimation methods cannot adapt to changes in noise in audio recordings, especially spectral variations between speech segments and non-speech segments, resulting in an inability to accurately reduce noise.
Audio input is divided into speech segments and non-speech segments by using a voice activity detector (VAD). The time-varying noise spectrum of the non-speech segments and the speech spectrum of the speech segments are estimated separately. The most suitable noise spectrum is selected for noise reduction by using a similarity metric.
It achieves adaptive estimation and reduction of noise in audio recordings in the presence of speech, improving the accuracy of noise estimation and the noise reduction effect.
Smart Images

Figure CN116324985B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 120,253, filed December 2, 2020; U.S. Provisional Application No. 63 / 168,998, filed March 31, 2021; and Spanish Patent Application No. P202030960, filed September 23, 2020, each of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure generally relates to audio signal processing, and more specifically to estimating the noise floor in an audio signal for noise reduction. Background Technology
[0004] Noise estimation is commonly used to reduce steady-state noise in audio recordings. Typically, noise estimation is obtained by analyzing the energy in each frequency band of an audio recording segment containing only noise. However, in some audio recordings, steady-state noise can change smoothly and / or abruptly over time. Some examples of such abrupt changes include audio recordings where background ambient noise changes abruptly over time (e.g., a fan in a room being turned on or off), and audio content obtained by editing together different audio recordings, each with different background noise levels (such as a podcast containing a series of interviews recorded at different locations). Additionally, noise changes often do not occur during sufficiently long non-speech segments, and therefore may not be detectable and estimated early in the audio recording.
[0005] Some existing methods use a single audio recording segment containing only noise to make a single estimate of the noise floor. Other existing methods analyze the entire audio recording that converges to a single low-level noise floor. However, a drawback of both of these methods is that they cannot adapt to varying noise levels or the spectrum. Other existing methods estimate the minimum envelope of energy in each frequency band and track the estimated minimum envelope over time (e.g., by smoothing the estimated minimum envelope using an appropriate time constant). However, these existing methods are typically used in real-time online audio signal processing architectures and cannot accurately react to sudden changes in noise in the audio recording. Summary of the Invention
[0006] An implementation method for adaptive noise estimation is disclosed.
[0007] In some embodiments, an adaptive noise estimation method includes: dividing an audio input into speech segments and non-speech segments using at least one processor; estimating a time-varying noise spectrum of the non-speech segment using the at least one processor for each frame in each non-speech segment; estimating a speech spectrum of the speech segment using the at least one processor for each frame in each speech segment; identifying one or more non-speech frequency components in the speech spectrum for each frame in each speech segment; comparing the one or more non-speech frequency components with one or more corresponding frequency components from a plurality of estimated noise spectra; and selecting an estimated noise spectrum from the plurality of estimated noise spectra based on the result of the comparison. In an embodiment, the method further includes: using the at least one processor to reduce noise in the audio input using the selected estimated noise spectrum.
[0008] In some embodiments, the method further includes: obtaining the probability of speech in each frame of the audio input, and identifying the frame as containing speech based on the probability.
[0009] In some embodiments, the time-varying noise spectrum is estimated by calculating a moving average of the power spectrum of the non-speech segment and averaging the power spectrum of the current non-speech segment and at least one past non-speech segment.
[0010] In some embodiments, during the non-speech segment, a time-varying estimated noise spectrum is fed to a noise reduction unit configured to reduce noise in the audio input using the selected estimated noise spectrum.
[0011] In some embodiments, for each speech segment, an estimated noise spectrum most likely representing noise in the current speech segment is determined using past estimated noise spectra preceding the speech segment, future estimated noise spectra following the speech segment, and the current speech frame.
[0012] In some embodiments, determining the estimated noise spectrum most likely representing the noise of the current speech segment further includes: obtaining an average noise spectrum from past noise spectra of past non-speech segments preceding the speech segment and future noise spectra of future non-speech segments following the speech segment, respectively; determining upper frequency limits for the past noise spectrum and the future noise spectrum; determining a cutoff frequency as the lowest of the two upper frequency limits; calculating a distance metric between frequency components in the speech spectrum and frequency components in the noise spectrum; and selecting the noise spectrum in the past noise spectrum or the future noise spectrum that has the smallest distance metric up to the cutoff frequency as the estimated noise spectrum of the audio input.
[0013] In some embodiments, the distance metric is averaged over a set of speech frames in a speech segment.
[0014] In some embodiments, speech components are estimated in speech segments of the audio signal, and then the speech components are subtracted from the actual speech components to obtain the remaining spectrum as an estimate of non-speech frequency components.
[0015] In some embodiments, an audio processor includes: a segmenter unit configured to segment an audio input into segments of overlapping frames; a plurality of buffers configured to store the segments of the overlapping frames; a spectrum analysis unit configured to calculate the spectrum of each segment stored in each buffer; a voice activity detector (VAD) configured to detect speech segments and non-speech segments in the audio input; and an averaging unit coupled to the output of the VAD and configured to calculate a speech spectrum for each speech segment identified by the VAD output and to calculate a noise spectrum for each non-speech segment identified by the VAD output.
[0016] In one embodiment, an audio processor includes: a VAD configured to detect speech segments and non-speech segments in an audio input; an averaging unit coupled to the output of the VAD and configured to obtain a speech spectrum for each speech segment identified by the VAD output and a noise spectrum for each non-speech segment identified by the VAD output; a similarity measurement unit configured to calculate a similarity measure between one or more frequency components in the current speech spectrum and one or more corresponding frequency components in each noise spectrum, and to select a noise spectrum from the noise spectrum based on the similarity measure; and a noise reduction unit configured to reduce noise in the audio input using the selected noise spectrum.
[0017] Other embodiments disclosed herein relate to a system, apparatus, and computer-readable medium. Details of the disclosed embodiments are set forth in the accompanying drawings and description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0018] The specific embodiments disclosed herein offer one or more of the following advantages. A method for adaptively estimating noise in an audio recording in the presence of speech is disclosed. In embodiments, adaptive noise estimation is performed offline on the audio recording to estimate noise variations by examining a given frame of the audio recording before and after it. An advantage over conventional adaptive noise estimation methods is that the noise floor beneath the speech is estimated by selecting from the best available candidate noise floor estimates computed before and after the current speech segment. Attached Figure Description
[0019] In the accompanying drawings, for ease of description, a specific arrangement or order of schematic elements is shown, such as those representing devices, units, instruction blocks, and data elements. However, those skilled in the art will understand that the specific order or arrangement of schematic elements in the drawings does not imply a requirement for a particular processing sequence or order, or process separation. Furthermore, the inclusion of schematic elements in the drawings does not imply that such elements are required in all embodiments, or that in some embodiments, features represented by such elements may not be included in or combined with other elements.
[0020] Furthermore, in the accompanying drawings, where connecting elements such as solid or dashed lines or arrows are used to illustrate connections, relationships, or associations between two or more other schematic elements, the absence of any such connecting element does not imply the impossibility of such connections, relationships, or associations. In other words, some connections, relationships, or associations between elements are not shown in the drawings to avoid obscuring the invention. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, where a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or more signal paths that may be necessary to influence the communication.
[0021] Figure 1 The diagram illustrates, according to some embodiments, a two-dimensional (2D) graph of an audio waveform, speech activity over time, and thresholds for determining non-speech segments of the audio waveform.
[0022] Figure 2 It is a 2D graph of speech activity over time according to some embodiments, a threshold for determining non-speech segments of the audio waveform, and noise segments of speech activity below the threshold.
[0023] Figure 3 The diagram illustrates the average speech spectrum corresponding to a speech segment and two noise spectra corresponding to non-speech segments preceding and following the speech segment, according to some embodiments.
[0024] Figure 4 This is a block diagram of a system for adaptive noise estimation and noise reduction according to some embodiments.
[0025] Figure 5 This is a flowchart of a process for background noise estimation and noise reduction according to some embodiments.
[0026] Figure 6 This is an implementation reference based on some embodiments. Figures 1 to 5 A block diagram of the system describing the features and processes.
[0027] The same reference numerals used in the various figures indicate similar elements. Detailed Implementation
[0028] In the following detailed description, several specific details are set forth to provide a thorough understanding of the various embodiments described. It will be apparent to those skilled in the art that the various embodiments described can be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Several features are described below, each of which can be used independently of each other or in any combination with other features.
[0029] Terminology Explanation
[0030] As used herein, the term “comprising” and variations thereof should be understood as open-ended terms meaning “including but not limited to”. Unless the context explicitly states otherwise, the term “or” should be understood as “and / or”. The term “based on” should be understood as “at least partially based on”. The terms “one example implementation” and “example implementation” should be understood as “at least one example implementation”. The term “another implementation” should be understood as “at least one other implementation”. The term “determine” should be understood as obtaining, receiving, calculating, estimating, predicting, or acquiring. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0031] System Overview
[0032] The disclosed embodiments use a Voice Activity Detection (VAD) classifier to segment the audio input into speech segments containing speech and non-speech segments not containing speech. In the non-speech segments, at each frame, the noise spectrum is estimated by averaging the energy of each frequency in the time region surrounding the current frame. In the speech segments, for each frame, the estimated noise spectrum of a previous or subsequent non-speech region in time is selected by identifying one or more non-speech frequency components in the speech spectrum. Using a similarity metric (e.g., distance between frequency components), one or more non-speech frequency components are compared to corresponding one or more frequency components in the estimated noise spectra of the previous and subsequent non-speech regions.
[0033] Figure 1 This is a two-dimensional (2D) graph illustrating an audio waveform, speech activity over time, and a threshold for determining non-speech segments of the audio waveform, according to an embodiment. For simplicity, the amplitude values of the audio waveform are not shown in the graph. Figure 1The diagram illustrates this. The horizontal axis represents time (e.g., milliseconds). Audio input (e.g., an audio file) containing audio recordings including speech is divided into overlapping frames. In this embodiment, VAD is used to obtain the probability of speech in each frame, and the audio input is subsequently divided into speech segments and non-speech segments based on a comparison of the speech probability with a threshold. In the example shown, the vertical axis represents the VAD value (the probability of speech being present), and the example VAD threshold indicated by the horizontal line is approximately 0.18. Figure 2 It shows Figure 1 The image shows a close-up of a noise segment in which the VAD value is below the VAD threshold.
[0034] Any suitable VAD algorithm for detecting speech segments and non-speech segments in an audio recording can be used, including but not limited to VAD algorithms based on: zero-crossing rate and energy measurement, linear energy detection, adaptive linear energy detection, pattern recognition, and statistical metrics.
[0035] In this embodiment, the noise spectrum in non-speech segments is estimated using Adaptive Speech-Aware Noise Estimation (AVANE) and by inferring the most similar robust noise estimate for the speech segment. AVANE calculates a moving average of the power spectrum of non-speech frames, and for each non-speech frame, the power spectrum of noise in the non-speech frame is calculated by averaging the power of the current non-speech frame and one or more past non-speech frames. In this embodiment, the number of past frames to be averaged is determined by a time constant. Any suitable moving average algorithm can be used, including but not limited to: arithmetic mean, exponential mean, smoothed mean, and weighted moving average.
[0036] AVANE generation uses a time-varying noise spectrum in two ways. First, during non-speech segments, the time-varying estimated noise is fed (e.g., buffer-by-buffer) to the noise reduction system. Second, during speech segments, the last AVANE estimate before the current speech segment and the first AVANE estimate after the current speech segment are fed together with the current speech frame to the inference component. The inference component determines which AVANE estimate is most likely to represent the noise in the current speech frame.
[0037] Alternative methods for AVANE estimation include, for example, spectral minima tracking in subbands as described in Doblinger, G. (1995), Computationally efficient speech enhancement by spectral minima tracking in subbands, Proc. EUROSPEECH '95, Madrid, pp. 1513-1516; or, for example, noise power spectral density estimation based on optimal smoothing and minimum statistics as described in Martin, R. (2001), Noise power spectral density estimation based on optimal smoothing and minimum statistics, IEEE Transactions on Speech and Audio Processing, 9(5) 504-512.
[0038] Two embodiments are proposed for estimating the underlying noise spectrum of a given speech segment. In the first embodiment, the speech components are estimated, and then the estimated speech components are subtracted from the actual speech components to obtain the residual spectrum as a noise estimate. This embodiment results in a direct estimate of the background noise, and is therefore unrelated to or not combined with AVANE. Assuming that the speech is dominated by harmonic components, the fundamental tone is estimated first, and then the harmonic components are identified. Based on a sine model and its parameter estimates, the harmonic components are subtracted from the speech signal to obtain the residual signal. This method is described, for example, in Stylianou, Y. (1996), Harmonic plus Noise Models for Speech combined with Statistical Methods for Speech and Speaker Modification, PhD Thesis, Telecom Paris. Another possibility is to identify and subtract a sine curve from a given short-time spectrum without fundamental frequency (F0) information. This method is described, for example, in Yeh, C. (2008), Multiple Fundamental Frequency Estimation of Polyphonic Recordings, Ph.D. thesis, University Paris.
[0039] In another embodiment, harmonic components are estimated and attenuated in the cepstral domain, as described, for example, in Z. Zhang, K. Honda, and J. Wei, "Retrieving Vocal-Tract Resonance and Anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique," ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 7359-7363, doi:10.1109 / ICASSP40776.2020.9054741.
[0040] The AVANE method assumes that the underlying noise spectrum is closer to the last AVANE before the speech segment or the first AVANE after the speech segment. In this embodiment, spectral segments where speech (e.g., high frequencies) is not dominant are identified, and a spectral similarity metric (e.g., a distance metric) between the speech spectrum and the AVANE is calculated, considering only the non-speech segments with the main noisy components of the spectrum. In this embodiment, the spectral similarity metric is based on the distance between the speech spectrum and the AVANE. Assuming a positive signal-to-noise ratio (SNR) (defined as the ratio of speech energy to noise energy in the speech band, expressed in decibels), further constraints can be added so that the selected AVANE is accepted only if the average speech spectrum in the region of interest (over a duration of the same length as the AVANE) is higher than the selected AVANE.
[0041] In embodiments where harmonic subtraction is used to calculate noise spectrum estimates, the spectral similarity measure may not be limited to the non-speech frequency region of the speech spectrum, but may be extended to the entire spectrum, or limited to frequencies above a specific speech frequency, such as the lowest frequency range in which harmonic estimation is effective for speech. Therefore, a similarity measure is calculated between the residual signal after subtracting harmonics from a speech segment and the AVANE estimates before and after the speech segment.
[0042] In an embodiment, given an audio frame, the energy spectrum of the audio frame is calculated and converted to a decibel scale. When the current audio frame is a speech frame (i.e., within a speech segment), the previously calculated average noise spectrum (in dB) before and after the speech segment is obtained from, for example, storage (e.g., memory, optical disc). Figure 3 The average speech spectrum and two noise spectra corresponding to non-speech segments before and after the speech segment are shown according to some embodiments.
[0043] Given two noise spectra and the current speech spectrum, calculate the upper frequency limit f of the noise spectrum. c The lower of the two upper limits is retained as the "cutoff" frequency f. cutof Next, a similarity metric is calculated across a segment from, for example, half the audio spectrum to a cutoff frequency (in this example, the sum of the absolute values of the differences (“distances”) between the speech spectrum and the two noise spectra). The noise spectrum with the smallest (as defined before) distance is retained as the current estimate of the noise spectrum of the audio recording. In an alternative embodiment, the distance metric can be calculated and averaged over a set of speech frames, and the noise spectrum that gives the lowest average distance can be selected as the current estimate of the noise spectrum.
[0044] Assume that audioframe is a vector of audio samples within a frame, and spectrum is the spectrum of the audio samples computed using the Fast Fourier Transform (FFT) of the audioframe:
[0045] spectrum=fft(audioframe). [1]
[0046] The spectrum can be converted to a dB scale using the following formula. dB :
[0047] spectrum dB =20log 10 (abs(spectrum)). [2]
[0048] If the current frame is a noisy frame, preserve its spectrum. dB And averaged with the past spectrum over a window of a given length (e.g., 5 seconds), hereinafter referred to as avg_spectrum. dB If the current frame is a speech frame, its spectrum is compared with past noise spectra and future noise spectra. In the following text, the speech spectrum is referred to as the speech_spectrum. dB The past noise spectrum and the future noise spectrum are respectively referred to as past_spectrum. dB and future_speech dB .
[0049] In some embodiments, past_spectrum dB and future_spectrum dB The upper frequency limit f of each of them c Determine by: 1) Selecting the range above which f is estimated. c1) The first frequency; 2) Divide the noise spectrum above the first frequency into blocks with a specified length and overlap (e.g., 50%); 3) In each block, calculate the average derivative, sorted in ascending order of the frequency of its corresponding block, and find the first derivative with a value less than a predefined negative value (e.g., -20dB); and 4) Calculate f c The average noise spectrum of the previous small area is used, and the average noise spectrum is used instead of f. c The above values are for the noise spectrum. Note that step (3) should be interpreted as a significant reduction in the noise spectrum, and the frequency of the corresponding block is considered to be the upper frequency limit.
[0050] Given a cutoff frequency f cutoff As the determined upper frequency limit f c The distance between the current speech spectrum and the noise spectrum is calculated as follows: (The lower of the two frequencies and above the speech frequency f1 are used.)
[0051]
[0052]
[0053] noise_spectrum selected =argmin(distance_past,distance_future). [4]
[0054] As shown in equation [4], f1 and f cutoff The frequency range between these values defines a spectral region where speech harmonics are almost nonexistent and background noise dominates. The minimum of distance_past and distance_future (given by argmin()) yields a noise spectrum that is closer to the current spectrum and is selected as a noise candidate. This method can be extended to multiple candidate noise spectra.
[0055] Note that in embodiments using harmonic subtraction to estimate and remove speech harmonics, the method described in equations 3a, 3b, and 4 can be extended to speech frequencies by replacing the starting index f1 with a lower frequency index (e.g., the lowest frequency of the speech or the lowest frequency at which the residual estimate is considered reliable).
[0056] Note that any method described herein capable of estimating noise in the presence of speech (e.g., the AVANE method) can be used to calculate the distance between the estimated spectrum and two known noise spectra by comparing the current frame with the estimate obtained from AVANE in adjacent non-speech segments and selecting either the past noise estimate or the future noise estimate.
[0057] Figure 4This is a block diagram of a system 400 for adaptive noise estimation and noise reduction according to an embodiment. A segmenter unit 401 segments an audio input (e.g., an audio file containing speech content) into overlapping frame segments and stores the resulting segments in multiple buffers 402, transforming the segments into a spectrum 405 via, for example, a Short Time Fourier Transform (STFT) block 403. A Voice Activity Detection (VAD) block 404 calculates the probability that a given audio frame contains speech. The spectrum 405 and the VAD output (speech probability) are fed to an averaging unit 406, which generates a current speech spectrum and multiple noise spectra 407 for each speech frame. The speech spectrum and the multiple noise spectra 407 are input to a similarity measurement unit 408, which selects one of the noise spectra (e.g., a distance metric based on equations [3a, 3b]) as the noise spectrum 410 to be used by the noise reduction block 409 to reduce noise in the audio input.
[0058] In some embodiments, the noise reduction unit 409 uses a selected noise spectrum 410 to reduce noise in the audio input by comparing the spectrum of the audio input with the selected noise spectrum 410 and applying gain reduction to those frequency bands where the energy of the input signal is less than the energy of the noise spectrum plus a predefined threshold.
[0059] Other embodiments
[0060] The following description of further embodiments focuses on the differences between the further embodiments and the previously described embodiments. Therefore, features common to both embodiments will be omitted from the following description, and it should therefore be assumed that features of the previously described embodiments can, or at least can, be implemented in the further embodiments, unless otherwise required by the following description of the further embodiments.
[0061] In some embodiments, multiple pre-calculated noise spectra are available.
[0062] noise_spectrum i , where i = 1,..,N, [5]
[0063] Furthermore, the similarity metric is the distance between the current speech spectrum and multiple noise spectra (scaled in dB), given by the following formula:
[0064]
[0065] The noise spectrum corresponding to the smaller distance was selected as:
[0066] noise_spectrum K Where K = argmin(distace) i [7]
[0067] Multiple noise spectra can be provided a priori, for example in applications where different noise conditions found in audio recordings are known and measured in advance, such as in a teleconference with multiple endpoints. Alternatively, multiple noise spectra can be determined by a clustering algorithm applied to multiple spectra of non-speech frames. The clustering algorithm can be, for example, k-means clustering applied to multiple non-speech spectrum vectors, or any other suitable clustering algorithm.
[0068] Online Implementation Examples
[0069] The above-described embodiments for offline computation can be extended to real-time, online, low-latency scenarios. Note that in this case, future noise spectra following the current speech frame cannot be used. When candidate noise spectra are provided a priori, a selection process is applied online for each speech frame using the available (stored) noise spectra. When candidate noise spectra are not provided a priori, noise spectra can be constructed online. For example, a first noise spectrum is obtained from a first non-speech frame. Upon receiving an additional non-speech frame, if its distance from each previously retained noise spectrum is greater than a predefined threshold, the noise spectrum of the non-speech frame is calculated and retained as an additional noise spectrum. Upon receiving an additional non-speech frame, the noise spectrum of the non-speech frame is calculated and clustered using a clustering algorithm (e.g., k-means clustering), and the resulting clusters are used as candidate noise spectra. The clustering process is repeated and improved each time a sufficient number of new non-speech frames are received, or each time a non-speech frame with a significant difference from existing clusters is received.
[0070] Music recording
[0071] In this embodiment, the audio recording includes music (or another type of audio content) rather than speech content. In this embodiment, the speech classifier VAD is replaced by a suitable music (or another type of) classifier.
[0072] Music plus voice recording
[0073] In this embodiment, the audio recording includes both speech and music. In this embodiment, it is desirable to remove noise from the speech and music portions while preserving the music signal. In this embodiment, the speech classifier is replaced by a multi-class classifier (e.g., a music and speech classifier) or two separate classifiers for music and speech. The probabilities of speech and music output by the classifiers are compared to predefined thresholds; when both the speech and music probabilities are less than the predefined thresholds, the frame is considered noise. The previously described method is then applied to estimate a suitable noise spectrum for the speech region and optionally also for the music region.
[0074] Example process
[0075] Figure 5 This is a flowchart of a process 500 for background noise estimation and noise reduction according to an embodiment. Process 500 can be used... Figure 6 The device architecture shown is used for implementation.
[0076] Process 500 begins by dividing the audio input into speech segments and non-speech segments (501), and for each frame in each non-speech segment, estimating the time-varying noise spectrum of the non-speech segment (503) and the speech spectrum of the speech segment (504).
[0077] The process 500 continues to: for each frame in each speech segment, identifying one or more non-speech frequency components in the speech spectrum (505), comparing the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra (506); and selecting an estimated noise spectrum from the plurality of estimated noise spectra based on the comparison results (507).
[0078] In some embodiments, the multiple estimated noise spectra include estimated noise spectra of past non-speech segments and estimated noise spectra of future non-speech segments. In some embodiments, the multiple estimated noise spectra can be determined by a clustering algorithm applied to multiple noise spectra of non-speech frames. The clustering algorithm can be, for example, k-means clustering applied to multiple non-speech spectrum vectors, or any other suitable clustering algorithm.
[0079] In some embodiments, process 500 may continue to reduce noise in the audio input using the selected estimated noise spectrum.
[0080] Example System Architecture
[0081] Figure 6 An implementation reference according to an embodiment is shown. Figures 1 to 5 A block diagram of an example system describing the features and processes. System 600 includes any device capable of playing audio, including but not limited to: smartphones, tablet computers, wearable computers, in-vehicle computers, game consoles, surround sound systems, and kiosks.
[0082] As shown, system 600 includes a central processing unit (CPU) 601 capable of executing various processes based on a program stored, for example, in read-only memory (ROM) 602 or a program loaded from, for example, storage unit 608 into random access memory (RAM) 603. RAM 603 also stores, as needed, data required by CPU 601 when executing various processes. CPU 601, ROM 602, and RAM 603 are interconnected via bus 609. Input / output (I / O) interface 605 is also connected to bus 604.
[0083] The following components are connected to I / O interface 605: input unit 606, which may include a keyboard, mouse, etc.; output unit 607, which may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 608, which includes a hard disk or other suitable storage device; and communication unit 609, which includes a network interface card such as a network card (e.g., wired or wireless).
[0084] In some implementations, the input unit 606 includes one or more microphones located at different positions (depending on the host device), which enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0085] In some embodiments, the output unit 607 includes a system having a variety of numbers of speakers. For example... Figure 6 As illustrated, the output unit 607 (depending on the capabilities of the host device) can present audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0086] Communication unit 609 is configured to communicate with other devices (e.g., via a network). Drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted on drive 610 so that computer programs read from it are installed into storage unit 608. Those skilled in the art will understand that although system 600 is described as including the components described above, in practice, some of these components may be added, removed, and / or replaced, and all such modifications or changes fall within the scope of this disclosure.
[0087] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 609, and / or installed from removable medium 611, such as... Figure 6 As shown.
[0088] Typically, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be implemented by control circuitry (e.g., with...) Figure 6The control circuitry executes the actions described herein, which are performed by the CPU in combination with other components. Some aspects may be implemented in hardware, while others may be implemented in firmware or software (e.g., control circuitry) that can be executed by a controller, microprocessor, or other computing device. Although various aspects of exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, other computing devices, or some combination thereof, as non-limiting examples.
[0089] Furthermore, the various boxes shown in the flowchart can be viewed as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform associated functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to perform the methods described above.
[0090] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0091] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that when executed by the processor of the computer or the processor of the other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0092] The listed example implementation (EEE)
[0093] Embodiments of this disclosure may relate to one of the listed embodiments (EEE) below.
[0094] EE1 is an audio processor comprising: a segmenter unit configured to segment an audio input into segments of overlapping frames; a plurality of buffers configured to store the segments of the overlapping frames; a spectrum analysis unit configured to calculate the spectrum of each segment stored in each buffer; a voice activity detector (VAD) configured to detect speech segments and non-speech segments in the audio input; an averaging unit coupled to the output of the VAD and configured to calculate a speech spectrum for each speech segment identified by the VAD output and to calculate a noise spectrum for each non-speech segment identified by the VAD output; a similarity measurement unit configured to calculate a similarity measure between one or more frequency components in the current speech spectrum and each noise spectrum, and to select a noise spectrum from the plurality of noise spectra based on the similarity measure; and a noise reduction unit configured to reduce noise in the audio input using the selected noise spectrum.
[0095] EEE2 is an audio processor as described in EEE1, wherein the VAD is configured to obtain the probability of speech in each frame of the audio input and to identify the frame as containing speech based on the probability.
[0096] EEE3 is an audio processor comprising: a voice activity detector (VAD) configured to detect speech segments and non-speech segments in the audio input; an averaging unit coupled to the output of the VAD and configured to obtain a speech spectrum for each speech segment identified by the VAD output and a noise spectrum for each non-speech segment identified by the VAD output; a similarity measurement unit configured to calculate a similarity measure between one or more frequency components in the current speech spectrum and one or more corresponding frequency components in each noise spectrum, and to select a noise spectrum from the noise spectrum based on the similarity measure; and a noise reduction unit configured to reduce noise in the audio input using the selected noise spectrum.
[0097] While this document contains numerous details of specific implementation, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Specific features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof. The logical flow depicted in the drawings does not require the specific order or ordered sequence shown to achieve the desired result. Additionally, other steps may be provided from the described flow, or steps may be removed, and other components may be added to or removed from the described system. Therefore, other implementations are within the scope of the appended claims.
Claims
1. An adaptive noise estimation method, comprising: Use at least one processor to divide the audio input into speech segments and non-speech segments; For each frame in each non-speech segment, the time-varying noise spectrum of the non-speech segment is estimated using the at least one processor; For each frame in each speech segment, the speech spectrum of the speech segment is estimated using the at least one processor; For each frame in each speech segment Identify one or more non-speech frequency components in the speech spectrum; The one or more non-speech frequency components are compared with one or more corresponding frequency components in a plurality of estimated noise spectra to obtain a distance metric for each noise spectrum; as well as The noise spectrum estimated with the minimum distance metric is selected as the noise spectrum estimated for the speech segment.
2. The method as described in claim 1, wherein, The multiple estimated noise spectra include estimated noise spectra of past non-speech segments and estimated noise spectra of future non-speech segments.
3. The method of claim 1 or 2, further comprising: Using the at least one processor, noise in the audio input is reduced using a selected estimated noise spectrum.
4. The method of claim 1 or 2, further comprising obtaining the probability of speech in each frame of the audio input, and identifying frames containing speech based on the probability.
5. The method as described in claim 1 or 2, wherein, The time-varying noise spectrum is estimated by calculating a moving average of the power spectrum of the non-speech segment and averaging the power spectrum of the current non-speech segment and at least one past non-speech segment.
6. The method as described in claim 1 or 2, wherein, During the non-speech segment, a time-varying estimated noise spectrum is fed to a noise reduction unit configured to reduce noise in the audio input using the selected estimated noise spectrum.
7. The method of claim 2, wherein, For each speech segment, the estimated noise spectrum that most likely represents the noise in the current speech segment is determined using past estimated noise spectrum before the speech segment, future estimated noise spectrum after the speech segment, and the current speech frame.
8. The method of claim 7, wherein, The noise spectrum for determining the estimated noise most likely representing the current speech segment also includes: The average noise spectrum is obtained from the past noise spectrum of past non-speech segments preceding the speech segment and the future noise spectrum of future non-speech segments following the speech segment, respectively. Determine the upper frequency limits of the past noise spectrum and the future noise spectrum; The cutoff frequency is determined to be the lowest of the two upper frequency limits; Calculate the distance metric between the frequency components in the speech spectrum and the frequency components in the noise spectrum; and The noise spectrum with the smallest distance metric up to the cutoff frequency from either the past noise spectrum or the future noise spectrum is selected as the estimated noise spectrum of the audio input.
9. The method of claim 8, wherein, The distance metric is averaged over a set of speech frames in a speech segment.
10. The method as claimed in claim 1 or 2, wherein, Speech components are estimated from the speech segment of the audio signal, and then the speech components are subtracted from the actual speech components to obtain the remaining spectrum as estimated non-speech frequency components.
11. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform the operation as described in any one of claims 1-9 of the preceding method.
12. An audio processor, comprising: The divider unit is configured to divide the audio input into speech segments and non-speech segments; The averaging unit is configured to estimate the speech spectrum for each speech segment and the time-varying noise spectrum for each non-speech segment; The similarity metric unit is configured as follows: Identify one or more non-speech frequency components in the speech spectrum; The one or more non-speech frequency components are compared with one or more corresponding frequency components in a plurality of estimated noise spectra to obtain a distance metric for each noise spectrum; as well as The noise spectrum estimated with the minimum distance metric is selected as the noise spectrum estimated for the speech segment.
13. The audio processor of claim 12, wherein, The multiple estimated noise spectra include estimated noise spectra of past non-speech segments and estimated noise spectra of future non-speech segments.
14. The audio processor of claim 12 or 13, further comprising: The noise reduction unit is configured to reduce noise in the audio input using a selected estimated noise spectrum.
15. The audio processor of claim 14, wherein, During the non-speech segment, the noise reduction unit is configured to receive the non-speech segment and reduce the noise in the audio input using a selected estimated noise spectrum.
16. The audio processor of claim 14, wherein, The noise reduction unit is configured to use a selected estimated noise spectrum to reduce noise in the audio input by comparing the spectrum of the audio input with the selected estimated noise spectrum and applying gain reduction to a frequency band where the energy of the audio input is less than the energy of the noise spectrum plus a predefined threshold.
17. The audio processor of claim 12 or 13, wherein, A voice activity detector (VAD) is configured to obtain the probability of speech in each frame of the audio input and to identify frames containing speech based on the probability.
18. The audio processor as claimed in claim 12 or 13, wherein, The averaging unit is configured to estimate the time-varying noise spectrum by calculating a moving average of the power spectrum of the non-speech segment and averaging the power spectrum of the current non-speech segment and at least one past non-speech segment.
19. The audio processor as claimed in claim 12 or 13, wherein, For each speech segment, the similarity measurement unit is configured to determine the estimated noise spectrum that most likely represents the noise in the current speech segment based on past estimated noise spectra preceding the speech segment, future estimated noise spectra following the speech segment, and the current speech frame.
20. The audio processor of claim 19, wherein, The similarity measurement unit is configured to determine the estimated noise spectrum that most likely represents the noise of the current speech segment by: The average noise spectrum is obtained from the past noise spectrum of past non-speech segments preceding the speech segment and the future noise spectrum of future non-speech segments following the speech segment, respectively. Determine the upper frequency limits of the past noise spectrum and the future noise spectrum; The cutoff frequency is determined to be the lowest of the two upper frequency limits; Calculate the distance metric between the frequency components in the speech spectrum and the frequency components in the noise spectrum; as well as The noise spectrum with the smallest distance metric up to the cutoff frequency from either the past noise spectrum or the future noise spectrum is selected as the estimated noise spectrum of the audio input.
21. The audio processor of claim 20, wherein, The similarity measurement unit is configured to average the distance measurement over a set of speech frames in a speech segment.
22. The audio processor of claim 12, wherein, The similarity measurement unit is configured to estimate one or more speech components in the speech segment of the audio input, and then subtract the one or more estimated speech components from the actual speech components to obtain the remaining spectrum as the estimated non-speech spectrum.
Citation Information
Patent Citations
Apparatus and method for removing noise of audio signalin portable terminal
KR1020070108598A
Noise estimation using an adaptive smoothing factor based on a teager energy ratio in a multi-channel noise suppression system
US20110099007A1