Content-aware audio noise management
By segmenting and estimating audio signals through a content-aware audio noise management system, the problem of oversuppression caused by the difficulty in distinguishing between speech, music and noise in existing technologies is solved, thereby improving the quality and signal-to-noise ratio of audio signals.
Patent Information
- Application Number
- CN202480047311.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-22
- Filing Date
- 2024-06-19
- Publication Date
- 2026-02-13
AI Technical Summary
Existing noise reduction technologies struggle to effectively distinguish between speech, music, and noise when processing audio signals with mixed content types, leading to excessive suppression of useful signals and impacting audio quality.
A content-aware audio noise management system is adopted. Through audio event segmentation and noise estimation, content-aware segmentation information is generated, and frequency bin data is used for noise suppression. The suppression gain is adjusted according to different audio content types to reduce the impact on useful signals.
It achieves more precise noise suppression in the noise reduction process of audio signals with mixed content types, improves audio quality and signal-to-noise ratio, and reduces interference with useful signals.
Smart Images

Figure CN121532827A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Various example embodiments relate to audio signal processing and enhancement. BACKGROUND
[0002] The material described in this section is not prior art to the claims in this application and is not admitted to be prior art by virtue of its inclusion in this section.
[0003] Noise reduction processing generally aims to significantly reduce background and / or broadband noise while minimizing the impact on audio signal quality as much as possible. Such processing can be configured to remove various noise components, including but not limited to tape hiss, microphone background noise, hum, wind noise, etc. The appropriate amount of noise reduction generally depends on the type of useful audio signal, the type of noise, and the acceptable loss of useful audio signal. In some cases, noise reduction processing can increase the signal-to-noise ratio (SNR) by 5 dB to 20 dB while maintaining high audio quality of the remaining useful audio signal.
[0004] With these and other considerations in mind, the disclosure presented herein was made. SUMMARY
[0005] Techniques for processing audio signals are described. Examples found herein provide systems, apparatuses, devices, and methods for performing content-aware audio noise management. Some embodiments can be beneficially used to automatically manage floor noise in mixed content types (e.g., including speech, music, and noise sounds). At least some embodiments can be applicable to applications such as real-time communication, content creation, and audio capture of user devices.
[0006] According to example embodiments, a noise management method is provided that includes performing audio event segmentation on an input audio signal to generate content-aware segmentation information, estimating a floor noise level associated with the input audio signal based on the content-aware segmentation information and further based on ranking frequency bin data corresponding to fixed length portions of the input audio signal, and applying noise suppression to the input audio signal to generate an output audio signal, the noise suppression being performed in frequency bins having a selected frequency resolution and based at least in part on the content-aware segmentation information and the estimated floor noise level.
[0007] In some embodiments of the above method, the audio event segmentation is performed in real time.
[0008] In some embodiments of any of the above methods, the includes processing frame-by-frame speech, music, and noise confidence to identify a continuous sequence of audio events corresponding to the input audio signal, the audio events selected from a plurality of event clusters, each event cluster in the plurality of event clusters associated with one or more classes selected from the group consisting of a vocal class, a music class, and a noise class.
[0009] In some embodiments of any of the above methods, the performing includes performing a moving average convergence-divergence process on the frame-by-frame speech, music, and noise confidence.
[0010] In some embodiments of any of the above methods, the estimating includes updating the fixed-length portion using a first-in-first-out buffer configured to receive the frequency bin data.
[0011] In some embodiments of any of the above methods, the estimating includes truncating the sorted frequency bin data to a size determined by an adaptive percentile estimator based on the content-aware segmentation information.
[0012] In some embodiments of any of the above methods, the adaptive percentile estimator is configured to select different respective percentiles for speech events, music events, and noise events, the determined size based on a selected percentile of the different respective percentiles.
[0013] In some embodiments of any of the above methods, the applying includes computing an attenuation gain value based on the computed per-bin signal-to-noise ratio values and further based on the estimated noise floor level.
[0014] In some embodiments of any of the above methods, for each of the frequency bins, the computing includes: computing a respective first attenuation gain value based on a respective one of the per-bin signal-to-noise ratio values; computing a respective second attenuation gain value based on the respective first attenuation gain value and further based on a respective one of the estimated noise floor levels; and computing a product of the respective first attenuation gain value and the respective second attenuation gain value.
[0015] In some embodiments of any of the above methods, the applying includes applying different respective aggressiveness of noise attenuation to different types of audio content based on the content-aware segmentation information.
[0016] In some embodiments of any of the above methods, the method further includes: converting the input audio signal to frequency bin data using a Fourier transform; and applying an inverse Fourier transform to the frequency bins to generate the output audio signal after applying the noise attenuation to the frequency bins.
[0017] According to another example embodiment, there is provided a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any of the methods described above.
[0018] According to yet another example embodiment, there is provided a noise management apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and program code are configured, with the at least one processor, to cause the apparatus at least to: perform audio event segmentation on an input audio signal to generate content-aware segmentation information; estimate a noise floor level associated with the input audio signal based on the content-aware segmentation information and further based on ranking frequency bin data corresponding to fixed length portions of the input audio signal; and apply noise suppression to the input audio signal to generate an output audio signal, the noise suppression being performed in frequency bins having a selected frequency resolution and based at least in part on the content-aware segmentation information and the estimated noise floor level.
[0019] Various aspects of the present disclosure relate to audio signal processing and provide improvements at least in the technical fields of audio processing, audio event classification and segmentation, noise estimation, noise suppression, etc.
[0020] Some embodiments disclosed herein can be generally described as technology, where the term “technology” can refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested in the context presented below.
[0021] Features and benefits other than those explicitly described above will become apparent from a reading of the following specification and a study of the drawings in which benefits of the present disclosure will become fully apparent. This Summary of the Invention is provided to introduce a selection of concepts in a simplified form and is not intended to identify key or essential features of the claimed subject matter lest the scope thereof. BRIEF DESCRIPTION OF DRAWINGS
[0022] These and other more detailed and specific features of various embodiments are fully disclosed in the following description, in which references are made to the accompanying drawings, which illustrate in: Figure 1 is a block diagram illustrating an example content-aware noise management system in which various aspects of the present disclosure can be practiced.
[0023] Figure 2 is a block diagram illustrating an example content-aware noise management system according to some aspects of the present disclosure. Figure 1 is a block diagram illustrating an example circuit implementation of the content-aware noise management system of
[0024] Figure 3is a block diagram illustrating a workflow implemented in an audio event classification and segmentation component of a content-aware noise management system in accordance with some aspects of the present disclosure. Figure 1
[0025] Figure 4 is a pictorial diagram illustrating an audio event clustering used in a workflow of a content-aware noise management system in accordance with some aspects of the present disclosure. Figure 3
[0026] Figure 5 is a flow diagram illustrating a noise management method in accordance with some aspects of the present disclosure.
[0027] Figure 6A illustrates a schematic block diagram of an example device architecture that can be used to implement various aspects of the present disclosure.
[0028] Figure 6B illustrates a schematic block diagram of an example CPU that can be used to implement various aspects of the present disclosure implemented in a device architecture of Figure 6A DETAILED DESCRIPTION
[0029] In the following description, numerous specific details are set forth to provide an understanding of one or more aspects of the present disclosure, such as audio device configurations, timing, operations, etc. It will be apparent, however, to one skilled in the art that these specific details are merely examples and that other aspects can be practiced without such specific details.
[0030] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless expressly indicated otherwise. Such terms are to be construed in a non-limiting fashion. For example, a reference to “A and / or B” can mean “A and B” or “A or B,” or “A and B and / or other items.” As another example, a reference to “A, B, and / or C” can mean “A, B, and C,” or “A, B, or C,” or “A, B, C, and / or other items.” As yet another example, a reference to “at least one of A, B, or C” can mean “at least one of A, at least one of B, or at least one of C,” or “at least one of A, at least one of B, at least one of C, and / or other items.” As used herein, the term “based on” is to be read as “based, at least in part, on.” As used herein, the terms “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” As used herein, the term “another implementation” is to be read as “at least one other implementation.” As used herein, the term “determining” is to be read as obtaining, receiving, calculating, estimating, estimating, predicting, or deriving. In addition, throughout the description and claims of this disclosure, the following terminology should be understood as having the following meaning unless otherwise expressly defined herein. Solely for the purposes of the following definitions, the following terms shall be construed as follows.
[0031] Various acronyms that can appear throughout this disclosure and associated claims and / or drawings are listed as follows. Other commonly used acronyms and technical terms can be excluded from this list for brevity. Accordingly, a short list of acronyms is provided below for the reader’s quick reference.
[0032] MACD - Moving Average Convergence Divergence.
[0033] FIFO - First In, First Out.
[0034] SNR - Signal to Noise Ratio.
[0035] FFT - Fast Fourier Transform.
[0036] IFFT - Inverse FFT.
[0037] Noise suppression is an audio process that removes background noise from captured signals. Millions of internet-connected devices are used in increasingly diverse everyday life scenarios that involve audio capture, such as calls, user content capture / creation, virtual music lessons, and live streaming of new or evolving forms of media content (e.g., podcasts, audiobooks, extended reality content, virtual worlds, etc.). In these scenarios, noise management systems that implement noise suppression are expected to provide not only clear speech, but also effectively enhance richer sound experiences (such as high-fidelity music, virtual meetings, live streaming, user-generated content, etc.).
[0038] Speech-centric noise management algorithms designed specifically for speech can disadvantageously result in over-suppression effects when applied to music or other audio signals with useful non-speech components. Accordingly, example embodiments disclosed herein provide content-aware audio noise management systems and methods that can be beneficially used, for example, to manage background noise in mixed content types. At least some embodiments can be applicable to applications such as real-time communication, content creation, and audio capture of user devices.
[0039] Figure 1 is a block diagram illustrating an example content-aware noise management system 100 in which various aspects of the present disclosure can be practiced. The content-aware noise management system 100 includes an audio event classification and segmentation component 130. The audio event classification and segmentation component 130 receives an input audio signal 102 and processes the input audio signal 102 to generate content-aware segmentation information. The content-aware noise management system 100 also includes a content-aware noise estimation component 110 and a content-aware noise suppression component 120. The audio event classification and segmentation component 130 is connected to the content-aware noise estimation component 110 and the content-aware noise suppression component 120 via an information path 128 through which the content-aware segmentation information is provided to the content-aware noise estimation component and the content-aware noise suppression component. The content-aware noise estimation component 110 and the content-aware noise suppression component 120 are also connected to each other via an intermediate path 112. In some examples, the audio event classification and segmentation component 130 generates the content-aware segmentation information in real-time.
[0040] As used herein, the term “real-time” refers to a computer-based process that controls or monitors a corresponding environment by receiving data, processing the received data, and generating a response that is fast enough to affect or characterize the environment without producing significant delay. In the context of control or processing software, real-time response is generally understood to be on the order of milliseconds, or sometimes microseconds.
[0041] The content-aware noise estimation component 110 receives the input audio signal 102 and processes the received input audio signal 102 based on the content-aware segmentation information to generate intermediate output signals, which are then provided to the content-aware noise suppression component 120 via path 112. The content-aware noise suppression component 120 receives the intermediate output signals and processes the received intermediate output signals based on the content-aware segmentation information to generate the output audio signal 122. In some examples, the processing implemented in the content-aware noise estimation component 110 and the content-aware noise suppression component 120 are configured to generate the output audio signal 122 in which noise suppression is performed in frequency subbands (bins) having a selected frequency resolution and is based at least in part on the content-aware segmentation information. The estimated noise data is provided in the intermediate output signals transmitted via path 112.
[0042] Figure 2 is a block diagram illustrating an example circuit implementation of a content-aware noise management system 100 in accordance with some aspects of the present disclosure. In the illustrated example, the content-aware noise management system 100 includes a fast Fourier transform (FFT) module 202. The FFT module 202 is a shared module for the content-aware noise estimation component 110 and the content-aware noise suppression component 120. In other implementations, the FFT module 202 can be replaced by another suitable time-to-frequency domain converter, such as a filter bank. In the illustrated example, the FFT module 202 is connected to the content-aware noise estimation component 110 and downstream modules of the audio event classification and segmentation component 130 via an input bus 204.
[0043] In operation, the FFT module 202 receives the input audio signal 102 and converts the received input audio signal 102 into corresponding frequency domain signals, which are transmitted downstream via bus 204. In some examples, the frequency domain signals are represented by a plurality of frequency bins. Example conversion operations performed in the FFT module 202 include: (i) dividing the audio signal waveform provided via the input audio signal 102 into windowed, semi-overlapping frames; and (ii) applying a Fourier transform to each frame to generate a corresponding bin spectrum block. The magnitudes of the contents of the spectrum blocks are collectively referred to as where, where, denotes a frequency index, and denotes a time index. The quantity is referred to as the frame of the frequency bin corresponding to time t.
[0044] In the illustrated example, the audio event classification and segmentation component 130 also includes a multi-classification module 240 and an energy module 244, both of which receive the frequency bins via the bus 204. The audio event classification and segmentation component 130 further includes an audio event segmentation module 248. The multi-classification module 240 and the energy module 244 are connected to the audio event segmentation module 248 via a first path 242 and a second path 246, respectively. In operation, the multi-classification module 240 processes the frequency domain signal (on the bus 204) to generate a corresponding stream of classification information, which is fed to the audio event segmentation module 248 via the first path 242. The energy module 244 processes the frequency domain signal (on the bus 204) to generate a corresponding stream of energy information, which is fed to the audio event segmentation module 248 via the second path 246. The audio event segmentation module 248 processes the stream of classification information received via the first path 242 and the stream of energy information received via the second path 246 to generate the aforementioned content-aware segmentation information.
[0045] In one example, the frequency bins corresponding to the current audio frame and a sequence of adjacent historical audio frames are processed in the multi-classification module 240 to generate corresponding classification information, which includes a speech confidence , a music confidence , and a noise confidence , as follows: (1) where denotes the total number of bins, and denotes the length of the historical frames used for classification purposes. Herein, the term "noise" refers to background noise, such as indoor noise, fan noise, or electrical noise, which does not include transient noise. The audio event segmentation module 248 uses the aforementioned confidence and energy information to generate content-aware segmentation information using a suitable segmentation function , for example, as follows: (2) where , , denote the sound energy, music energy, and noise energy, respectively. In one example, the segmentation function is based on a moving average convergence / divergence (MACD) process, which will be described in more detail below, for example, with reference to Figure 3 and Equations (16) to (24).
[0046] In the example shown, the content-aware noise estimation unit 110 also includes a serially connected First-In-First-Out (FIFO) buffer 212, a sorting module 216, and a percentile estimator 220. The FIFO buffer 212 has a fixed size, which is selected to include, in addition to the current frame of the bin, a portion of the bin's data. A historical frame. In operation, the FIFO buffer 212 receives a frequency domain signal (on bus 204) and continuously refreshes the frequency bin data stored therein by replacing the oldest frame of the bin with the latest frame bin provided via the frequency domain signal (on bus 204). The FIFO buffer 212 is connected to the sorting module 216 via a third path 214. The sorting module 216 is further connected to the percentile estimator 220 via a fourth path 218. The percentile estimator 220 additionally receives the aforementioned content-aware segmentation information from the audio event classification and segmentation unit 130.
[0047] In one example, sorting module 216 applies a sorting function to the current contents of FIFO buffer 212 obtained via third path 214. The mathematical expression for this sorting operation is as follows: (3) The sorting operation is performed separately for each bin frequency. The sorted frequency bin data generated by the sorting module 216 is represented by the left side of equation (3) and is provided to the percentile estimator 220 via the fourth path 218.
[0048] In a noisy segment, the energy in each frequency bin represents the noise level, but the minimum and maximum levels can differ significantly. The minimum level typically underestimates the noise floor and is quite sensitive to outliers. The maximum level can represent transient impulses and therefore may lead to an overestimation of the noise floor. Additionally, the maximum level can also be susceptible to outliers. Therefore, a certain percentile of the sorted FIFO buffer often better represents the noise floor in a noisy segment. For example, in a speech segment, the FIFO buffer is not permanently occupied by speech, for example, due to the inter-sentence silences inherent in any speech. Therefore, for a portion of the time (e.g., quantized as a percentage), the energy in each bin frequency is at the noise floor level in the speech segment. In a musical segment, no single energy in each frequency bin may be at the noise floor level because harmonics from certain instruments (such as the didgeridoo) may persist for a relatively long time. In this case, even the minimum level may overestimate the noise floor. Therefore, in some cases, noise estimation can be stopped in a musical segment to avoid subsequent oversuppression of useful sounds.
[0049] Different content types typically have different respective proportions of silence. Accordingly, in some examples, the percentile estimator 220 is configured to adaptively estimate the floor noise percentile of the ordered FIFO buffer based on segments classified as speech, music, or noise . A corresponding mathematical representation of the percentile estimation operation is as follows: (4) wherein, is a classification-sensitive percentile estimation function applied by the percentile estimator 220 to the ordered FIFO buffer based on the segment classification provided by the percentile estimator 220 from the content-aware segmentation information received by the percentile estimator 220 from the audio event classification and segmentation component 130.
[0050] The percentile estimator 220 is further configured to truncate the ordered FIFO buffer to a size of less than a time buffer length of a duration . This truncation operation can be mathematically represented as follows: (5) wherein the duration is expressed as: (6) The left-hand side of equation (5) provides the floor noise estimate value computed by the adaptive percentile estimator 220.
[0051] The content-aware noise estimation component 110 further comprises a floor noise update module 224 and a noise update controller 228. The floor noise update module 224 is connected to the percentile estimator 220 via a fifth path 222, and is further connected to the noise update controller 228 via a control path 226. The noise update controller 228 also receives the content-aware segmentation information. In operation, the floor noise update module 224 performs a recursive floor noise update based on the floor noise estimate value computed by the percentile estimator 220 and received therefrom via the fifth path 222. In one example, the recursive floor noise update is performed in accordance with equation (7): (7) wherein, is a dynamic smoothing factor provided by the noise update controller 228 to the floor noise update module 224 via the control path 226. For example, as described above, the updated floor noise estimate value obtained by the floor noise update module 224 is emitted to the content-aware noise suppression component 120 via the intermediate output path 230. The frequency domain signal and the updated floor noise estimate value are denoted as Figure 1a first signal component and a second signal component of the emitted intermediate output signal.
[0052] In one example, the noise update controller 228 computes the dynamic smoothing factor as follows: (8) where, is a constant; and is a time-dependent update coefficient obtained as follows: (9a) where, is a threshold value selected to constrain the capture of transient sounds and to reduce or substantially avoid errors in the background noise estimate. The quantity represents the dynamic range between the maximum and minimum horizontal bins. In one example, is computed as follows: (9b) Equations (3) through (9) apply to inputs having magnitudes obtained via a time- frequency transform, such as a Fourier transform. Those of ordinary skill in the relevant art will readily understand how to adapt equations (3) through (9) to alternative input formats, such as magnitude power or magnitudes expressed in dB. In the preferred implementation, the time-frequency transform produces a frequency representation of the input audio signal 102, such as the frequency-domain signal on bus 204, having sufficient frequency resolution to support processing of audio signals having rich harmonic content. For example, some musical instruments have many harmonics with fundamental frequencies below about 100 Hz.
[0053] The content-aware noise suppression component 120 includes an SNR computation module 260, a first suppression gain computation module 264, a second suppression gain computation module 270, an aggressiveness controller 274, a multiplication module 278, a gain application module 282, and an inverse FFT (IFFT) module 290. The input signals received by the content-aware noise suppression component 120 include the above-described intermediate output signal and the content-aware segmentation information. The first signal component of the intermediate output signal is applied to the SNR computation module 260 and the gain application module 282. The second signal component of the intermediate output signal is applied to the SNR computation module 260 and the second suppression gain computation module 270. The output signal generated by the IFFT module 290 is the output audio signal 122.
[0054] The SNR computation module 260 is connected to the first suppression gain computation module 264 via a sixth path 262. The first suppression gain computation module 264 is further connected to the second suppression gain computation module 270 and the multiplication module 278 via a seventh path 266. The second suppression gain computation module 270 is further connected to the multiplication module 278 via an eighth path 272. The aggressive controller 274 is connected to the multiplication module 278 via a ninth path 276. The multiplication module 278 is further connected to the gain application module 282 via a tenth path 280. The gain application module 282 is further connected to the IFFT module 290 via an eleventh path 284.
[0055] In one example, the SNR computation module 260 is configured to compute the SNR value based on the following mathematical expression: (10) wherein the values of the first and second signal components of the intermediate output signal provide the values of and respectively. The computed SNR value is then sent to the first suppression gain computation module 264 via the sixth path 262.
[0056] The first suppression gain computation module 264 is configured to compute the first suppression gain based on the following mathematical expression: (11) wherein the SNR value is received from the first suppression gain computation module 264 via the sixth path 262. In different examples, the mapping function may be designed based on different suitable non-linear curves. The computed first suppression gain value is used to compute the residual signal which is sent to the second suppression gain computation module 270 and the multiplication module 278 via the seventh path 266.
[0057] The second suppression gain computation module 270 is configured to remove the residual noise from the residual signal based on the following mathematical expression: (12) wherein the second suppression gain value is denoted by ; and is a second mapping function. The values of the residual signal are received from the first suppression gain computation module 264 via the seventh path 266. These values are received together with the second signal component of the intermediate output signal (on the path 230). The second mapping function is configured to compare and the value of the difference, and based on the value of the difference In different examples, the second mapping function may be designed based on a suitable non-linear curve. The value of the difference is sent to the multiplication module 278 via an eighth path 272.
[0058] The aggressiveness controller 274 is configured to apply different aggressiveness settings to different types of content to reduce the detrimental effects of possible over-suppression of the relevant content. In one example, the aggressiveness controller 274 obtains a content-aware gain As follows: (13) wherein, denotes a content-aware gain value; and is a content-aware gain function. The value of the difference is sent to the multiplication module 278 via a ninth path 276.
[0059] The multiplication module 278 is configured to compute a final noise-suppression gain As follows: (14) The value of the difference is sent to the gain application module 282 via a tenth path 280. The gain application module 282 is configured to apply the received gain value to the bin to compute a corresponding output bin as follows: (15) The output bin with the original phase is then directed to the IFFT module 290 via an eleventh path 284. The IFFT module 290 generates the output audio signal 122 via an IFFT operation applied to the frame of output bins .
[0060] Figure 3 is a block diagram illustrating a workflow 300 implemented in the audio event classification and segmentation component 130 of the content-aware noise management system 100 according to some aspects of the present disclosure. The input 302 of the workflow 300 includes the above-mentioned input audio signal 102 and / or frequency-domain signal (on the bus 204). The output 332 of the workflow 300 includes the above-mentioned content-aware segmentation information.
[0061] Workflow 300 includes an audio analysis and classification block 310, a moving average convergence-divergence (MACD) block 320, and an audio event segmentation block 330. In audio analysis and classification block 310, the audio signal of each frame of input 302 is analyzed for energy and / or classified to generate raw confidence(s). In MACD block 320, the energy or confidence is post-processed using a suitable MACD algorithm. In audio event segmentation block 330, the audio events of input 302 are segmented based on the real-time MACD processing results to generate output 332.
[0062] The operation of MACD block 320 includes computing the difference between a short-term smoothing and a long-term smoothing of a particular value to predict a trend momentum. In one example, the MACD processing implemented in MACD block 320 is as follows based on a value Four outputs are recursively generated: (16) (17) (18) (19) wherein, denotes a short-term moving average; denotes a long-term moving average; denotes the difference between the short-term moving average and the long-term moving average; and denotes a moving average of the difference. In some examples, denotes a raw confidence generated by multi-classification module 240, such as a speech confidence , a music confidence , and a noise confidence , or a raw frame energy . Coefficients denote a short-term smoothing factor, a long-term smoothing factor, and a smoothing factor for convergence / divergence, respectively. These coefficients are hyperparameters of the MACD algorithm.
[0063] In some examples, the audio signal waveform is first divided into windowed semi-overlapping frames, and then the frame data is converted to the frequency domain using a filter bank or a time-frequency transformer, such as FFT module 202. As explained previously, the magnitudes of the contents of the spectral blocks are collectively referred to as wherein, denotes a frequency index, and denotes a time index.
[0064] In some examples of workflow 300, energy module 244 is used to compute the total energy (denoted in dB) of each frame according to the following mathematical expression: (20) wherein is the total number of bins. For different recording devices, the absolute values can be different, e.g., due to differences in preset system gain, recording distance, and / or loudness of the loudspeaker. Therefore, it is not recommended to directly base the threshold for determining an audio event on However, based on the MACD processing described above, denotes the difference between the short-term and long-term moving average of which can be used to determine general audio events in a more robust way, e.g., because does not directly depend on the absolute value of
[0065] Based on the above considerations, the MACD processing of implemented in the MACD block 320 can be represented as follows: (21) wherein and are pair factors defining a trade-off between sensitivity to onset / offset and segmentation latency. The coefficients and can be selected based on the specific properties of the MACD application.
[0066] In some examples of the workflow 300, the operation of the audio event segmentation block 330 comprises segmenting general audio events based on the following condition: (22) wherein is a threshold for determining an event onset, and is a threshold for determining an event end.
[0067] Figure 4 Fig. 4 is a schematic diagram illustrating audio event clusters used in the workflow 300 according to some aspects of the present disclosure, with pictorial illustrations. In the illustrated example, various audio events are assigned to a plurality of clusters distributed over several audio classes including a vocal class 410, a music class 420, and a noise class 430. The vocal class 410 includes a non-speech cluster 412 and a speech cluster 414. The music class 420 includes a vocal cluster 422, an instrument cluster 424, and a mixed cluster 426. The noise class 430 includes a selective noise cluster 432, a transient noise cluster 434, and an “other noise” cluster 436. Figure 4Examples of certain clusters are listed in Table 1 as illustrations. Note that certain clusters can span more than one class. For example, the human voice cluster 422 spans both the human voice class 410 and the music class 420. In other examples, other suitable clustering schemes can also be used.
[0068] In practice, due to the above-mentioned overlap between clusters and classes, it is challenging to identify the cluster to which an audio signal belongs with 100% accuracy. For real-time communication systems (such as systems that carry audio or video calls, live broadcasts, etc.), steady-state room background noise, device fan noise, electrical device noise, etc. are types of noise that can be identified and suppressed with relatively high confidence. However, for transient noise, it is often more difficult to definitively determine user preference - to suppress or to share. Thus, in designing a noise classifier, a subset of the entire noise class 430 can be selected as training data for noise classification. In this way, any overlap between noise and speech or music can be mitigated to obtain a more robust selective noise classifier.
[0069] In some examples of the framework 300, the current frame and the sequence of adjacent history frames are fed into the multi-classification module 240 to generate raw confidences related to specific classes, such as speech confidence , music confidence , and noise confidence , as described above with reference to Figure 2 and Equation (1). In such examples, the MACD processing of the raw speech confidence implemented in the MACD block 320 can be represented as follows: (23) Then, the speech events can be segmented in the audio event segmentation block 330 using the following condition: (24) where , , and are optional thresholds depending on the specific application.
[0070] In some examples, the audio event segmentation block 330 is configured to perform segmentation of music events and noise events using similar methods (e.g., as exemplified by Equations (23) to (24)). In such examples, the smoothing factors , , and the thresholds , , and The selection of values for the classes 410, 420, 430 will typically differ for different classes in order to appropriately balance the sensitivity and latency aspects of the framework 300.
[0071] Figure 5 is a flowchart illustrating an example noise management method 500 in accordance with some aspects of the present disclosure. The method 500 can be implemented using the content-aware noise management system 100. The method 500 can be performed by a processor, which can be configured to perform the method 500 via machine-executable instructions. The method 500 can be divided into individual blocks or partitions, such as blocks 502, 504, 506, 508, and 510. Figure 5 The individual process blocks illustrated in the method 500 provide examples of various methods disclosed herein, and it should be understood that some blocks can be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, the processing of the individual blocks can begin at block 502, which can be described as a process, method, step, block, operation, or function.
[0072] At block 502, “convert the input audio signal(s) to frequency bin data,” the method 500 includes converting the input audio signal 102 to frequency bin data. In some examples, the operations of block 502 include performing a Fourier transform, such as the Fourier transform implemented with the FFT module 202. Processing can proceed from block 502 to block 504.
[0073] At block 504, “perform audio event segmentation on the converted input audio signal,” the method 500 further includes performing audio event segmentation on the converted input audio signal. The operations of block 504 include generating content-aware segmentation information (on path 128) corresponding to the input audio signal 102. In some examples, the audio event segmentation is performed in real-time. In some examples, the content-aware segmentation information (on path 128) includes frame-by-frame speech, music, and noise confidence. The operations of block 504 further include processing the frame-by-frame speech, music, and noise confidence to identify a contiguous sequence of audio events corresponding to the input audio signal 102. The audio events are selected from a plurality of event clusters, each event cluster of the plurality of event clusters being associated with one or more classes selected from the group consisting of the human voice class 410, the music class 420, and the noise class 430. In some examples, the processing operations are configured to mitigate false selections from occurring in the identified contiguous sequence. In some examples, the operations of block 504 further include moving average convergence-divergence processing 320 of the frame-by-frame speech, music, and noise confidence.
[0074] At block 506, "estimate noise floor level based on segmentation information," the method 500 further includes estimating a noise floor level. The noise floor level (on path 230) is associated with the input audio signal 102 and is computed based on the content-aware segmentation information (on path 128) and further based on ranking frequency bin data (on bus 204) corresponding to a fixed length portion of the input audio signal 102. The operations of block 506 include updating the fixed length portion using the FIFO buffer 212 configured to receive the frequency bin data (on bus 204). The operations of block 506 further include truncating the ranked frequency bin data 218 to a size determined by the adaptive percentile estimator 220 based on the content-aware segmentation information (on path 128). In some examples, the adaptive percentile estimator 220 is configured to select different respective percentiles for speech events, music events, and noise events, and the determined size is based on a selected percentile of the different respective percentiles. The operations of block 506 further include updating the noise floor level (on path 230) using the update controller 224. In some examples, the update controller 224 is configured to start and stop updating based on audio event data included in the content-aware segmentation information (on path 128).
[0075] At block 508, "apply noise suppression to frequency bin data," the method 500 further includes applying noise suppression to the input signal 102 represented by the frequency bin data. In some examples, the noise suppression is performed in frequency bins having a selected frequency resolution and is based at least in part on the content-aware segmentation information (on path 128) and the estimated noise floor level (on path 230). The operations of block 508 include computing bin-wise SNR values based on amplitudes of the frequency bins and further based on the estimated noise floor level (on path 230). The operations of block 508 further include computing suppression gain values based on the computed bin-wise SNR values and further based on the estimated noise floor level (on path 230). In some examples, such computation includes: (i) computing a respective first suppression gain value based on a respective bin-wise SNR value of the bin-wise SNR values; (ii) computing a respective second suppression gain value based on the respective first suppression gain value and further based on a respective noise floor level of the estimated noise floor level; and (iii) computing a product of the respective first suppression gain value and the respective second suppression gain value. The operations of block 508 further include applying different respective aggressiveness of the noise suppression to different types of audio content based on the content-aware segmentation information (on path 128).
[0076] In some examples, at least some of the operations of block 506 are performed using frequency bins having a first selected frequency resolution. In contrast, at least some of the operations of block 508 are performed using frequency bins having a second selected frequency resolution that is different from the first selected frequency resolution. The second selected frequency resolution can be higher than the first selected frequency resolution. In such examples, the operations of block 508 can also include converting the estimated noise floor level (on path 230) from the first selected frequency resolution to the second selected frequency resolution.
[0077] At block 510, "applying an inverse Fourier transform to the noise-suppressed frequency bins to generate an output audio signal(s)", the method 500 further includes applying an inverse Fourier transform to the output frequency bins. The operations of block 510 also include concatenating time-domain signal portions generated via the inverse Fourier transform to generate the output audio signal 122.
[0078] Figure 6A FIG. illustrates a schematic block diagram of an example device architecture 600 (e.g., apparatus 600) that can be used to implement various aspects of the present disclosure. The architecture 600 includes, but is not limited to, server and client devices, systems, and methods as described with reference to Figures 1 to 5 As shown, the architecture 600 includes a central processing unit (CPU) 601 capable of performing various processes according to programs stored in, for example, a read only memory (ROM) 602 or loaded from, for example, a storage unit 608 to a random access memory (RAM) 603. The CPU 601 can be, for example, an electronic processor 601. In the RAM 603, data required when the CPU 601 performs various processes is also stored as necessary. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output interface 605 is also connected to the bus 604.
[0079] The following components are connected to the I / O interface 605: an input unit 606 that can include a keyboard, a mouse, and the like; an output unit 607 that can include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 608 that includes a hard disk or another suitable storage device; and a communication unit 609 that includes a network interface card such as a wired or wireless network card.
[0080] In some implementations, the input unit 606 includes one or more microphones located at different positions (depending on the host device) that enable capturing of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0081] In some implementations, output unit 607 includes a system with a variety of numbers of speakers. Output unit 607 (depending on the capabilities of the host device) can render audio signals in various formats, such as mono, stereo, immersive, binaural, and other suitable formats.
[0082] In some embodiments, communication unit 609 is configured to communicate with other devices (e.g., via a network). Drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted on drive 610 so that computer programs read from it are installed into storage unit 608. Those skilled in the art will understand that although apparatus 600 is described as including the above-described components, some of these components may be added, removed, and / or replaced in various applications, and all such modifications or alterations fall within the scope of this disclosure.
[0083] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 609, and / or installed from removable medium 611, such as... Figure 6A As shown.
[0084] Figure 6B The illustration shows an example in Figure 6A A schematic block diagram of a CPU 601 implemented in architecture 600. In the example shown, CPU 601 includes an electronic processor 620 and a memory 621. The electronic processor 620 is electrically and / or communicatively connected to the memory 621 for bidirectional communication. The memory 621 stores software 622. In some examples, the memory 621 may be located internally to the electronic processor 620, such as as internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, the memory 621 may be located externally to the electronic processor 620, such as in ROM 602, RAM 603, flash memory, or removable medium 611, or another non-transitory computer-readable medium contemplated for architecture 600. In some instances, the electronic processor 620 runs software 622 stored in the memory 621 to perform the functions described above. Figures 1 to 5 The description includes various methods and operations related to content-aware noise management.
[0085] Those of ordinary skill in the related arts will readily recognize that the present invention is in no way limited to the above-described example embodiments. Rather, numerous modifications and variations are possible and are considered to be within the scope of the appended claims. Various aspects and embodiments of the present disclosure can also be understood from the following enumerated example embodiments (EEEs), which are not claims and can represent systems, methods, and apparatuses all arranged in accordance with various aspects of the present disclosure.
[0086] EEE1. A noise management method comprising: performing audio event segmentation on an input audio signal to generate content-aware segmentation information; estimating a noise floor level associated with the input audio signal based on the content-aware segmentation information and further based on ordering frequency bin data corresponding to fixed length portions of the input audio signal; and applying noise suppression to the input audio signal to generate an output audio signal, the noise suppression being performed in frequency bins having a selected frequency resolution and based at least in part on the content-aware segmentation information and the estimated noise floor level.
[0087] EEE2. The method of EEE1, wherein the audio event segmentation is performed in real time.
[0088] EEE3. The method of claim EEE1 or EEE2, wherein the content-aware segmentation information comprises frame-wise speech, music, and noise confidences.
[0089] EEE4. The method of EEE3, wherein the performing comprises processing the frame-wise speech, music, and noise confidences to identify a contiguous sequence of audio events corresponding to the input audio signal, the audio events being selected from a plurality of event clusters, each event cluster of the plurality of event clusters being associated with one or more classes selected from the group consisting of a vocal class, a music class, and a noise class.
[0090] EEE5. The method of EEE4, wherein the processing is configured to mitigate an occurrence of a false selection in the identified contiguous sequence.
[0091] EEE6. The method of any of EEE3-EEE5, wherein the performing comprises a moving average convergence-divergence processing of the frame-wise speech, music, and noise confidences.
[0092] EEE7. The method of any of EEE1-EEE6, wherein the estimating comprises updating the fixed length portions using a first-in-first-out buffer configured to receive the frequency bin data.
[0093] EEE8. The method of any of EEE1-EEE6, wherein the estimating comprises truncating the sorted frequency bin data to a size determined with an adaptive percentile estimator based on the content-aware partitioning information.
[0094] EEE9. The method of EEE8, wherein the adaptive percentile estimator is configured to select different respective percentiles for speech events, music events, and noise events, the determined size being based on a selected one of the different respective percentiles.
[0095] EEE10. The method of any of EEE1-EEE6, wherein the estimating comprises updating the noise floor level using an update controller configured to start and stop the updating based on audio event data included in the content-aware partitioning information.
[0096] EEE11. The method of any of EEE1-EEE10, wherein the applying comprises computing a per-bin signal-to-noise ratio value based on the frequency bin’s magnitude and the estimated noise floor level.
[0097] EEE12. The method of EEE11, wherein the applying further comprises computing an attenuation gain value based on the computed per-bin signal-to-noise ratio value and further based on the estimated noise floor level.
[0098] EEE13. The method of EEE12, wherein, for each of the frequency bins, the computing comprises computing a respective first attenuation gain value based on a respective one of the per-bin signal-to-noise ratio values, computing a respective second attenuation gain value based on the respective first attenuation gain value and further based on a respective one of the estimated noise floor levels, and computing a product of the respective first attenuation gain value and the respective second attenuation gain value.
[0099] EEE14. The method of any of EEE1-EEE13, wherein the applying comprises applying different respective aggressiveness settings for the noise attenuation to different types of audio content based on the content-aware partitioning information.
[0100] EEE15. The method of any of EEE1-EEE14, wherein the estimating is performed using the frequency bins having a first selected frequency resolution; and wherein the applying is performed using the frequency bins having a second selected frequency resolution different from the first selected frequency resolution.
[0101] EEE16. The method according to EEE15, wherein the second selected frequency resolution is higher than the first selected frequency resolution.
[0102] EEE17. The method according to EEE16, wherein the estimation further includes converting the estimated background noise level from the first selected frequency resolution to the second selected frequency resolution.
[0103] EEE18. The method according to any one of EEE1 to EEE17 further includes: converting the input audio signal into the frequency bin data using a Fourier transform; and applying the noise suppression to the frequency bin to the frequency bin and then applying an inverse Fourier transform to the frequency bin to generate the output audio signal.
[0104] EEE19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including any one of EEE1 to EEE18.
[0105] EEE20. A noise management apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured, together with the at least one processor, to cause the apparatus to at least: perform audio event segmentation on an input audio signal to generate content-aware segmentation information; estimate a noise floor level associated with the input audio signal based on the content-aware segmentation information and further based on sorting frequency bin data corresponding to a fixed-length portion of the input audio signal; and apply noise suppression to the input audio signal to generate an output audio signal, the noise suppression being performed in a frequency bin having a selected frequency resolution and at least in part based on the content-aware segmentation information and the estimated noise floor level.
[0106] Regarding the processes, systems, methods, heuristics, etc., described herein, it should be understood that although the steps of these processes, etc., have been described as being performed in a specific ordered sequence, these processes can be practiced using the described steps performed in a different order than that described herein. Furthermore, it should be understood that some steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should in no way be construed as limiting the claims.
[0107] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0108] All terms used in the claims are to be given their broadest interpretation consistent with their ordinary meaning and the context in which they are used, unless an otherwise expressly limited interpretation is clearly indicated in the specification. In particular, the use of a singular term, such as "a", "the", "said", etc., shall be read to encompass the plural unless the context clearly indicates otherwise.
[0109] The abstract of the disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or the meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments have more features than are explicitly recited in each claim. Rather, as the appended claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. The claims are hereby hereby incorporated into the Detailed Description, causing that section to operate also as a transclusion of the claims. Thus, the claims are part of the disclosure.
[0110] While the present disclosure includes reference to illustrative embodiments, the specification herein is not intended to be interpreted as limiting. Various modifications to the described embodiments, as well as other embodiments within the scope of the present disclosure, are contemplated and are considered to be within the scope of the present disclosure, for example, as expressed in the claims.
[0111] Some embodiments can be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0112] Some embodiments can be realized in a method and an apparatus for practicing the method. Some embodiments can also be realized in the form of program code recorded in a tangible medium, such as a magnetic record medium, an optical record medium, a solid state memory, a floppy diskette, a CD-ROM, a hard disk drive, or any other non-transitory machine readable storage medium, wherein, when the program code is loaded into a machine, such as a computer, and executed, it becomes an apparatus for practicing the patent invention(s). Some embodiments can also be realized in the form of program code, such as being stored in a non-transitory machine readable storage medium, including being loaded into and / or executed by a machine, such as a computer or a processor, wherein, when the program code is loaded into the machine and executed, it becomes an apparatus for practicing the patent invention(s). Program code segments, when implemented on general purpose processors, combine with the processors to provide unique devices that operate akin to specific logic circuits.
[0113] Unless specifically stated otherwise, each numerical value and range should be interpreted as approximately as if the word "about" preceded the respective value or range.
[0114] The use of the diagram number and / or the reference characters in the claims is intended to identify one or more possible embodiments of the claimed subject matter, in order to facilitate interpreting the claims. Such use does not limit the scope of the claims to the corresponding embodiments in which the diagram number and / or reference characters are used.
[0115] Although the elements in the method claims, if any, are recited in a particular order, unless otherwise specified, the order of elements in the claims does not inherently imply a specific order of performing the elements, and the method can be performed in any order that is practicable.
[0116] Reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. The appearance of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all mutually exclusive or alternative embodiments to the other. This is also true for the term "implementation."
[0117] Unless otherwise provided herein, the use of ordinal adjectives, such as "first," "second," "third," etc., to refer to a number of similar objects, indicates the different instance of the similar objects being referred to, and is not meant to imply that the similar objects must be in a certain order, in a certain ranking, or in any other manner in time, in space, in ranking, or in any other manner.
[0118] Unless otherwise indicated herein, the conjunctive term "if' can be further or alternatively interpreted as meaning "when," or "upon," or "in response to a determination," or "in response to a detection," depending on the corresponding context. For example, the phrase "if it is determined that" or "if [stated condition] is detected" can be interpreted to mean "upon determining that" or "in response to determining that" or "upon detecting [stated condition]" or "in response to detecting [stated condition]."
[0119] Also for the purposes of this description, the terms "couple," "coupling," "coupled," "connect," "connecting," or "connected" mean any of the ways in which the energy is transferred between two or more elements as known in the art or later developed, and the insertion of one or more additional elements is contemplated, although this is not necessary. In contrast, the terms "directly coupled," "directly connected," and the like imply the absence of such additional elements.
[0120] As used herein with respect to elements and standards, the term "compliant" means that the element communicates with other elements in a manner that is all or part of the manner specified by the standard, and is recognized by the other elements as being sufficient to enable communication with the other elements in the manner specified by the standard. A compliant element need not operate internally in the manner specified by the standard.
[0121] The functions of the various elements shown in the figures, including any functional blocks labeled as "processors" and / or "controllers," can be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions can be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which can be shared. Moreover, explicit use of the term "processor" or "controller" should not be construed to refer exclusively to hardware capable of executing software, and can implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and non volatile storage. Other hardware, conventional and / or custom, can also be included. Similarly, any switches shown in the figures are conceptual only. Their function can be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0122] As used in this application, the terms "circuit" and "circuitry" can refer to one or more or all of the following: (a) a purely hardware electromechanical implementation (as in a purely analog circuit and / or digital circuit); (b) combinations of hardware circuits and software, such as (if applicable): (i) combinations of analog and / or digital hardware circuits with software / firmware and (ii) portions of hardware processor(s) with software (including digital signal processors), software, and memory that work together to cause an apparatus, such as a mobile phone or server, to perform various functions; and (c) hardware circuitry alone (in conjunction with human software that needs software to operate but can not need software when operating, e.g., a microprocessor or a portion of a microprocessor). This definition of circuit applies to all uses of this term in this application (including all claims). As a further example, as used in this application, the term "circuit" also covers an implementation that is an only hardware circuit or an only processor (or multiple processors) or a portion of a hardware circuit or processor with its accompanying software and / or firmware. The term circuit, for example, covers a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device, as applicable to a particular claim element.
[0123] Those of ordinary skill in the art will realize that any of the block diagrams herein represent a conceptual view of illustrative circuitry embodying the principles of the present disclosure. Similarly, it will be appreciated that any flow chart, flow diagram, state transition diagram, pseudocode, and / or the like represent various processes which can be substantially represented in computer readable medium and so executed by a computer or processor, whether or not the computer or processor is explicitly shown.
[0124] The "SUMMARY" in the detailed description is intended to introduce some example embodiments and additional embodiments are described in the "DETAILED DESCRIPTION" and / or with reference to one or more of the accompanying drawings. The "SUMMARY" is not intended to identify key or essential elements of the claimed subject matter, nor does it limit the scope of the claimed subject matter.
Claims
1. A noise management method (500), comprising: performing (504) audio event segmentation on an input audio signal to generate content-aware segmentation information; estimating (506) a floor noise level associated with the input audio signal based on the content-aware segmentation information and further based on ordering frequency bin data corresponding to fixed length portions of the input audio signal; and applying (508) noise suppression to the input audio signal to generate an output audio signal, the noise suppression being performed in frequency bins having a selected frequency resolution and being based at least in part on the content-aware segmentation information and the estimated floor noise level. The audio event segmentation is performed in real time.
2. The method of claim 1, wherein, The content-aware segmentation information includes frame-wise speech, music, and noise confidences.
3. The method of claim 1 or 2, wherein, The performing (504) includes processing the frame-wise speech, music, and noise confidences to identify a contiguous sequence of audio events corresponding to the input audio signal, the audio events being selected from a plurality of event clusters, each event cluster of the plurality of event clusters being associated with one or more classes selected from the group consisting of a vocal class, a music class, and a noise class.
4. The method of claim 3, wherein, The processing is configured to mitigate an occurrence of a false selection in the identified contiguous sequence.
5. The method of claim 4, wherein, The performing (504) includes a moving average convergence-divergence processing of the frame-wise speech, music, and noise confidences.
6. The method of any one of claims 3 to 5, wherein, The estimating (506) includes updating the fixed length portions using a first-in-first-out buffer configured to receive the frequency bin data.
7. The method of any one of claims 1 to 6, wherein, The estimating (506) includes truncating the ordered frequency bin data to a size determined by an adaptive percentile estimator (220) based on the content-aware segmentation information.
8. The method of any one of claims 1 to 6, wherein, The adaptive percentile estimator is configured to select different respective percentiles for speech events, music events, and noise events, the determined size being based on a selected one of the different respective percentiles.
9. The method of claim 8, wherein, The estimating (506) includes updating the floor noise level using an update controller configured to start and stop the updating based on audio event data included in the content-aware segmentation information.
10. The method of any one of claims 1 to 6, wherein, The applying (508) includes computing bin-wise signal-to-noise ratio values based on amplitudes of the frequency bins and the estimated floor noise level.
11. The method of any one of claims 1 to 10, wherein, The applying (508) further includes computing suppression gain values based on the computed bin-wise signal-to-noise ratio values and further based on the estimated floor noise level.
12. The method of claim 11, wherein, For each of the frequency bins, the computing includes:
13. The method of claim 12, wherein, computing a respective first suppression gain value based on a respective one of the bin-wise signal-to-noise ratio values; computing a respective second suppression gain value based on the respective first suppression gain value and further based on a respective one of the estimated floor noise levels; and computing a product of the respective first suppression gain value and the respective second suppression gain value. The applying (508) includes applying different respective aggressiveness settings of the noise suppression to different types of audio content based on the content-aware segmentation information.
14. The method of any one of claims 1 to 13, wherein, 15. The method of any one of claims 1 to 14, wherein, said estimating (506) is performed using the frequency bins having a first selected frequency resolution; and wherein said applying (508) is performed using the frequency bins having a second selected frequency resolution different from the first selected frequency resolution.
16. The method of claim 15, wherein, said second selected frequency resolution is higher than the first selected frequency resolution.
17. The method of claim 16, wherein, said estimating (506) further comprises converting the estimated noise floor level from the first selected frequency resolution to the second selected frequency resolution.
18. The method of any one of claims 1 to 17, further comprising: converting (502) the input audio signal into the frequency bin data using a Fourier transform; and applying (510) an inverse Fourier transform to the frequency bins after applying the noise suppression to the frequency bins to generate the output audio signal.
19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1 to 18.
20. A noise management apparatus (100), comprising: at least one processor (601); and at least one memory (621) including program code (622), wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to perform the following operations: perform (504) an audio event segmentation on an input audio signal (102) to generate content-aware segmentation information; estimate (506) a noise floor level associated with the input audio signal based on the content-aware segmentation information and further based on ranking frequency bin data corresponding to fixed length portions of the input audio signal; and apply (508) noise suppression to the input audio signal to generate an output audio signal (122), the noise suppression being performed in frequency bins having a selected frequency resolution and based at least in part on the content-aware segmentation information and the estimated noise floor level.