Automatic detection and attenuation of speech articulatory noise events

By performing frame segmentation and feature parameter detection on the speech recording signal, the system automatically identifies and attenuates mouth clicks and plosives, solving the problems of low noise removal efficiency and cumbersome manual editing in existing technologies, and improving recording quality.

CN116670755BActive Publication Date: 2026-02-06DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180062729.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-29
Filing Date
2021-08-11
Publication Date
2026-02-06
Estimated Expiration
2041-08-11

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and automatically remove noise caused by pronunciation in speech recordings, such as mouth clicking and plosive sounds, leading to an unpleasant listening experience, and manual editing is cumbersome.

Method used

By segmenting the input audio signal into multiple frames, the system detects and classifies speech pronunciation noise events, including mouth clicks and plosives, using feature parameters, and performs automatic attenuation processing based on these feature parameters.

Benefits of technology

It achieves efficient and automatic noise removal from speech recordings, improving the listening experience and avoiding tedious manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116670755B_ABST
    Figure CN116670755B_ABST
Patent Text Reader

Abstract

A method of performing automatic audio enhancement on an input audio signal including at least one speech sound noise event is described. The method includes segmenting the input audio signal into a plurality of audio frames, obtaining at least one feature parameter from the audio frames, and determining a respective type of the speech sound noise event and a respective time-frequency range associated with the speech sound noise event within the input audio signal based at least in part on the obtained feature parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to the following prior applications: Spanish application P202030864 (reference number: D20066ES), filed August 12, 2020, and U.S. Provisional Application 63 / 107,012 (reference number: D20066USP1), filed October 29, 2020, which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the general field of performing automatic audio enhancement, such as automatic detection and attenuation, on speech articulation noise events (e.g., mouth clicks, speech plosives, etc.). Background Technology

[0004] On various media platforms, the increasing volume of speech content, often of varying quality, has reached a point where manual editing alone seems insufficient. Effective automatic speech enhancement can maintain natural speech and reduce editing workload.

[0005] Generally speaking, speech enhancement algorithms can handle two types of unwanted "noise": noise generated by background sources and noise generated by pronunciation.

[0006] Plosives belong to the second type. Plosives typically occur when an air blast is generated from the mouth (e.g., during the pronunciation of syllables containing "p" or "t") and the impact of this air blast causes a large vibration of the microphone diaphragm. In the context of this disclosure, the term "plosive" is used broadly to include any air blast emitted from the mouth that causes a large vibration of the microphone diaphragm (e.g., including short fricatives like "f" and "z").

[0007] Even when speech is recorded in a well-controlled acoustic environment, plosives can often produce a sudden low-frequency boost, known as a "poof," which can lead to an unpleasant listening experience.

[0008] Several recording techniques have been proposed to reduce the intensity of plosive sounds, such as using pop filters or deflectors, and off-axis speaking. However, due to practical reasons, pop reduction is not as effective as expected: for example, it may fail to stabilize the speaker's or actor's posture, or physical filters may reduce the emotional connection with the audience. Therefore, signal processing tools are needed to improve the quality of such recordings. The process of detecting and attenuating plosive sounds is often referred to as "plosive removal" (or sometimes as "deplosive" or "deplosive processing").

[0009] Mouth clicks are another type of transient sound caused by speech sounds produced using the tongue / teeth / lips mixed with saliva. These mouth clicks can occur in both speech and non-speech parts and are usually audible through headphones / earphones for high SNR recordings. Mouth clicks are typically short, usually lasting between 10ms and 100ms, and can also appear as several consecutive transients.

[0010] In the context of professional recordings such as TV / movie / game dialogues, the quality of speech without clicking sounds can be extremely demanding. Now, even for user-generated content, the widespread use of headsets / headphones has made mouth clicking sounds quite noticeable.

[0011] Several recording techniques have been proposed to reduce mouth clicks by professional male / female voice actors. However, in most cases, there is no way to control the speaker's mouth / lip movements. Manual editing in post-production can be tedious, making it impractical to process hundreds or thousands of dialogues. Therefore, signal processing tools are needed to more efficiently correct mouth clicks. The process of detecting and attenuating mouth clicks is often referred to as "mouth click removal" or simply "click removal" (or sometimes as "de-clicking" or "de-clicking processing").

[0012] Therefore, broadly speaking, the focus of this disclosure is to propose techniques for performing automatic audio enhancement (including, but not limited to, detection and attenuation) on audio signals that include one or more speech articulation noise events (e.g., mouth clicks, speech plosives, etc.). Summary of the Invention

[0013] In view of the foregoing, this disclosure generally provides a method for performing automatic audio enhancement on an input audio signal including at least one speech pronunciation noise event, as well as corresponding apparatus, programs, and computer-readable storage media having the features of the respective independent claims.

[0014] According to one aspect of this disclosure, a method is provided for performing automatic audio enhancement on an input audio signal that includes at least one speech articulation noise event. As those skilled in the art will understand and recognize, automatic audio enhancement can involve any suitable audio enhancement means, including (but not limited to) the automatic detection and attenuation of (multiple) speech articulation noise events(s) within the input audio signal. Here, the term speech articulation noise event can be understood broadly, for example, as referring to a noise event that is related to or caused (i.e., generated) in some way by speech articulation.

[0015] Specifically, the method may include segmenting the input audio signal (e.g., by using one or more suitable windows) into multiple audio frames (e.g., 100 ms in size). The method may further include obtaining (e.g., determining, calculating, extracting, etc.) at least one feature parameter from the (segmented) audio frames. In some possible example implementations, the feature parameter thus obtained may be considered as being associated with the type of speech articulation noise event (to be detected). That is, in some possible example implementations, depending on the type of speech articulation noise event (to be detected), it may be necessary to obtain different feature parameters from the audio frames (e.g., in the sense that feature parameters can be selected based on the speech articulation noise event to be detected). The method may further include determining (e.g., detecting, calculating, etc.) the corresponding type of speech articulation noise event within the input audio signal and the corresponding range (e.g., time and / or frequency range) associated with the speech articulation noise event, at least in part, based on the obtained feature parameter.

[0016] As configured above, the proposed method provides an efficient and flexible mechanism for identifying (detecting) multiple potential speech articulation noise events (e.g., artifacts) included within the input audio signal. This facilitates appropriate further enhancement (post-processing) (e.g., attenuation). Consequently, the tedious manual editing / processing previously required for identifying and attenuating multiple noise events in the audio signal can be largely avoided. Simultaneously, the listening experience (from the listener's perspective) can be significantly improved.

[0017] In some example implementations, the determined range may include at least one boundary of the determined speech articulation noise event in the time domain and / or spectral domain. That is, the range so determined by the proposed method may include information indicating one or more boundaries of the (detected) speech articulation noise event. More particularly, as those skilled in the art will understand and recognize, such boundaries may be in the time domain, the spectral domain, or both.

[0018] In some example implementations, the method may further include attenuating the speech articulation noise event based on a determined type and range of the speech articulation noise event. As those skilled in the art will understand and recognize, attenuation can be performed by any suitable means, such as by applying an appropriate attenuation gain based on the determined type and range of the speech articulation noise event.

[0019] In some example implementations, speech articulation noise events may include at least one of the following: mouth click events or speech plosive events. As mentioned above, broadly speaking, there are generally two possible types of unwanted / undesirable "noise" that speech enhancement algorithms typically seek to address: noise generated by background sources and noise generated by articulation. Plosives belong to the second type. These plosives occur when air bursts are generated from the mouth (such as during the articulation of syllables containing "p" or "t") and cause significant oscillations of the microphone diaphragm in the presence of wind impact. As indicated above, in the context of this disclosure, the term "plosive" is broadly used to include any air burst emanating from the mouth that causes significant oscillations of the microphone diaphragm (e.g., including short fricatives like "f" and "z"). Even for speech content recorded in a well-controlled acoustic environment, plosives often produce a sudden low-frequency boost, a so-called "plop," resulting in an unpleasant listening experience. Mouth clicks, on the other hand, are another type of transient sound caused by speech articulation using the tongue / teeth / lips mixed with saliva. These mouth clicks can occur in both speech and non-speech segments and are typically audible through headphones / earphones for high SNR recordings. The mouth clicks are usually short, typically lasting between 10 ms and 100 ms, and may also appear as several consecutive transients. Of course, as those skilled in the art will understand and recognize, the methods proposed can also be applied to detect (and optionally attenuate) any other suitable speech articulation noise events.

[0020] In some example implementations, speech articulation noise events may include one or more mouth click events. Specifically, one or more mouth click events may include at least one of the following: non-verbal click events, verbal click events, or lip smacking events. Broadly speaking, as those skilled in the art will understand and recognize, lip smacking can be considered a specific type of non-verbal click in some cases, which often occur just before the start of speech. Lip smacking is often intentional and therefore manifests as a strong and prolonged transient event. In the context of the methods proposed by this disclosure, lip smacking events can often be detected separately from non-verbal click events.

[0021] In some example implementations, after segmenting the input audio signal into multiple audio frames, the method may further include classifying (e.g., determining) the audio frames as speech frames or non-speech frames. That is, the segmented audio frames may be individually determined as speech frames (i.e., containing speech) or non-speech frames (i.e., not containing speech) based, for example, whether the audio frame contains speech. As those skilled in the art will understand and recognize, such classification can be performed in any suitable manner.

[0022] In some example implementations (without intended limitation), the input audio signal can be identified and segmented into speech frames and non-speech frames using a speech activity detector (VAD). That is, the VAD can be used to identify whether each (segmented) audio frame / block (e.g., short audio frame / block) contains speech. Mouth clicks present in non-speech portions can be referred to as "non-speech clicks," and mouth clicks present in speech portions can be referred to as "speech clicks," and these two types of mouth clicks are detected separately. As explained above, lip smacking is a specific type of non-speech click (typically occurring just before the start of speech), and this specific non-speech click can be detected separately from non-speech clicks in the context of this disclosure.

[0023] In some example implementations, the split can be performed by using two different window sizes. In particular, one of the two window sizes can be shorter (smaller) than the other.

[0024] In some example implementations, a shorter (smaller) window size can be (primarily) used to detect speech click events in speech frames, and a longer window size can be (primarily) used to detect non-speech click events in non-speech frames. This allows for efficient and reliable detection of both short and long transient events. In some possible implementations, sufficiently small (one or more) jump sizes can optionally be used to achieve fine temporal resolution, as those skilled in the art will recognize.

[0025] In some example implementations, obtaining at least one feature parameter from an audio frame may include: for each audio frame, obtaining at least one kurtosis metric based on the temporal sample amplitude of the audio frame. Additionally, determining the corresponding type and range of speech articulation noise events in the input audio signal based on the obtained feature parameter may include: comparing the obtained kurtosis metric with a predefined kurtosis threshold; and if the kurtosis metric exceeds the predefined kurtosis threshold, determining that the audio frame includes a mouth click event, and determining the start and end boundaries of the mouth click event based on the corresponding positions where the kurtosis metric rises above and falls below the predefined kurtosis threshold. Notably, by using a kurtosis metric, an estimation (e.g., determination) of a first (coarse) range of (multiple) mouth click events can be achieved efficiently, enabling further refinement if necessary.

[0026] In some example implementations, obtaining at least one feature parameter from an audio frame may include: for each speech frame, obtaining a corresponding first kurtosis metric of the residual approximation without speech harmonic components and the (time-domain) sample amplitude of the residual approximation. Additionally, determining the corresponding type and range of speech articulation noise events in the input audio signal based on the obtained feature parameter may include: comparing the obtained first kurtosis metric with a first predefined kurtosis threshold; and if the first kurtosis metric exceeds the first predefined kurtosis threshold, determining that the speech frame includes a speech click event, and determining the start and end boundaries of the speech click event based on the corresponding positions where the first kurtosis metric rises above and falls below the first predefined kurtosis threshold. As described above, by using the kurtosis metric, a first (coarse) range of (multiple) mouth click events can be estimated (e.g., determined) in an efficient manner, enabling further refinement if necessary.

[0027] In some example implementations, the residual approximation without speech harmonic components can be a second-order waveform difference.

[0028] In some example implementations, the method may further include obtaining a second kurtosis metric from the residual sample amplitude of the speech frame. Specifically, the type and extent of the speech articulation noise event may be determined based on the second kurtosis metric relative to the first kurtosis metric. As a non-limiting example, determining the type and extent of the speech articulation noise event based on the second kurtosis metric relative to the first kurtosis metric may involve determining the type and extent of the speech articulation noise event based on the difference between the second kurtosis metric and the first kurtosis metric.

[0029] In some example implementations, the method may further include refining (e.g., limiting) the determined (coarse) range of the speech click event by: locating a sample location with the largest second-order difference within the determined range of the speech click event; and determining the refined range of the speech click event by applying a predefined speech click event duration (e.g., 5 ms) around the located sample location (e.g., in front of and behind it, possibly centered). As another non-limiting example, the refined range of the speech click event may be determined to be half the predefined speech click event duration (e.g., 2.5 ms) before the located sample location and half the predefined speech click event duration (e.g., 2.5 ms) after the located sample location. Of course, any other suitable measures may be employed depending on the respective implementation.

[0030] In some example implementations, the method may further include determining the range of the speech click event based on the minimum / maximum rate of change calculated from local minima and local maxima in the speech frame. Broadly speaking, this range determination (or refinement) process can generally be viewed as a method for detecting rapid modulation within a (coarse) click range. Specifically, in some possible implementations, the corresponding zero-crossing rate, referred to below as the "minimum / maximum rate of change," can be used to characterize the speed of modulation by converting the local minima / maximum to, for example, -1 and +1 values.

[0031] In some example implementations, obtaining at least one feature parameter from an audio frame may include: for each non-verbal frame, obtaining a corresponding third kurtosis metric of the temporal sample amplitude in the non-verbal frame. Additionally, determining the corresponding type and range of speech phonation noise events in the input audio signal based on the obtained feature parameter may include: comparing the obtained third kurtosis metric with a second predefined kurtosis threshold; and if the third kurtosis metric exceeds the second predefined kurtosis threshold, determining that the non-verbal frame includes a non-verbal click event; and determining the start and end boundaries of the non-verbal click event based on the corresponding positions where the third kurtosis metric rises above and falls below the second predefined kurtosis threshold.

[0032] In some example implementations, the method may further include merging two adjacent nonverbal click events (e.g., for attenuation purposes) into a single verbal click event if the two adjacent nonverbal click events are within a predefined gap threshold. Generally, nonverbal clicks tend to be relatively long (e.g., 50 ms). Therefore, in some cases, merging adjacent clicks within a predefined gap or threshold (e.g., 25 ms) can be beneficial.

[0033] In some example implementations, the method may further include: for a non-verbal click event identified in a non-verbal frame immediately preceding a verbal frame, calculating a high / low frequency peak ratio as the amplitude ratio between the maximum peak value above a predefined frequency and the maximum peak value below a predefined frequency; and if the calculated high / low frequency peak ratio is higher than a predefined ratio threshold, determining the non-verbal click event as a lip-smacking event.

[0034] In some example implementations, the high / low frequency band peak ratio can be calculated as the amplitude ratio between the maximum peak value above a predefined frequency (e.g., 1.5 kHz) and the maximum peak value below a predefined frequency but above another predefined low frequency (e.g., 100 Hz). Generally, the predefined frequency can be chosen as the harmonic-dominant threshold frequency. Of course, as those skilled in the art will understand and recognize, any other suitable calculation method can be employed depending on the various implementations and / or requirements.

[0035] In some example implementations, the method may further include refining the range of the determined smacking events based on the high / low frequency band peak ratio, spectral slope, and / or energy envelope.

[0036] In some example implementations, the range of refined lip-smacking events determined may include: extending the end position of the lip-smacking event determined by using a third kurtosis metric, provided that the following conditions are met: the high / low frequency band peak ratio is higher than a predefined ratio threshold, the spectral slope is lower than a predefined slope threshold, and / or the energy in the energy envelope is reduced.

[0037] In some example implementations, the method may further include determining speech articulation noise events based on the centroid (COG) calculated for the speech frame according to another predefined threshold, to distinguish mouth-clicking events from speech transients. Generally, speech transients can often share similarities in nature with mouth-clicking sounds, but can typically be different in magnitude or spectral characteristics. Speech transients can be identified based on the evolution of the VAD and / or COG (signal average time) of the short-time speech waveform (the waveform of a short frame in the time domain), thus avoiding false alarms detected as mouth-clicking sounds.

[0038] In some example implementations, the method may further include attenuating one or more determined mouth click events based on corresponding spectral gains, which are obtained from the spectral envelope of an audio frame containing the detected mouth click events and a target envelope calculated based on a corresponding reference frame.

[0039] In some example implementations, for each detected mouth click event, the reference frame may include audio frames preceding and following the audio frame containing the detected mouth click event. Further, the target envelope can be calculated by interpolating the spectral envelope of the reference frame. Of course, as those skilled in the art will understand and recognize, any other suitable calculation method may also be employed depending on the appropriate implementation and / or requirements.

[0040] In some example implementations, attenuation can be applied to frequency bands above a predefined high-frequency threshold (e.g., 4 kHz). More specifically, in some possible implementations, further constraints can be optionally applied to speech clicks to allow only high-frequency attenuation (e.g., above 4 kHz) in order to avoid unintentional modification of speech harmonics.

[0041] In some example implementations, the method may further include replacing one or more determined mouth click events based on corresponding adjacent audio frames. More specifically, in some possible implementations, for the correction of speech clicks, it is also possible to use autoregressive modeling or a granular approach similar to pitch-synchronized waveform modeling. That is, given the location of the click event, the local periods on the left and right can be estimated. By comparing adjacent periods, a “waveform slice” matching the relative click location within the period can be used to replace clicks with simple cross-gradients. In some possible implementations, to select the left or right period for correction, the period with the smaller waveform difference can simply be selected. Of course, as those skilled in the art will understand and recognize, any other suitable means may be employed depending on the respective implementation and / or requirements.

[0042] In some example implementations, speech articulation noise events may include at least one speech plosive event. Additionally, obtaining at least one feature parameter from an audio frame may include obtaining a corresponding low-frequency energy (LFE) metric for each audio frame to identify outliers.

[0043] In some example implementations, the LFE metric can be calculated in the time domain or the spectral domain. As those skilled in the art will understand and recognize, depending on the appropriate implementation and / or requirements, any suitable means can be used to calculate the LFE metric. As a non-limiting example, in some possible implementations, for the time domain case, the LFE can be calculated as the root mean square (RMS) energy of the low-pass filtered signal. In some possible implementations, for example, the low-pass filter can be a fourth-order Butterworth filter with a predefined cutoff frequency of, for example, 80 Hz. In some other possible implementations, for the spectral domain case, the LFE can be calculated as the RMS energy below the cutoff frequency based on the spectrum.

[0044] In some example implementations, the method may further include determining the range of speech plosive events based on outliers identified from the LFE metric and a threshold calculated based on the LFE metric, or based on an LFE ratio calculated from previous and current audio frames.

[0045] In some example implementations, the method may further include obtaining a corresponding maximum zero-crossing metric (ZCM) for each audio frame to refine the range of speech plosive events already determined based on the LFE metric. Specifically, the ZCM metric can be viewed as indicating the length of the maximum interval between consecutive zero-crossings within an audio frame. In some possible implementations, the ZCM metric may be further normalized by the window size (e.g., the size of the window used to segment the audio frames).

[0046] In some example implementations, the method may further include attenuating a determined speech plosive event. The attenuation can be performed in the time domain or the spectral domain.

[0047] In some example implementations, time-domain attenuation can be performed by applying a high-pass filter (e.g., a Butterworth high-pass filter). Specifically, in some possible implementations, the cutoff frequency of the filter can be determined based on the ZCM metric of audio frames within the range of a defined speech plosive event; and the order of the filter can be determined based on the LFE metric of audio frames within the range of a defined speech plosive event. Of course, as those skilled in the art will understand and recognize, depending on the various implementations and / or requirements, any other suitable high-pass filter, or more generally, any other suitable time-domain attenuation, can be determined and used.

[0048] In some example implementations, spectral attenuation can be performed by using an overlapped summative short-time Fourier transform (STFT) with adaptive spectral slope and frequency.

[0049] In some example implementations, spectral domain attenuation may involve processing audio frames with a Fast Fourier Transform (FFT), applying an attenuation gain with an adaptive slope and frequency, applying an inverse FFT, windowing, and overlapping summation to produce an attenuated output audio signal. Specifically, in some possible implementations, the frequency may be determined based on a ZCM metric of audio frames within a defined range of speech plosive events; and the slope may be determined based on an LFE metric of audio frames within a defined range of speech plosive events. Of course, as those skilled in the art will understand and recognize, any other suitable spectral domain attenuation may be employed depending on the appropriate implementation and / or requirements.

[0050] In some example implementations, the method may further include applying noise spectrum estimation to limit the attenuation gain to prevent oversuppression. That is, in some possible implementations, noise spectrum estimation may be used to limit the gain reduction so that the attenuation does not affect the overall spectral distribution of the noise spectrum, particularly in the low-frequency region.

[0051] As configured above, the method proposed in this disclosure typically attenuates faster pops with higher cutoff frequencies, thus effectively adapting to the pitch of the speaker's speech. Furthermore, the method also attenuates stronger pops with steeper cutoff frequency slopes, thus effectively adapting to weak yet intense plosives.

[0052] In some example implementations, the method may further include applying a content classifier (e.g., VAD) to audio frames to distinguish between speech frames and non-speech frames in order to determine speech plosive events. More specifically, in some possible implementations, when the techniques described above are applied to content including music or both speech and music, the proposed algorithm may be sensitive to low-frequency transients (such as low-frequency transients generated by kicking drums or bass). To address this issue, in some possible implementations, a content classifier (e.g., a speech / music activity detector) that computes the probability p(n) that a given frame n contains speech can be used to modify detection or attenuation parameters to ensure that the musical content is unaffected by deplosive processing.

[0053] In some example implementations, spectral domain attenuation may involve: generating multiple approximately equivalent rectangular bandwidth (ERB) spaced frequency bands below a predefined frequency threshold and multiple frequency bands above the predefined frequency threshold, which lies within the frequency range of a defined speech plosive event; applying multiple attenuation gains to the audio signal in each of the frequency bands, wherein the attenuation gains are calculated based on energy calculated for each frequency band; and feeding the attenuated audio samples to a synthesis filter bank to generate an output audio signal. This spectral domain attenuation is typically used when computational complexity allows, compared to the spectral domain attenuation described above.

[0054] In some example implementations, the attenuation gain in each frequency band can be further constrained to prevent the energy of that band from falling below the estimated noise floor in that band. In other words, in some possible implementations, the (attenuation) gain can be clipped to ensure that the power in each frequency band is not reduced below the estimated noise floor in the corresponding band. Generally, this will avoid audible noise drop when there are pops in the presence of significant background noise. As those skilled in the art will understand and recognize, noise (or noise floor) can be estimated using any suitable means.

[0055] In some example implementations, the method may further include calculating time-smoothed low-frequency energy estimates for audio samples above the estimated noise floor to distinguish speech plosive events in the input audio signal from higher-frequency content.

[0056] In some example implementations, the method may further include calculating a speech harmonic protection metric in the spectrum of the input audio signal; and calculating an attenuation gain based on the speech harmonic protection metric and a time-smoothed low-frequency energy estimate.

[0057] In some example implementations, the speech harmonic protection metric can be a periodic metric or a tonality metric.

[0058] In some example implementations, periodicity measures in the spectrum can be calculated from the cepstrum of the audio samples before the final frequency band calculation of the analysis filter bank.

[0059] In some example implementations, the tonality measure in the spectrum can be calculated based on the main lobe of the spectral peak compared to the main lobe of the sine peak, prior to the final bandwidth calculation of the analytical filter bank.

[0060] In some example implementations, the method may further include constraining the calculated attenuation gain based on the immediately adjacent lower frequency band. As a non-limiting example, the gain may be constrained such that for frequency bands above a certain threshold (e.g., 70 Hz), the attenuation of the gain cannot exceed that of the immediately adjacent lower frequency band. Generally, this would enforce attenuation or reduction to follow the physical decrease in plosive energy with frequency. That is, when the energy of a lower frequency band decreases significantly, if the next higher frequency band has more energy, it is more likely to be genuine speech energy rather than plosive-related energy. Broadly speaking, very low frequency bands (below, for example, 70 Hz) may not follow this trend; for example, excessive 60 Hz power hum can make a frequency band louder, or a DC blocking filter can attenuate the lowest frequency band, and this should not limit the attenuation of plosive energy.

[0061] According to another aspect of this disclosure, a method is provided for performing automatic audio enhancement on an input audio signal to detect and / or attenuate at least one speech articulation noise event contained therein. As those skilled in the art will understand and recognize, automatic audio enhancement may involve any other suitable audio enhancement means. In particular, the speech articulation noise event may include at least one speech plosive event.

[0062] More specifically, the method may include: generating multiple approximately equivalent rectangular bandwidth (ERB) spaced frequency bands below a predefined frequency threshold and multiple frequency bands above the predefined frequency threshold, which lies within the frequency range of speech plosive events, by using an analysis filter bank. The method may further include: applying multiple attenuation gains to the audio signal in each of the frequency bands, wherein the attenuation gains are calculated based on energy calculated for each frequency band. The method may further include feeding attenuated audio samples to a synthesis filter bank to generate an output audio signal.

[0063] As described above, broadly speaking, the proposed method provides an efficient and flexible mechanism for identifying (detecting) and attenuating (multiple) possible / potential speech articulation noise events (e.g., speech plosive events) included within the input audio signal. Therefore, the tedious manual editing / processing previously required for identifying and attenuating (multiple) noise (e.g., plosive) events in the audio signal can be largely avoided. Simultaneously, the listening experience (from the listener's perspective) can be significantly improved.

[0064] In some example implementations, the attenuation gain in each frequency band can be further constrained to prevent the energy of that band from falling below the estimated noise floor in that band. In other words, in some possible implementations, the (attenuation) gain can be clipped to ensure that the power in each frequency band is not reduced below the estimated noise floor in the corresponding band. Generally, this will avoid audible noise drop when there are pops in the presence of significant background noise. As those skilled in the art will understand and recognize, noise (or noise floor) can be estimated using any suitable means.

[0065] In some example implementations, the method may further include calculating time-smoothed low-frequency energy estimates for audio samples above the estimated noise floor to distinguish speech plosive events in the input audio signal from higher-frequency content.

[0066] In some example implementations, the method may further include calculating a speech harmonic protection metric in the spectrum of the input audio signal; and calculating an attenuation gain based on the speech harmonic protection metric and a time-smoothed low-frequency energy estimate.

[0067] In some example implementations, the speech harmonic protection metric can be a periodic metric or a tonality metric.

[0068] In some example implementations, periodicity measures in the spectrum can be calculated from the cepstrum of the audio samples before the final frequency band calculation of the analysis filter bank.

[0069] In some example implementations, the tonality measure in the spectrum can be calculated based on the main lobe of the spectral peak compared to the main lobe of the sine peak, prior to the final bandwidth calculation of the analytical filter bank.

[0070] In some example implementations, the method may further include constraining the calculated attenuation gain based on the immediately adjacent lower frequency band. As a non-limiting example, the gain may be constrained such that for frequency bands above a certain threshold (e.g., 70 Hz), the attenuation of the gain cannot exceed the attenuation of the immediately adjacent lower frequency band. Generally, this would enforce attenuation or reduction to follow the physical decrease in plosive energy with increasing frequency. That is, when the energy of a lower frequency band decreases significantly, if the next higher frequency band has more energy, it is more likely to be genuine speech energy rather than plosive-related energy. Broadly speaking, very low frequency bands (below, for example, 70 Hz) may not follow this trend; for example, excessive 60 Hz power hum can make a frequency band louder, or a DC blocking filter can attenuate the lowest frequency band, and this should not limit the attenuation of plosive energy.

[0071] In some example implementations, the input audio signal can be processed continuously at a predefined look-ahead frame (window) size (e.g., 50ms).

[0072] According to another aspect of this disclosure, an apparatus is provided including a processor and a memory coupled to the processor. The processor may be adapted to cause the apparatus to perform all the steps of the example methods described throughout the disclosure.

[0073] According to a further aspect of this disclosure, a computer program is provided. The computer program may include instructions that, when executed by a processor, cause the processor to perform all the steps of the example methods described throughout the disclosure.

[0074] According to another aspect, a computer-readable storage medium is provided. This computer-readable storage medium can store the aforementioned computer program.

[0075] It will be understood that the apparatus features and method steps can be interchanged in various ways. In particular, the details of the disclosed methods (various methods) can be implemented by the corresponding apparatus (or system), and vice versa, as will be recognized by those skilled in the art. Furthermore, any statement made above with respect to the methods (various methods) is to be understood to equally apply to the corresponding apparatus (or system), and vice versa. Attached Figure Description

[0076] The following explanation of exemplary embodiments of this disclosure is based on the accompanying drawings, in which:

[0077] Figure 1A This is a schematic illustration of a simplified diagram showing an example of a nonverbal clicking sound according to an embodiment of the present disclosure.

[0078] Figure 1B This is a schematic illustration of a simplified diagram showing an example of a speech click sound according to an embodiment of the present disclosure.

[0079] Figure 1C This is a schematic illustration of a simplified example of smacking lips according to an embodiment of the present disclosure.

[0080] Figure 2 This is a schematic illustration of a simplified diagram showing an example of speech click detection and refinement according to an embodiment of the present disclosure.

[0081] Figure 3 This is a schematic diagram illustrating an example of speech click detection and refinement according to another embodiment of the present disclosure.

[0082] Figure 4 This is a schematic illustration of a simplified diagram showing an example of lip-smacking detection according to an embodiment of the present disclosure.

[0083] Figure 5 This is a schematic illustration of a simplified diagram showing an example of spectral attenuation according to an embodiment of the present disclosure.

[0084] Figure 6 This is a schematic block diagram illustrating an example of a functional overview of the technology according to embodiments of the present disclosure.

[0085] Figure 7 This is a schematic illustration of a simplified diagram showing an example comparison between the maximum zero-crossing value (ZCM) and the zero-crossing rate (ZCR).

[0086] Figure 8 This is a schematic illustration of a simplified diagram showing an example of attenuation of speech plosives according to an embodiment of the present disclosure.

[0087] Figure 9 This is a schematic block diagram illustrating an example of a functional overview of the technology according to embodiments of the present disclosure.

[0088] Figure 10 This is a schematic block diagram illustrating another example of a functional overview of the technology according to embodiments of the present disclosure.

[0089] Figure 11 This is a schematic flowchart illustrating an example of a method according to an embodiment of the present disclosure.

[0090] Figure 12 This is a schematic flowchart illustrating an example of a method according to another embodiment of the present disclosure.

[0091] Figure 13 This is a schematic block diagram illustrating yet another example of a functional overview of the technology according to embodiments of the present disclosure, and

[0092] Figure 14 This is a block diagram of an apparatus for performing a method according to an embodiment of the present disclosure. Detailed Implementation

[0093] The accompanying drawings and the following description are for illustrative purposes only and relate to preferred embodiments. It should be noted that, based on the discussion below, alternative embodiments of the structures and methods disclosed herein will be readily recognized as feasible alternatives that can be employed without departing from the claimed principles.

[0094] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that similar or identical reference numerals may be used in the drawings where feasible, and these reference numerals may denote similar or identical functions. The drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.

[0095] Furthermore, in drawings where connecting elements (such as solid or dashed lines or arrows) are used to illustrate connections, relationships, or associations between two or more other schematic elements, the absence of any such connecting elements does not imply the absence of connections, relationships, or associations. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, where a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or more signal paths (as may be required) to influence communication.

[0096] As indicated above, the increasing volume of speech content, often of varying quality, on various media platforms has reached a point where manual editing no longer seems like a viable solution. Automatic speech enhancement, when done well, can typically maintain the naturalness of the speech and save editing time.

[0097] Broadly speaking, speech enhancement algorithms typically attempt to address two types of unwanted "noise" events: noise generated by background sources and noise generated by vocalization. In particular, plosives and mouth clicks both belong to the second type.

[0098] More specifically, on the one hand, speech plosives typically occur when an air blast is generated from the mouth (such as during the pronunciation of syllables containing "p" or "t") and causes a large vibration of the microphone diaphragm in the presence of wind impact. As stated above, in the context of this disclosure, the term "plosive" can be used broadly to include any air blast emitted from the mouth that causes a large vibration of the microphone diaphragm (e.g., including short fricatives like "f" and "z"). Even in speech content recorded in a well-controlled acoustic environment, plosives can often produce a sudden low-frequency boost, the so-called "plop," resulting in an unpleasant auditory experience. This can be seen from, for example... Figure 8 The simplified diagram 8200 shows an illustrative example of speech plosive events (in particular, the white portion in the low-frequency part, which will be discussed in more detail later).

[0099] Several recording techniques have been proposed to reduce plosive intensity, such as using pop filters or deflectors, off-axis speaking, etc. However, pop reduction is not as effective as expected due to practical reasons: for example, the speaker's or (voice) actor's posture cannot be fixed. Therefore, signal processing tools are needed to improve the quality of such recordings. Two main feasible methods exist for automatic plosive detection, including detection based on simple features and detection based on telephone calls (multidimensional features for speech recognition). While telephone-based detection seems to have its advantages in identifying the precise time span of plosive events, it is more complex and therefore requires more computational resources. Detection based on simple features is usually local and does not require fine-tuning of plosive event boundaries. Another feasible solution typically provides three user parameters (sensitivity / intensity / frequency limits) for its plosive removal module. However, to obtain the best results, users may need to manually edit the automated curves of these parameters, as the intensity and frequency of plosives vary in the same recording, and users may want to attenuate these plosives accordingly. As a result, this process can be time-consuming.

[0100] On the other hand, mouth clicks are typically transient sounds caused by the mixing of saliva with the tongue / teeth / lips during speech production. These mouth clicks usually occur in both speech and non-speech parts and are generally audible through headphones / earphones for high SNR recordings. Mouth clicks are usually short, typically lasting between 10ms and 100ms, and can also appear as several consecutive transients. In the context of professional recordings such as TV / movie / game dialogues, the absence of mouth clicks can be considered extremely demanding in terms of speech quality. Now, even for user-generated content, mouth clicks often become quite noticeable due to the prevalence of headsets / earphones.

[0101] In the context of this disclosure, the proposed methods generally seek to address three types of mouth clicks, namely: 1) nonverbal clicks; 2) verbal clicks; and 3) lip smacking (which can also be considered a special type of nonverbal click).

[0102] Now refer to the attached diagram, Figure 1A An example of a nonverbal clicking sound is illustrated schematically (e.g., at approximately 0.1 s); Figure 1B An example of a speech click is schematically illustrated (particularly shown at approximately 0.7056s at the end of the leftmost loop, indicated by a circle); and Figure 1C An example of lip smacking is illustrated schematically (specifically shown as a strong transient just before a speech segment, at approximately 2.1 seconds).

[0103] Several recording techniques have been proposed to reduce mouth clicks in professional male / female voice actors. However, in most cases, there is no way to control the speaker's mouth / lip condition. Manual editing for post-production can be tedious, making it impractical to process hundreds or thousands of dialogues. Therefore, signal processing tools are needed to more efficiently correct mouth clicks. However, there is currently very little available academic research on mouth click detection. Lip smacking detection can be considered a similar problem, but the transient energy is usually much larger, so the corresponding methods may not be directly applicable to small transients such as mouth clicks. Furthermore, in the context of digital audio restoration, "click removal" is often used to remove impulse noise typically present in phonograph recordings. When the duration of the damaged audio is long, the problem becomes a general signal interpolation / extrapolation issue.

[0104] In view of this, this disclosure proposes a method for performing automatic audio enhancement on one or more input audio signals including one or more such speech pronunciation (related or induced) noise events. More specifically, this disclosure seeks to provide a method for automatically detecting and attenuating speech plosives and mouth clicks, as well as other noise events included in the input audio signals, thereby avoiding manual editing while maintaining or even improving audio quality for the listener.

[0105] First, methods related to "removing click sounds" according to embodiments of this disclosure will be discussed.

[0106] In a broad sense, the method for automatic detection and attenuation of mouth clicks described in this disclosure mainly comprises two key aspects. Firstly, the detection algorithm typically targets mouth clicks in non-speech regions and mouth clicks in speech regions, respectively. A kurtosis metric of the waveform amplitude is typically used as the primary standard applicable to both the original waveform and its second-order difference, where the second-order difference serves as an approximation of the non-harmonic signal portion. The roughly detected click location is further refined to more accurately define the click sample region. Secondly, mouth click attenuation is typically based on spectral gain attenuation, which is obtained by interpolating the spectral envelope across short time frames containing the (detected) click.

[0107] Now refer to Figure 6 The “click removal” method is discussed in more detail, and the figure generally provides a schematic functional overview of the (click removal) technique according to embodiments of the present disclosure.

[0108] More specifically, as shown in box 6010, the input audio signal can be provided, for example, as an input file or stream (or in any other suitable form). Depending on its form (e.g., format), the input audio signal may need to undergo a suitable segmentation process to be divided into, for example, multiple (short-time) audio frames (e.g., having equal or different frame sizes).

[0109] It is worth noting that, prior to proceeding to subsequent de-clicking processing, an optional denoising process (shown as dashed box 6020) can be applied to the input signal to better reveal potential mouth clicks.

[0110] Then, a speech activity detector (VAD), as illustrated in box 6030, can identify each short time frame (audio frame) of the speech signal as containing or not containing speech. This allows for separate processing of mouth clicks in the speech portion (e.g., frames) and the non-speech portion. Mouth clicks present in the non-speech portion are generally referred to as "non-speech clicks" (e.g., such as...). Figure 1A (As shown) and the mouth clicks present in the speech segment are called "speech clicks" (e.g., as shown) Figure 1B As shown above, these two lip-smacking sounds are detected separately. In the context of this disclosure, lip-smacking is generally considered a specific type of nonverbal smacking sound that typically occurs just before the start of speech. Lip-smacking is often intentional and therefore manifests as a strong and prolonged transient event (e.g., as shown). Figure 1C(As shown in the diagram). Therefore, using two (different) window sizes can be considered advantageous for detecting both short and long transient events. Specifically, in some possible implementations, the shorter (smaller) window size can be (primarily) used to detect speech click events in speech frames, and the longer window size can be (primarily) used to detect non-speech click events in non-speech frames. In this way, both short and long transient events can be detected efficiently and reliably. Additionally, in some possible implementations, a sufficiently small jump size can also be used to achieve fine temporal resolution.

[0111] On the other hand, for the detection of non-verbal clicking events, although the energy is weak, they are generally stronger than background noise and can therefore be identified by transient detection algorithms. In this disclosure, it is generally proposed to use a (first) kurtosis metric k of the short-time waveform (time domain) amplitude. W (Box 6040) is used to identify and distinguish between kurtosis (also known in some cases as large outliers) and flat distributions. Then the kurtosis metric k can be used. W A comparison with a predefined threshold (box 6100) is made to detect (or determine) mouth clicks (in the present case, nonverbal clicks), as shown in box 6060. The start and / or end positions of the detected nonverbal click events can then be simply defined as the positions where their kurtosis rises above and / or falls below the predefined threshold. Generally, nonverbal clicks may tend to be relatively long (e.g., 50 ms), and therefore, in some cases, merging (e.g., for attenuation purposes) adjacent click events within a predefined interval / threshold (e.g., 25 ms) can be beneficial.

[0112] On the other hand, regarding the detection of speech clicks, it is generally considered that mouth clicks in spoken speech tend to exhibit rapid modulation and are therefore more difficult to detect. Ideally, if speech harmonics are well modeled, it might rely on the residual waveform (minus the harmonics) to detect any abrupt changes. However, this would typically involve using robust F0 (also known as fundamental frequency) / harmonic estimation algorithms that could increase the complexity of the detection algorithm. Therefore, in this disclosure, it is generally proposed to approximate the removal of slowly changing signal components (harmonics) using a second-order sample difference (box 6050), allowing the potential transients to be revealed. Similar to the detection of non-speech clicks, a (second) short-time kurtosis metric k can be calculated for the difference (residual) waveform. D (again, box 6040). However, as those skilled in the art will understand, other forms of residual signals besides the second-order sample difference can be used at this stage, as long as they allow for the identification of potential transients.

[0113] In some possible implementations, the kurtosis metric k can be related to (or relative to) (the first) kurtosis metric. W Evaluation (Second) Kurtosis metric k D More specifically,

[0114] k R =k D -α×k W (1)

[0115] Where α is (for example, predefined) a weighting parameter.

[0116] Since speech crackling typically occurs within the audible portion of a speech pattern, the harmonic energy can be quite strong and therefore exhibits a smooth amplitude distribution (which usually means k...). W (This will be relatively small). Therefore, this implicitly avoids the speech transient (which usually means a peaked amplitude distribution or, in other words, k). W The large sound detected was a clicking sound from the mouth. In other words, k R The click of speech will be relatively loud, but the transient of speech will be relatively small, which allows the two to be distinguished.

[0117] Furthermore, speech clicks may tend to be very short, and therefore refining the (rough) click event location defined above with better sample precision may often be necessary.

[0118] A simple approach could be to locate the maximum second-order difference (which typically implies the fastest change) within the coarse click range detected by kurtosis. Then, a predefined speech click duration, such as 5 ms, could be used to determine refined start and / or end positions around the sample location of the fastest change. As those skilled in the art will understand and recognize, this can be achieved by any suitable means. For example (without limitation), such a speech click duration (e.g., 5 ms) could simply be evenly divided before and after the sample location of the fastest change in the sense that the interval corresponding to the speech click duration is centered at the sample location of the fastest change.

[0119] exist Figure 2 An example of such a refinement process is illustrated in the diagram. Specifically, in... Figure 2 In the example, waveform 2100 typically shows the original input audio waveform, while waveform 2200 typically shows the second-order difference waveform obtained from the original waveform 2100. Then, as explained above, the refined range 2300 of the nonverbal click event can be determined based on the second-order difference waveform 2200.

[0120] Another possible refinement method could be to detect the fastest modulation within the coarse click range. In particular, by converting local minimum / maximum values ​​into, for example, –1 and +1 values ​​(or any other suitable values, such as those with different signs and equal magnitudes), the corresponding zero-crossing rate (ZCR) (also referred to below as the “minimum / maximum rate of change”) can be used to characterize the modulation speed.

[0121] exist Figure 3 An example of this refinement process is illustrated in the diagram. Specifically, in... Figure 3 In the example, similar to Figure 2 In the example shown, waveform 3100 typically represents the original input audio waveform. However, in this refinement process, instead of using the second-order difference, a minimum / maximum rate of change waveform 3200 is obtained from the original waveform 3100. Subsequently, the refinement ranges 3310, 3320, and 3330 of the nonverbal click event can be determined based on the minimum / maximum rate of change waveform 3200, as... Figure 3 As shown in the figure.

[0122] In some possible implementations, kurtosis thresholds and minimum / maximum rates of change can be used in combination to detect speech clicks with better accuracy.

[0123] Regarding the lip-smacking detection described above, the lip-smacking event typically manifests as a strong transient event that usually occurs just before speech (e.g., Figure 1C (As shown in the example). To distinguish the smacking event from the two aforementioned click events (i.e., verbal clicks and regular nonverbal clicks), it is conceivable to rely on, for example, by using spectral features to verify a sudden change in resonance. In this disclosure, the use of spectral slope (hereinafter also referred to as "SpS") and high / low frequency band peak ratio (hereinafter also referred to as "ratioHL") is generally proposed.

[0124] Generally, in some possible implementations, the feature ratioHL can be calculated as a frequency higher than the predefined frequency freq. HL (e.g., 1.5kHz) maximum peak value and below freq HL The amplitude ratio between the maximum peak values. In some possible implementations, a frequency higher than the (predefined) low frequency freq is further selected. L The maximum peak value in a lower frequency band (e.g., 100 Hz) to avoid low-frequency noise may be preferred.

[0125] In some possible implementations, for a nonverbal click sound detected just before speech, if ratioHL > th R (where, th) RIf a threshold is set (which could be a predefined threshold), then it can subsequently be considered a candidate for smacking lips (e.g., such as...). Figure 6 (as shown in box 6070).

[0126] Typically, when a lip smacking occurs, the high / low frequency peak ratio ratioHL may tend to become larger, and the spectral slope may tend to become steeper due to high-frequency resonance. Since lip smacking events are usually much larger than small (regular) mouth clicks (e.g., typically lasting 100 ms), it is often possible to refine the start / end positions of (multiple) events based on features including ratioHL, SpS, and energy envelope.

[0127] In some possible implementations, the initial (coarse) ending position can be continuously extended (i.e., via k) provided that one of the following conditions is met. W Detected): 1) ratioHL>th R ;2)SpS <th S , among which, th S It is a predefined threshold; and 3) energy reduction.

[0128] Additional verification of the end position of the expansion can be performed by comparing the skewness before and after refining the event location. That is, the expansion of the event may only add samples with smaller amplitudes, making the sample amplitude distribution "more skewed".

[0129] Of course, as those skilled in the art will understand and recognize, any other suitable implementation may be adopted as appropriate.

[0130] Figure 4 This is a schematic illustration of a simplified diagram showing an example of lip-smacking detection according to an embodiment of the present disclosure. Specifically, Figure 4 The waveforms in the figure show, in a general and illustrative manner, the original waveform, spectral slope (SpS), energy, and high / low frequency band peak ratio (ratioHL).

[0131] In some cases, it may be necessary or desirable to avoid detecting speech transients as clicks. Specifically, speech transients can often share some degree of similarity in nature with mouth clicks, but on the other hand, they can typically be different in magnitude and / or spectral characteristics. Therefore, speech transients can be actively identified and thus falsely detected as mouth clicks can be avoided based on the evolution of the VAD and / or center of gravity (COG, which can often be considered as the signal averaging time) of the short-time speech waveform.

[0132] In some possible implementations, COG can be calculated as follows:

[0133]

[0134] The beginning of the transient on the right side of the window suggests that it can be used with the help of COG > th. COG (where, for example, th) COG =0.2) is a positive value for transient detection. More specifically, when VAD indicates no speech, nonverbal clicks will be processed regardless of COG. Conversely, when VAD indicates speech, if any COG is close to the start of the click event and is above th... COG If the value is set to 0, the clicking sound will not be processed.

[0135] In a broader sense, the reason for using a “normalized” metric (i.e., COG) is to more effectively handle speech transients, while using a “non-normalized” metric (i.e., kurtosis) often helps to select a range of degrees of transientity to correct for.

[0136] After mouth clicking sounds (including non-verbal clicking sounds, verbal clicking sounds, and lip smacking sounds) have been detected, the attenuation (or correction) of these clicking sounds (i.e., clicking sound removal processing) can be the next step.

[0137] More specifically, the click removal processing proposed in this disclosure is generally based on the observed spectral envelope (hereinafter referred to as "E") and the target envelope (hereinafter referred to as "E"). T The spectral gain attenuation obtained () Figure 6 (in the frame 6090), as in Figure 6 As illustrated in box 6080. More specifically, in some possible implementations, given the start / end positions of the click sound(s), it is typically proposed that a block before the click sound (with envelope E0) and after the click sound (with envelope E1) be considered as a reference frame. The spectral envelopes of these two reference frames can then be used to estimate the target envelope covering each short-time block of the click sound event. Then, in some possible implementations, the target envelope can be simply calculated as a linear interpolation of the two reference envelopes. Accordingly, the spectral gain is then defined by dividing the target envelope by the observed envelope, where the constraint is that only attenuation is allowed. That is, for each segment k at a given frame b across a total of B frames, the attenuation gain can be calculated as:

[0138] in,

[0139] Of course, as those skilled in the art will understand and recognize, any other suitable implementation may be adopted as appropriate.

[0140] In particular, for speech clicks, another constraint can be optionally applied to allow only high-frequency attenuation (e.g., above 4kHz) in order to avoid unintentionally modifying speech harmonics.

[0141] In some possible implementations, when residual estimation (with harmonic components removed) is available (e.g., as...), Figure 13 As illustrated in box 13040, envelope attenuation can be applied to the residual signal and then the harmonic components can be added back as the processed output (e.g., as shown in box 13040). Figure 13 (illustrated in block 13090).

[0142] In some possible implementations, other algorithms, such as autoregressive modeling or granular methods similar to pitch-synchronized waveform modeling, can also be used for speech click correction. Specifically, given the location of the click event, local periods on the left and right can be estimated. By comparing adjacent periods, a “waveform slice” matching the relative click location within a period can be used to replace clicks with simple cross-gradients. To select the left or right period for correction, the period with the smaller waveform difference can simply be chosen. In cases where continuous clicks are present, the above methods may sometimes be less effective, and more generative methods may then become better options.

[0143] Figure 5 This is a schematic illustration of an example of spectral attenuation according to an embodiment of the present disclosure, wherein the observed spectral waveform, the processed spectral waveform, the observed envelope, and the target envelope are illustrated illustratively, respectively. (As can be seen from...) Figure 5 As seen in the example, the spectral region of attenuation (detection) of the click sound. For completeness, however, it should be noted that even as currently... Figure 5 The examples shown may relate to "de-clicking," and similar or analogous attenuation concepts can also be applied to "de-plosive" scenarios. In some implementations, this may involve, for example, smoothing the envelope of the residual spectrum, as those skilled in the art will recognize.

[0144] Next, methods related to "removing plosive sounds" according to embodiments of this disclosure will be discussed.

[0145] Similar to the above, the method for automatic detection and adaptive attenuation of speech plosives described in this disclosure, in a broad sense, also mainly includes two key aspects. Firstly, the feature is measured using the maximum zero-crossing value (ZCM). Compared to the zero-crossing rate (ZCR) metric, it can be seen that ZCM only takes the maximum zero-crossing length. Therefore, ZCM can generally be considered robust to noise crossover information, especially when used in an averaging manner as in the case of ZCR. Secondly, accurate detection of plosive event boundaries can be performed based on low-frequency energy (LFE) and ZCM. Specifically, outliers from the observed low-frequency energy distribution (e.g., for all short frames across a file or recording) can be selected as possible (annoying) plosive events, and ZCM can then be used to refine the event temporal location / boundary. Finally, plosive attenuation can generally be performed in the time or spectral domain based on high-pass filtering, where the filter order is adapted to the LFE and the filter frequency is adapted to the ZCM of the detected plosive.

[0146] Now, refer to Figure 9 and / or Figure 10 The "plosive removal" method is discussed in more detail, and these two figures provide a schematic functional overview of the (plosive removal) technique according to embodiments of the present disclosure. In a broader sense, Figure 9 This can be considered a larger example, while Figure 10 This can be viewed as a more detailed example of a particular possible implementation. Therefore, Figure 9 and Figure 10 The examples shown may simultaneously exhibit some degree of similarity (e.g., in some blocks) and difference (e.g., in some other blocks), as will be understood and recognized by those skilled in the art.

[0147] More specifically, as shown in blocks 9010 or 10010, an input audio signal is provided and can be segmented / divided into multiple (short-time) overlapping audio frames (e.g., having equal frame sizes). This can be achieved in any suitable manner, as will be understood and recognized by those skilled in the art. For example, in some possible implementations, this segmentation of the audio frames can be achieved by performing short-time frame analysis using a Hamming window. In particular, in some possible implementations, the frame size can be set large enough to allow reliable extraction of zero-crossing maximum values. Similarly, the overlap size can be set large enough to track short-time features with fine temporal resolution.

[0148] Subsequently, two short-time features (or sometimes called feature parameters) can be calculated, namely: the low-frequency energy (LFE) illustrated in box 9020 or 10020 and the zero-crossing maximum value (ZCM) illustrated in box 9040 or 10050.

[0149] LFE can be calculated in the time domain or the frequency domain and using any suitable means. In some possible implementations, for the time domain case, LFE can be calculated as the root mean square (RMS) energy of the low-pass filtered signal. In some possible implementations, the low-pass filter can be a 4th-order Butterworth filter with a predefined cutoff frequency of, for example, 80 Hz. On the other hand, in some other implementations, for the frequency domain case, LFE can be calculated as the RMS energy below the cutoff frequency based on the spectrum.

[0150] As mentioned above, ZCM is typically the length of the maximum interval between consecutive zero crossings within a short frame, which may be further normalized by the window size. Notably, the techniques presented in this disclosure generally do not rely on ZCR commonly used in plosive sound detection mechanisms.

[0151] Since low-frequency sudden pops are often the primary problem, pop detection can be initiated by identifying outliers in the observed LFE distribution (boxes 9030 or 10030). In some possible implementations, outliers can be identified based on the concept / principle of standard scores:

[0152]

[0153] Where x is the LFE sample value, μ is its mean, and σ represents the standard deviation.

[0154] If any outliers exist, they can be passed to the next threshold detection stage. Otherwise, it can be assumed that there are no potentially (annoying) pops that require further processing. In a non-restrictive example, outliers can be indicated by z > 1 (or any other suitable value).

[0155] In some possible implementations, the adaptive threshold th LFE The outliers that can be detected can be used to select the dominant component according to the following formula:

[0156] th LFE =α×(maxLFE-th) Z )+th Z (5)

[0157] Where maxLFE is the maximum LFE, and

[0158] th Z =μ+z0×σ (6)

[0159] It is worth noting that here, th zA predefined factor z0 is suitable for values ​​higher than the standard deviation of the mean. The multiplication factor α in equation (5) can be set to adjust the detection sensitivity. In some possible implementations, the multiplication factor α can be set according to the global plosive removal amount parameter as follows:

[0160] α = 1 - amount, where 0 ≤ amount ≤ 1 (7)

[0161] In scenarios requiring low-latency online (real-time) processing, the aforementioned statistical thresholds may not be reliably estimated. Therefore, in some cases, the LFE ratio can be used alternatively for the current frame n according to the following formula:

[0162]

[0163] Otherwise, the ratio can be calculated based on the previously valid LFE.

[0164] The detection function can then be expressed as R > 1 + f(α), where f(α) is a customizable mapping function. Correspondingly, the detection function can also be simply written as R > 1 + α.

[0165] In some possible implementations, frames exceeding a detection threshold can be used to define the signal region considered as a plosive event to be attenuated, which also implicitly defines the (initial) time position of the start and / or end of the plosive event (box 9030 or 10040). However, the event boundaries may need to be further refined (box 9050 or 10060), typically because the actual plosive may start and / or end with very low energy. Therefore, in some possible implementations, for example, a ZCM metric (box 9040 or 10050) can be used to extend the frame boundaries, where ZCM < 0.1 (or any other suitable value).

[0166] Furthermore, similar to the "remove click" scenario, in some cases where two plosive events can overlap or are very close, they can be merged into a single plosive event (e.g., for further "remove plosive" processing).

[0167] Figure 7 An example of a comparison between ZCM and ZCR is illustrated schematically. Specifically, as can be seen from... Figure 7 As seen in the examples, the ZCM diagram 7100 is generally less noisy than the ZCR diagram 7200, and is therefore better suited for identifying potential plosive events.

[0168] After the speech plosive events and their corresponding range / location / boundaries within the audio frame have been identified (box 9080), the attenuation (or correction) of these plosives (i.e., plosive removal processing) can be the next step (box 9110). In some possible specific implementations (e.g., as...), Figure 10 As shown in the diagram, attenuation can be performed by using a high-pass filter (e.g., as illustrated in box 10070).

[0169] In particular, similar to the "removal of clicks" scenario, speech plosives can also be attenuated in the time domain or the spectral domain.

[0170] In a general sense, in some possible implementations, time-domain attenuation can be achieved using a Butterworth high-pass filter (or any other suitable means) with adaptive order and frequency; while spectral-domain attenuation can be achieved using an overlapped summative short-time Fourier transform (STFT) (or any other suitable means) with adaptive spectral slope and frequency.

[0171] Specifically, for both time-domain and spectral-domain attenuation, the attenuation frequency (box 9070) or, in some possible implementations, the filter (cutoff) frequency freq. C (For example, as illustrated in box 10072) can be set to adapt to the "velocity" (box 9070) of the plosive event, which can typically be defined as 1-max(ZCM). plosive ), where ZCM used here is normalized between 0 and 1, and max(ZCM) plosive ) is the maximum ZCM from the start frame to the end frame of the plosive sound event. The mapping can then be defined as:

[0172] freq C =minFreq+speed×(maxFreq-minFreq) (9)

[0173] In some possible implementations, the cutoff frequency freq C It can be further constrained to a predefined range, such as [minFreq = 100Hz, maxFreq = 150Hz]. Of course, any other suitable range may be used depending on the specific implementation and / or requirements.

[0174] For time-domain attenuation, the order of the Butterworth filter can be adapted to the intensity of the plosive event (box 9060). Specifically, the plosive intensity *st* can be defined in some possible implementations as follows:

[0175] st=g(max(LFE plosive )-th Z (10)

[0176] Where max(LFE) plosive ) is the maximum LFE from the start frame to the end frame of the plosive sound event; g(x) is a customizable mapping function mainly used to ensure 0≤st≤1, which can be achieved by simply applying a normalization factor.

[0177] Then, the attenuation gain (as illustrated in box 9090) or, in some possible cases, the filter order (as illustrated in box 10071) can be obtained by mapping:

[0178] order=round(minOrder+st×(maxOrder-minOrder)) (11)

[0179] In some possible implementations, the order can be further constrained to a predefined range, such as [minOrder = 2, maxOrder = 12]. Of course, any other suitable range may be used depending on the specific implementation and / or requirements.

[0180] Furthermore, in some possible implementations, a 10ms crossover gradient region can be used to generate a smooth transition from the input signal to the filtered signal.

[0181] On the other hand, for spectral domain attenuation, the input short-time signal can be processed using a Fast Fourier Transform (FFT) in some possible implementations, followed by applying an attenuation gain with an adaptive cutoff frequency and slope, applying an inverse FFT, and finally applying windowing and overlapping to generate the (attenuated) output. Of course, as those skilled in the art will understand and recognize, any other suitable attenuation mechanism can also be applied depending on the implementation.

[0182] The spectral low-cutoff / high-pass gain slope can also be estimated based on the plosive intensity. In some possible implementations, the target reduction gain for each plosive event can be defined as:

[0183]

[0184] Among them, st mean It is the average intensity of the input signal. That is, the goal is usually to reduce the intensity of the plosive sound to an average level without over-suppressing it.

[0185] When the LFE ratio is used to represent intensity, this ratio can be directly expressed as the target gain. Although in some cases the target gain is expressed in dB (negative for reductions), the attenuation gain slope can be defined as:

[0186] slope = -targetGaindB ×β (13)

[0187] This maps the target gain to a slope (dB per octave as a positive value), and β is a scaling factor used to control the aggressiveness. For values ​​below x C Each frequency band x (at freq) C The attenuation gain (in dB) of the segment can then be calculated as:

[0188] gain dB [x] = (log2 x - log2(0.5*x)) C ))×slope-slope (14)

[0189] In some possible implementations, noise spectrum estimation can be used to limit gain reduction so that attenuation does not affect the overall spectral distribution curve in the low-frequency region.

[0190] Therefore, broadly speaking, the proposed method generally attenuates faster pops with higher cutoff frequencies, thus effectively adapting to the pitch of the speaker's speech. The method also attenuates stronger pops with steeper cutoff frequency slopes, thus effectively adapting to weak but intense plosives.

[0191] It is worth noting that when the techniques described above are applied to content including music or a combination of speech and music, the algorithm may be sensitive to low-frequency transients (such as those generated by kicking drums or bass). To address this issue, in some possible implementations, a content classifier (e.g., a speech / music activity detector) that computes the probability p(n) that a given frame n contains speech (or does not contain speech) can be used to modify detection or attenuation parameters, thereby ensuring that the musical content is not affected by de-plosive processing. In some possible implementations, p(n) > th p (where, th) p Frames with a predefined threshold can be removed from the LFE and ZCM pools to ensure relevant plosive detection and attenuation. p(n) can also be used to dynamically modify the amount parameter, for example, by multiplying it by the logical mapping function f(p(n)), where, for example, f(x) = 1 / (1 + κ*e -(x-0.5) ) is a continuous function that approaches 0 and 1 respectively when x approaches 0 and 1 respectively. κ usually represents the steepness parameter of the mapping.

[0192] In some implementations, particularly when computational complexity allows, another embodiment of frequency / spectral domain attenuation may be employed and will now be described in more detail.

[0193] Specifically, it can be proposed to first use an analytical filter bank to generate (approximately) equivalent rectangular band (ERB) spaced frequency bands below a (predefined) frequency threshold (e.g., approximately 500 Hz) in the plosive tone frequency region, and additionally generate one or more frequency bands above this frequency threshold (e.g., 500 Hz) to cover the remaining frequency range. At each time point t, the energy in each of these frequency bands b (denoted as e(b,t)) is used to control the reduction process to generate a series of gains g(b,t) applied to each filtered signal. The results are then fed into a synthesis filter bank to produce an output signal with reduced plosive tone energy.

[0194] More specifically, in some possible implementations, the plosive reduction gain in each frequency band g(b,t) can first be calculated based on the energy of the frequency band from the output of the compression curve as follows:

[0195] g1(b,t)=C(e dB (b,t)) (15)

[0196] in,

[0197] e dB (b,t)=1010g 10 (e(b,t)) (16)

[0198] In some possible implementations, a compression curve having a threshold T, knee width W, and compression ratio R (where all quantities are expressed in decibels) can be described as follows:

[0199]

[0200] As those skilled in the art will understand and recognize, any suitable values ​​for the threshold T, knee width W, and compression ratio R can be used. In an illustrative example, T = -65, W = 10, and R = 6 can be used. The compression curve will then be 0 dB at low energy and may only show a decay as energy increases. It should also be understood that T can be dynamically adapted to the time-smoothed speech energy envelope over time.

[0201] In some possible implementations, the gain can then be further clipped to ensure that the power in each frequency band does not drop to the estimated noise floor in the band, according to the following formula (expressed as...). Or expressed in dB )the following:

[0202]

[0203] This will generally prevent the noise drop that can be heard when there is a popping sound in the presence of significant background noise.

[0204] One possible way to estimate noise is:

[0205]

[0206] Here, a negative value for t1 implies the use of estimated history, and a positive value for t2 may, in some cases, require some time delay compensation for causality and can therefore be set to 0. In some possible implementations, a good estimate can be given by -t1 = t2 = 300 ms. In some possible implementations, removing values ​​below 80 dB from the minimum calculation can also be useful, as they will generally not represent background noise during speech (and are more likely to be generated by noise gates).

[0207] In some cases, when the lowest frequency is around, for example, 80 Hz, a difficult situation may be distinguishing the unwanted low-frequency energy of a plosive event from the expected low-frequency energy in a vowel sound. Depending on the implementation, several tools can typically be used to resolve these conditions. More specifically, in some possible implementations, time-smoothed low-frequency energy estimates of the signal above the noise floor (which seeks to maintain compression gain) and tonality measures (or, in some possible implementations, some kind of periodicity measure) that detect the repetition kurtosis of vowels and reduce the gain can be used. These can be implemented as follows:

[0208]

[0209] Here, b = B corresponds to a frequency band centered, for example, 200 Hz. This estimate can then be time-smoothed using an exponential smoother with, for example, an attack time of 50 ms and a release time of 100 ms, thus giving a smoothed LFE. S Finally, subtracting the estimated background noise will give:

[0210]

[0211] This can then be further thresholded and scaled to a useful range, for example, according to the following formula, to produce a factor:

[0212] f uf =(min(max(LFE) n ,30),40)-30) / 10 (22)

[0213] In some possible implementations, the tonality (or, in some cases, a periodicity measure) can be estimated before conversion to the filter bank domain. In some possible implementations, the filter bank can compute the FFT values ​​of overlapping windowed audio signals. For ease of illustration, in some possible implementations, it can be assumed that the power p(k) in the FFT segment is available and segments k = 0 up to k = K will be used, where, for example, K corresponds to 500 Hz at a given sampling rate.

[0214] Periodic metrics (e.g., cepstrum in some possible implementations) can then be computed on these segments, as follows:

[0215]

[0216] in, This can be a forward or inverse Fourier transform. This can be considered a form of autocorrelation. Generally speaking, vowels can be expected to have a periodicity of approximately 100 Hz or less. Therefore, in some possible implementations, C can be considered. p The first 100Hz and find the minimum C p (min) and the maximum value C that appears after this minimum value in the first 100Hz. p (max). In some possible implementations, this vowel is then amplitude-limited and scaled to a tonality measure:

[0217] tonality=(min(max(C p (max)-C p (min), 0), 6)) / 6 (24)

[0218] In some possible implementations, the tonality metric may alternatively be calculated by searching for the largest spectral peak in p(k) within, for example, a frequency range of 60 Hz to 250 Hz, requiring that the peak be a reasonably sinusoidal peak (the main lobe should be sufficiently narrow and deep). For example, the tonality metric may be scaled (e.g., linearly) from 0 to 1, since the depth at the peak center plus or minus the range of 60 Hz is from 5 dB to 15 dB.

[0219] For example, this value can be time-smoothed with an attack time of 75ms and a release time of 300ms, thus giving a smoothed tone.

[0220] This (smoothed) tonality measure and the f calculated above lf This can be further combined into a gain scaling factor:

[0221] g3 = g2 × f lf +(1-f lf )×(1-tonalitys ) 2 (25)

[0222] It should be noted that the periodicity / tonality measure described above can also be referred to as a "speech harmonic protection measure" in the context of this disclosure. Furthermore, periodicity and tonality measures can be used interchangeably.

[0223] The gain can then be further constrained so that for frequency bands above a certain (predetermined) threshold (e.g., 70Hz), the gain attenuation cannot exceed that of the immediately adjacent lower frequency band, according to the following formula:

[0224] g4(b,t)=max(g3(b,t),g3(b-1,t)) (26)

[0225] Among them, b has a frequency band center frequency higher than, for example, 70 Hz.

[0226] In a broad sense, the methods described above typically enforce reduction to follow the physical decrease in plosive energy with increasing frequency. Specifically, when the energy of a lower frequency band decreases significantly, if the next higher frequency band has more energy, it is more likely to be genuine speech energy rather than plosive-related energy. Generally, in some possible implementations, very low frequency bands (below, for example, 70 Hz) may not follow this trend; for example, excessive 60 Hz power hum can make a band louder, or a DC blocking filter can attenuate the lowest frequency band, and this should not limit the attenuation of plosive energy.

[0227] Finally, in some possible implementations, these gains g4(b,t) can be further time-smoothed with, for example, an attack time of 20 ms and a release time of 50 ms to produce a final gain g(b,t) that will be applied to the filtered signal (e.g., a sub-band signal). In some implementations, for example, the final gain can be applied in a band-by-band manner.

[0228] Figure 8 This is a schematic illustration of a simplified diagram showing an example of attenuation of speech plosives according to an embodiment of the present disclosure. Specifically, as can be seen from... Figure 8 As seen in the corresponding attenuation diagram 8100, speech plosive events (the white area in the low-frequency part of the comparison diagram 8200) have been effectively attenuated.

[0229] Figure 11 This is a schematic flowchart illustrating an example of a method 11000 for performing automatic audio enhancement on an input audio signal including at least one speech pronunciation noise event according to an embodiment of the present disclosure.

[0230] In particular, the method 11000 described herein can be applied to perform automatic audio enhancement (e.g., detection, attenuation, etc.) for speech plosive noise events or mouth click noise events.

[0231] More specifically, method 11000 can begin at step S11010 by segmenting the input audio signal (e.g., by using one or more suitable windows) into multiple audio frames (e.g., 100ms in size). Method 11000 can then continue at step S11020 by obtaining (e.g., determining, calculating, extracting, etc.) at least one feature parameter from the (segmented) audio frames. In some possible example implementations, the feature parameter thus obtained can be considered as being associated with the type of speech articulation noise event (to be detected). That is, in some possible example implementations, depending on the type of speech articulation noise event (to be detected), different feature parameters will have to be obtained from the audio frames. Finally, method 11000 can continue at step S11030 by determining (e.g., detecting, calculating, etc.) the corresponding type of speech articulation noise event within the input audio signal and the corresponding range (e.g., time and / or frequency range) associated with the speech articulation noise event, at least in part, based on the obtained feature parameters.

[0232] As described above, broadly speaking, the proposed method 11000 provides an efficient and flexible mechanism for identifying (detecting) multiple possible / potential speech articulation noise events (e.g., artifacts) included within the input audio signal. This facilitates appropriate further enhancement (post-processing) (e.g., attenuation). Consequently, the manual editing / processing previously required for identifying and attenuating multiple noise events in the audio signal can be largely avoided. Simultaneously, the listening experience can be significantly improved.

[0233] Figure 12 This is a schematic flowchart illustrating an example of a method 12000 for performing automatic audio enhancement on an input audio signal to detect and / or attenuate at least one speech articulation noise event contained therein, according to another embodiment of this disclosure. The speech articulation noise event may, in particular, include at least one speech plosive event. Therefore, it is conceivable that the method 12000 described herein may be particularly suitable for performing automatic audio enhancement (e.g., detection, attenuation, etc.) for speech plosive noise events.

[0234] Specifically, method 12000 can begin at step S12010 by generating multiple approximately equivalent rectangular bandwidth (ERB) spaced frequency bands below a predefined frequency threshold and multiple frequency bands above the predefined frequency threshold, which is within the frequency range of speech plosive events. Method 12000 can then continue at step S12020 by applying multiple attenuation gains to the audio signal in each of the frequency bands, wherein these attenuation gains are calculated based on energy calculated for each frequency band. Finally, method 12000 can further continue at step S12030 by feeding attenuated audio samples to a synthesis filter bank to generate an output audio signal.

[0235] As described above, broadly speaking, the proposed method 12000 provides an efficient and flexible mechanism for identifying (detecting) and attenuating (multiple) possible / potential speech articulation noise events (e.g., speech plosive events) included within the input audio signal. Therefore, the manual editing / processing previously required for identifying and attenuating (multiple) noise (e.g., plosive) events in the audio signal can be largely avoided. Simultaneously, the listening experience can be significantly improved.

[0236] Incidentally, it should be noted that although the methods / techniques for de-clicking and de-popping sound processing appear to be illustrated separately, those skilled in the art will understand and recognize that at least some of the techniques described above can be used interchangeably.

[0237] As an illustrative and non-limiting example, in some possible implementations, the filter bank approach (described above in the context of de-plosive processing) can also be applied to de-clicking, where the spectral envelope can be defined by the ERB band energy and a similar multi-band compression scheme (a compressor ratio determined by the target attenuation gain and the corresponding attack / release time) can be applied. It can be noted that the effective ERB bands may extend up to the Nyquist limit of de-clicking techniques, but for plosive processes, they are limited to low frequencies (e.g., 500 Hz). Furthermore, the “residuals” also used for de-plosive processing (described above as being used only for de-clicking) can be utilized as an alternative to cepstral-based periodicity metrics. It can be noted that the residuals used for de-plosive processing cannot use second-order sample differences, but must use some other suitable estimation.

[0238] Figure 13 An example is illustrated where the goal is to combine techniques for de-clicking and de-popping in a (single) feature overview.

[0239] In particular, it is important to note that Figure 13Function blocks 13010, 13020, and 13030 in the code are generally similar to or analogous to each other. Figure 6 Function blocks 6010, 6020, and 6030 in the code allow for the omission of their repeated descriptions for brevity. It is important to further note that... Figure 13 The dashed boxes shown often indicate that the corresponding functional steps may be optional, as will be described in more detail below.

[0240] As described above, for de-plosive processing, ERB band analysis (dashed box 13050) can be applied to detect the corresponding speech artifact (in this case, a speech plosive event) (as illustrated in box 13060) and then attenuate this speech artifact (box 13070). On the other hand, for de-clicking scenarios, an ERB correlation procedure (or in some cases, a filter bank method) can be performed after a speech artifact (in this case, a mouth click event) has been detected (box 13060). In such a case, such an ERB correlation procedure can also be referred to as ERB band synthesis for attenuating the detected mouth click (box 13070) (as illustrated in dashed box 13080). As explained above, when the filter bank method (which is described above in the context of de-plosive processing) is to be applied to de-clicking, the spectral envelope can be defined by the ERB band energy and a multi-band compression scheme (a compressor ratio determined by the target attenuation gain and the corresponding attack / release time or envelope interpolation) can be applied. As will be understood and recognized by those skilled in the art, any other or further suitable process may be employed depending on the various implementation methods and / or requirements.

[0241] Furthermore, as described above and also in Figure 13 As shown, the techniques described herein can be further (optionally) utilized as “residuals” for both de-clicking and de-plosive processing (e.g., by removing speech harmonic components, as illustrated in dashed box 13040) (wherein the residual is used as an alternative to a periodicity / tonality metric). However, it should be noted that in such cases (i.e., using residuals), the harmonics may ultimately have to be restored or added back, for example, after envelope attenuation has been applied to the residual signal (as illustrated in dashed / optional box 13090).

[0242] This disclosure also relates to an apparatus for performing the methods and techniques described throughout the disclosure. Figure 14An example of such a device 14000 is shown. The device 14000 includes a processor 14010 and a memory 14020 coupled to the processor 14010. The memory 14020 may store instructions for the processor 14010. The processor 14010 may receive audio data 14030 as input. The audio data 14030 may have the properties described above in the context of a corresponding method for performing automatic audio enhancement on an input audio signal to detect and / or attenuate at least one speech pronunciation noise event contained therein. The processor 14010 may be adapted to perform the methods / techniques described throughout the disclosure. Accordingly, the processor 14010 may output denoised (e.g., de-clicking, de-plosive) audio data 14040. In some further possible embodiments, the processor 14010 may also be made capable of receiving further inputs (e.g., control parameters, etc.). Figure 14 (not shown in the image), for example, to control audio enhancement processing behavior.

[0243] explain

[0244] Computing devices implementing the techniques described above can have the following example architectures. Other architectures are also possible, including those with more or fewer components. In some implementations, the example architectures include one or more processors (e.g., dual-core). The components include a processor, one or more output devices (e.g., an LCD), one or more network interfaces, one or more input devices (e.g., a mouse, keyboard, touchscreen display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can exchange communication and data via one or more communication channels (e.g., a bus), which can utilize various hardware and software to facilitate the transfer of data and control signals between the components.

[0245] The term "computer-readable medium" refers to a medium that participates in providing instructions to a processor for execution, including but not limited to non-volatile media (e.g., optical discs or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media include, but are not limited to, coaxial cables, copper wires, and optical fibers.

[0246] Computer-readable media may further include an operating system (e.g., The system comprises an operating system, a network communication module, an audio interface manager, an audio processing manager, and a real-time content distributor. The operating system can be multi-user, multi-processor, multi-tasking, multi-threaded, real-time, etc. The operating system performs basic tasks, including but not limited to: recognizing input from network interfaces and / or devices and providing output to network interfaces and / or devices; recording and managing files and directories on computer-readable media (e.g., memory or storage devices); controlling peripheral devices; and managing traffic on one or more communication channels. The network communication module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP and HTTP).

[0247] The architecture can be implemented in parallel processing or peer-to-peer infrastructure, or on a single device with one or more processors. The software may include multiple software components or may be a single code body.

[0248] The described features can be advantageously implemented in one or more computer programs that can execute on a programmable system including at least one programmable processor coupled to receive and transfer data and instructions from a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform an activity or produce a result. Computer programs can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0249] For example, suitable processors for executing instruction programs include both general-purpose processors and special-purpose processors, as well as a single processor or one or more processors or cores in any type of computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data files, or operatively coupled to and communicating with them; such devices include disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly representing computer program instructions and data include all forms of non-volatile memory, such as semiconductor memory devices like EPROM, EEPROM, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by or incorporated into an ASIC (Application-Specific Integrated Circuit).

[0250] To provide user interaction, features can be implemented on a computer having a display device for showing information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or retina display device. The computer may have touch surface input devices (e.g., touchscreens) or a keyboard and pointing devices such as a mouse or trackball, through which the user can provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0251] The features can be implemented in a computer system that includes back-end components, such as data servers, or middleware components, such as application servers or internet servers, or front-end components, such as client computers with graphical user interfaces or internet browsers, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include, for example, LANs, WANs, and computers and networks that form the internet.

[0252] A computing system may include clients and servers. Clients and servers are typically located far apart and interact via a communication network. The client-server relationship arises from computer programs running on respective computers and having a client-server relationship relative to each other. In some embodiments, the server transmits data (e.g., HTML pages) to the client device (e.g., to display data to a user interacting with the client device and to receive user input from the user). Data generated at the client device (e.g., the result of user interaction) can be received from the client device at the server.

[0253] A system of one or more computers may be configured to perform specific actions by means of software, firmware, hardware, or a combination thereof installed on the system that will cause the system to perform actions in operation. One or more computer programs may be configured to perform specific actions by means of instructions that, when executed by a data processing device, cause the device to perform those actions.

[0254] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or claim, but rather as descriptions of features characteristic of particular embodiments of a particular invention. Specific features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in a particular combination and even initially claimed, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0255] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or to perform all illustrated operations to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0256] Unless otherwise specifically stated, it is obvious from the following discussion that, throughout this public discussion, terms such as “processing,” “computing,” “calculating,” “determining,” and “analyzing” are used to refer to the actions and / or processes by which data represented as physical (e.g., electronic) quantities are manipulated and / or transformed into other data similarly represented as physical quantities by a computer or computing system or similar electronic computing device.

[0257] Throughout this disclosure, references to “one example embodiment,” “some example embodiments,” or “example embodiment” mean that a particular feature, structure, or characteristic described in connection with an example embodiment is included in at least one example embodiment of this disclosure. Therefore, the phrases “in one example embodiment,” “in some example embodiments,” or “in an example embodiment” appearing throughout this disclosure do not necessarily refer to the same example embodiment. Furthermore, in one or more example embodiments, particular features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those skilled in the art based on this disclosure.

[0258] As used herein, unless otherwise specified, ordinal adjectives such as “first,” “second,” “third,” etc., are used to describe common objects only to indicate different instances of similar objects and are not intended to imply that the objects described must be in a given order in time, space, hierarchy, or any other way.

[0259] Similarly, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered restrictive. The use of “including,” “comprising,” or “having,” and variations thereof, is intended to cover the items listed thereafter and their equivalents, as well as additional items. Unless otherwise specified or limited, the terms “installation,” “connection,” “support,” and “coupled,” and variations thereof, are used extensively and cover direct and indirect installation, connection, support, and coupling.

[0260] In the claims below and in the description herein, the terms *comprising*, *comprised of*, or *which comprises* are open-ended terms meaning that at least the following element / feature is included, but not excluding other elements / features. Therefore, when the term *comprising* is used in a claim, it should not be construed as limited to the means, elements, or steps listed thereafter. For example, the expression of a device including A and B should not be limited to a device that includes only elements A and B. As used herein, the terms *including*, *which includes*, or *that includes* are also open-ended terms meaning that at least the element / feature following the term is included, but not excluding other elements / features. Therefore, *including* is synonymous with *comprising* and means *comprising*.

[0261] It should be recognized that in the foregoing description of exemplary embodiments of this disclosure, various features of this disclosure are sometimes combined in a single exemplary embodiment / figure or its description in order to simplify the disclosure and aid in understanding one or more of the inventive aspects. However, the approach of this disclosure should not be construed as reflecting an intention in the claims to require more features than expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in fewer than all features of a single foregoingly disclosed exemplary embodiment. Therefore, the claims following this specification are hereby expressly incorporated, wherein each claim is an independent, separate exemplary embodiment of this disclosure.

[0262] Furthermore, while some of the exemplary embodiments described herein include features that are included in other exemplary embodiments but not others, as those skilled in the art will understand, combinations of features from different exemplary embodiments are intended to be within the scope of this disclosure and to form different exemplary embodiments. For example, any exemplary embodiment of the claimed embodiments in the appended claims can be used in any combination.

[0263] Numerous specific details are set forth in the description provided herein. However, it should be understood that exemplary embodiments of this disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail to avoid obscuring the understanding of this specification.

[0264] Therefore, although the mode considered to be the best mode of this disclosure has been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of this disclosure, and all such changes and modifications falling within the scope of this disclosure are intended to be claimed. For example, any formulas given above merely represent processes that can be used. Functions can be added or removed from the block diagram, and operations can be interchanged between functional blocks. Steps can be added or removed from the methods described within the scope of this disclosure.

[0265] Various aspects and implementations of this disclosure can also be understood from the following enumerated exemplary embodiments (EEEs), which are not claims.

[0266] EEE 1. A method for detecting and attenuating mouth clicks in speech recordings based on the following:

[0267] a. Divide the audio into speech frames and non-speech frames;

[0268] b. Calculate the second-order waveform difference of the speech frames;

[0269] c. Detect the clicking sound of the mouth based on the kurtosis of each short-time waveform;

[0270] d. Calculate the target spectral gain based on interpolation of the spectral envelope between the start and end of the click sound; and

[0271] e. Apply gain to each frame and perform overlap-addition and recombination.

[0272] EEE 2. The method as described in EEE 1, wherein the recognition of speech and non-speech frames is provided by an existing VAD (Voice Activity Detector).

[0273] EEE 3. The method described in EEE 1, wherein optional denoising may be applied to the input signal to better reveal potential mouth clicks.

[0274] EEE 4. The method as described in EEE 1a, wherein two window sizes are used separately to detect speech clicks (short) and non-speech clicks (long).

[0275] EEE5. As described in EEE 1c, wherein the kurtosis k of the original waveform is calculated. W The kurtosis of the difference between the second-order waveform and the second-order waveform is used for mouth click detection.

[0276] EEE 6. The method as described in EEE 1c, wherein a mouth click sound is detected by means of predefined kurtosis thresholds. These thresholds are relative to k. W and k D They can be different.

[0277] EEE 7. As described in EEE 5, wherein, based on k W To detect nonverbal clicking sounds and based on k with weighted parameter α D -α*k W To detect speech clicks.

[0278] EEE 8. As described in EEE 7, speech transients can be further excluded from kurtosis-based detection. This only requires attention to speech clicks.

[0279] EEE 9. The method described in EEE 8, wherein speech transients can be detected based on the centroid (average time) of a short-time signal.

[0280] EEE 10. As described in EEE 7, wherein the start of a mouth click is defined as when the kurtosis rises above a threshold and the end of a mouth click is defined as when the kurtosis falls below a threshold. Therefore, a mouth click event typically covers several consecutive short frames.

[0281] EEE 11. The method as described in EEE 7, wherein the duration of the nonverbal clicks tends to be longer, and therefore merging adjacent nonverbal clicks is preferred.

[0282] EEE 12. The method as described in EEE 7, wherein a nonverbal click sound that occurs just before the start of speech is considered a smacking candidate.

[0283] EEE 13. The method as described in EEE 12, wherein the end position of the smacking event is extended based on the following features: spectral slope, high / low peak ratio, and energy envelope.

[0284] EEE 14. The method as described in EEE 13, wherein the high / low peak ratio is defined as the amplitude ratio between the maximum peak value in the high frequency band and the maximum peak value in the low frequency band.

[0285] EEE 15. The method as described in EEE 14, wherein the high / low frequency bands are separated by a predefined frequency (e.g., 1.5 kHz).

[0286] EEE 16. The method as described in EEE 7, wherein the speech clicks tend to be shorter, and therefore refining the start / end sample positions is preferred.

[0287] EEE 17. The method described in EEE 16, wherein the simple refinement method is to locate the maximum second-order waveform difference (maxD) within the range of the initial click detected by kurtosis. For example, a predefined speech click duration of 2 ms can then be used to determine the refinement start / end position around maxD.

[0288] EEE 18. As described in EEE 16, wherein the alternative refinement method is the "minimum / maximum rate of change". The zero-crossing rate (cZCR) of the transformed waveform is -1 / +1 at local minimum / maximum values ​​and 0 elsewhere. Frames with cZCR above a threshold define the refinement location.

[0289] EEE 19. The method as described in EEE 1d, wherein spectral gain attenuation is calculated based on the observed spectral envelope and the target spectral envelope. The spectral gain inherited from the spectral envelope defines the frequency-dependent gain value at each spectral segment.

[0290] EEE 20. The method as described in EEE 19, wherein the target spectral envelope can be estimated by linear interpolation of the spectral envelope between two “clean” frames (containing no click events) at the end of each click event.

[0291] EEE 21. The method as described in EEE 19, wherein the spectral gain at each short time frame is defined as the target envelope divided by the observed envelope.

[0292] EEE 22. The method as described in EEE 19, wherein the spectral gain is limited to attenuation only. The resulting amplification gain is forced to be 1.

[0293] EEE 23. The method as described in EEE 19, wherein the spectral gain of the speech frame is adapted to a spectral region above a predefined audible frequency (e.g., 4 kHz).

[0294] EEE 24. A method for detecting and attenuating unwanted plosive events in speech content recordings based on the following:

[0295] a. Divide the audio into overlapping frames;

[0296] b. Analyze the low-frequency energy (LFE) and zero-crossing maximum (ZCM) of each frame;

[0297] c. Detecting plosive events with precise start / end time positions; and

[0298] d. Attenuate plosive sounds by using a high-pass filter with adaptive order and cutoff frequency.

[0299] EEE 25. The method as described in EEE 24b, wherein the LFE can be calculated in the time domain or in the spectral domain having a predefined cutoff frequency.

[0300] EEE 26. The method as described in EEE 25, wherein the time-domain LFE can be calculated as the RMS energy of a low-pass filtered version of the input signal.

[0301] EEE 27. The method as described in EEE 25, wherein the spectral domain LFE can be calculated as the RMS energy of a short-time spectrum below the cutoff frequency.

[0302] EEE 28. The method described in EEE 24b, wherein ZCM is the maximum interval of consecutive zero crossings normalized by the window size.

[0303] EEE 29. The method as described in EEE 24a, wherein the frame size is set large enough to extract reliable values ​​of zero-crossing maximum values. The overlap size is set large enough to track short-term features with fine temporal resolution.

[0304] EEE 30. The method described in EEE 24c, wherein the plosive detection is based on outliers selected from the LFE distribution across all short-time frames of a file.

[0305] EEE 31. The method as described in EEE 30, wherein outliers are detected by standard scores and an adaptive threshold is used to select dominant outliers.

[0306] EEE 32. The method as described in EEE 31, wherein the threshold is adapted to the difference between the maximum LFE and the standard score threshold multiplied by a scaling factor.

[0307] EEE 33. The method described in EEE 32, wherein the scaling factor can be obtained from the global plosive removal amount control [0,1].

[0308] EEE 34. The method as described in EEE 24c, wherein the plosive detection for low-latency use cases is based on the LFE ratio between two adjacent frames.

[0309] EEE 35. As described in EEE 34, a predefined threshold is used for detection, which can be defined as 1 plus the detection sensitivity.

[0310] EEE 36. The method as described in EEE 32 or EEE 34, wherein consecutive frames exceeding a threshold define the time span of the plosive event.

[0311] EEE 37. The method as described in EEE 24c, wherein the initial plosive event boundary is defined by the method of EEE 36.

[0312] EEE 38. The method described in EEE 36, wherein the initial event boundaries are further refined based on ZCM.

[0313] EEE 39. The method as described in EEE 38, wherein the start position and end position are extended until the ZCM drops below a predefined threshold.

[0314] EEE 40. The method as described in EEE 24d, wherein the attenuation process can be performed in the time domain or the spectral domain.

[0315] EEE 41. The method as described in EEE 40, wherein the filter frequency is adaptive to ZCM under frequency constraints within a predefined range.

[0316] EEE 42. The method as described in EEE 40, wherein the time-domain attenuation uses a Butterworth filter whose filter order is adaptive to the intensity of the low-frequency energy under a predefined range of order constraints.

[0317] EEE 43. The method as described in EEE 42, wherein the filtered output is cross-gradually alternating with the original input signal at event boundaries with a predefined transition duration.

[0318] EEE 44. The method as described in EEE 40, wherein the spectral domain attenuation uses a standard STFT overlapping additive frame.

[0319] EEE 45. The method as described in EEE 44, wherein the spectral attenuation gain slope is adapted to the intensity of the low-frequency energy.

[0320] EEE 46. The method as described in EEE 45, wherein the gain slope is expressed in dB per octave below the cutoff frequency.

[0321] EEE 47. The method as described in EEE 44, wherein the attenuation gain can be limited by the estimated noise spectrum to prevent oversuppression.

[0322] EEE 48. The method as described in EEE 32, wherein the scaling factor may incorporate the probability of speech obtained from the content classifier. The resulting factor is then used to weight the detection threshold accordingly to avoid processing non-speech frames.

[0323] EEE 49. A method for detecting and attenuating mouth clicking sounds in audio data, comprising:

[0324] Receive multiple audio frames representing audio data;

[0325] Calculate one or more short-time waveforms based on multiple audio frames;

[0326] Detect one or more mouth click sounds based on the kurtosis of one or more short-time waveforms;

[0327] A set of target spectral gains is calculated, at least in part, based on the interpolation of the spectral envelope between the start and end of one or more detected mouth click sounds; and

[0328] One or more mouth clicks are attenuated by applying the target spectral gain to multiple audio frames and performing overlapping, additive, and resynthesizing.

[0329] EEE 50. The method as described in EEE 49 further includes:

[0330] Each of the multiple audio frames is classified as either a speech frame or a non-speech frame; and where:

[0331] Calculating one or more short-time waveforms based on multiple audio frames includes:

[0332] Calculate the original waveform obtained from the audio content; and

[0333] Calculate the second-order waveform difference of speech frames;

[0334] Detecting one or more mouth clicking sounds includes:

[0335] Use the raw waveform obtained from the audio content to detect one or more mouth clicks in non-verbal frames; and

[0336] Use the second-order waveform difference of speech frames to detect one or more mouth clicks in the speech frame.

[0337] EEE 51. The method as described in EEE 49 or 50 further includes: denoising the audio frame before calculating one or more short-time waveforms.

[0338] EEE 52. The method of any one of EEE 50 to 51, wherein classifying each of a plurality of audio frames into a speech frame or a non-speech frame is performed by an existing speech activity detector.

[0339] EEE 53. The method as described in any one of EEE 49 to 52, wherein, according to a first predefined kurtosis threshold (K) T1 This is used to detect one or more mouth clicks in a speech frame.

[0340] EEE 54. The method as described in EEE 53, wherein, according to a second predefined kurtosis threshold (K) different from the first predefined kurtosis threshold... T2 This is used to detect one or more mouth clicks in a speech frame.

[0341] EEE 55. The method of any of EEE 49 to 54, wherein speech transients are detected and speech transients are excluded from kurtosis-based mouth click detection.

[0342] EEE 56. The method as described in EEE 55, wherein speech transients are detected at least in part based on the centroid (average time) of the original waveform obtained from the audio content (e.g., a short-time signal based on the audio content).

[0343] EEE 57. The method as described in any one of EEE 49 to 56, wherein the onset of the corresponding mouth click is defined as when the kurtosis rises to K. T The above time, and the corresponding mouth click sound, are defined as the end when the kurtosis drops to K. T The following times.

[0344] EEE 58. The method of any one of EEE 50 to 57, wherein detecting one or more mouth clicks of a nonverbal frame includes merging nonverbal clicks separated by a first duration.

[0345] EEE 59. The method of any one of EEE 50 to 58, wherein detecting one or more mouth clicks of a speech frame further comprises: refining the start and end positions of each corresponding mouth click in the one or more mouth clicks of the speech frame.

[0346] EEE 60. The method as described in EEE 59, wherein the refinement start position and end position include:

[0347] Locate the maximum second-order waveform difference (MD) within the approximate clicking range of the corresponding mouth click sound detected by kurtosis; and

[0348] The fine-grained start or stop position of the corresponding mouth click is defined based on the predefined duration of the speech click.

[0349] EEE 61. The method as described in EEE 59, wherein the refinement start position and end position include:

[0350] The fine-grained start or stop position of the corresponding mouth click is defined based on the zero-crossing rate (cZCR) of the transformed waveform. (For example, the transformed waveform maps the local minimum / maximum values ​​of the observed waveform to -1 / 1 and all other values ​​to 0.)

[0351] EEE 62. The method as described in EEE 48, wherein the set of target spectral gains is calculated based at least in part on the observed spectral envelope and the target spectral envelope.

[0352] EEE 63. The method as described in EEE 62, wherein the target spectral envelope is estimated by linear interpolation of the spectral envelope between two “clean” frames (e.g., surrounding frames that do not contain any click events) at the end of each click event.

[0353] EEE 64. The method as described in EEE 62, wherein the target spectral gain at each short time frame is defined as the target envelope divided by the observed envelope.

[0354] EEE 65. The method as described in EEE 64, wherein the target spectral gain is limited to attenuation only. (For example, the resulting amplification gain is forced to 1.)

[0355] EEE 66. The method as described in EEE 64, wherein the set of target spectral gains of the speech frame is applicable to a spectral region above a predefined audible frequency.

[0356] EEE 67. A method for detecting and attenuating unwanted plosive events in audio including speech content based on:

[0357] Divide the audio into multiple overlapping frames;

[0358] Determine the low-frequency energy of each of multiple overlapping frames;

[0359] Determine the maximum zero-crossing value of at least one of multiple overlapping frames;

[0360] Detect multiple plosive events with precise start / end time positions;

[0361] The output audio is generated by attenuating multiple plosive events using an adaptive high-pass filter, wherein the order and cutoff frequency of the adaptive high-pass filter are adapted to each of the multiple plosive events.

[0362] EEE 68. The method described in EEE 67, wherein the low-frequency energy is the RMS energy of a low-pass filtered version of the input signal.

[0363] EEE 69. The method is similar to that of EEE 67, where the maximum zero-crossing value is the maximum interval between consecutive zero-crossing points normalized by the window size.

[0364] EEE 70. The method as described in EEE 67, wherein the frame size is set large enough to extract reliable values ​​of zero-crossing maximum values. The overlap size is set large enough to track short-time features with fine temporal resolution.

[0365] EEE 71. The method as described in EEE 67, wherein detecting multiple plosive events includes detecting outliers in the low-frequency energy distribution across all short-time frames of a file based on a first threshold.

[0366] EEE 72. The method of any one of EEE 67 to 71, wherein detecting a plurality of plosive events further comprises:

[0367] The threshold for LFE outlier detection is calculated based on the standard score; and

[0368] A second threshold, different from the first threshold, is applied (e.g., an adaptive threshold for selecting the dominant component).

[0369] EEE 73. The method as described in EEE 72, wherein the second threshold is adapted to the difference between the maximum outlier and the first threshold.

[0370] EEE 74. The method as described in EEE 73, wherein consecutive frames exceeding an adaptive threshold define the time span of a plosive event.

[0371] EEE 75. The method as described in EEE 73, wherein the global attenuation effect [0,1] is mapped to an adaptive threshold scaled by a certain factor.

[0372] EEE 76. The method as described in EEE 67, wherein the initial plosive event boundary (e.g., start / stop position) is defined by the method of EEE 73.

[0373] EEE 77. The method as described in any one of EEE 67 to 74 further includes: refining the location of the plosive event (e.g., the initial boundary) based on the maximum zero-crossing value.

[0374] EEE 78. The method as described in EEE 77 further includes: extending the start and end positions of the plosive event until the maximum zero-crossing value drops below a predefined threshold.

[0375] EEE 79. The method as described in EEE 67, wherein generating output audio includes cross-gradienting at the boundaries of multiple plosive events with a predefined transition duration.

[0376] EEE 80. The method described in EEE 67, wherein the filter order adapts to the intensity of low-frequency energy within a predefined end range.

[0377] EEE 81. The method as described in EEE 67, wherein the cutoff frequency adapts to the maximum zero-crossing value within a predefined cutoff frequency range.

[0378] EEE 82. The method as described in EEE 75, further comprising:

[0379] Obtain the speech probabilities of one or more of multiple overlapping frames from the content classifier; and

[0380] Reduce the number of detections when the corresponding probability is less than the first classification threshold (e.g., by changing the global decay effect).

[0381] EEE 83. The method as described in EEE 75 further includes:

[0382] Obtain the speech probabilities of one or more of multiple overlapping frames from the content classifier; and

[0383] Frames from detected plosive events are removed when the corresponding probability is less than the second classification threshold.

[0384] EEE 84. The method as described in EEE 67, wherein using an adaptive high-pass filter to attenuate multiple plosive events includes:

[0385] The first plosive event in a plurality of plosive events is filtered using a first filter order and a first cutoff frequency; and

[0386] The second plosive event among multiple plosive events is filtered using a second filter order and a second cutoff frequency, wherein at least one of the second filter order and the second cutoff frequency is different from the first filter order and the first cutoff frequency, respectively.

[0387] EEE 85. The method described in EEE 67, wherein the adaptive high-pass filter is a Butterworth filter.

[0388] EEE 86. A non-transitory computer-readable storage medium storing one or more programs comprising instructions that, when executed by one or more processors, perform the method as described in any one of EEE 67 to 85.

[0389] EEE 87. An electronic device comprising one or more processors and a memory storing one or more programs including instructions that, when executed by the one or more processors, cause the device to perform a method as described in any one of EEE 67 to 85.

[0390] EEE 88. A method for performing automatic audio enhancement on an input audio signal including at least one speech pronunciation noise event, the method comprising:

[0391] The input audio signal is divided into multiple audio frames;

[0392] Obtain at least one feature parameter from the audio frame; and

[0393] The type of speech articulation noise event within the input audio signal and the corresponding time-frequency range associated with the speech articulation noise event are determined, at least in part, based on the obtained characteristic parameters.

[0394] EEE 89. The method according to EEE 88, wherein the determined range includes at least one boundary of the determined speech pronunciation noise event in the time domain and / or spectral domain.

[0395] EEE 90. The method according to EEE 88 or 89 further includes:

[0396] The speech pronunciation noise event is attenuated based on its determined type and range.

[0397] EEE 91. The method according to any one of the preceding EEEs, wherein the speech articulation noise event includes at least one of the following: a mouth click event or a speech plosive event.

[0398] EEE 92. The method according to EEE 91, wherein the speech pronunciation noise event includes one or more mouth click events; and wherein the one or more mouth click events include at least one of the following: a non-speech click event, a speech click event, or a lip smacking event.

[0399] EEE 93. The method according to EEE 92, wherein, after segmenting the input audio signal into multiple audio frames, the method further includes:

[0400] The audio frames are classified as speech frames or non-speech frames.

[0401] EEE 94. The method according to EEE 93, wherein the input audio signal is identified by using a speech activity detector (VAD) and the input audio signal is segmented into speech frames and non-speech frames.

[0402] EEE 95. The method according to any one of EEE 92 to 94, wherein the segmentation is performed by using two different window sizes, one of which is shorter than the other.

[0403] EEE 96. The method according to EEE 95 when attached to EEE 93 or 94, wherein a shorter window size is used to detect speech click events in the speech frame, and a longer window size is used to detect non-speech click events in the non-speech frame.

[0404] EEE 97. The method according to any one of EEE 91 to 96, wherein obtaining at least one feature parameter from the audio frame comprises:

[0405] For each audio frame, at least one kurtosis metric is obtained based on the temporal sample amplitude of the audio frame, and

[0406] The determination of the corresponding type and range of the speech pronunciation noise event in the input audio signal based on the obtained feature parameters includes:

[0407] Compare the obtained kurtosis metric with a predefined kurtosis threshold; and

[0408] If the kurtosis metric exceeds the predefined kurtosis threshold, the audio frame is determined to include a mouth click event, and the start and end boundaries of the mouth click event are determined based on the corresponding positions where the kurtosis metric rises above and falls below the predefined kurtosis threshold.

[0409] EEE 98. The method according to any one of EEE 93 to 97, wherein obtaining at least one feature parameter from the audio frame comprises:

[0410] For each speech frame, obtain the corresponding residual approximation without speech harmonic components and the corresponding first kurtosis measure of the sample amplitude of the residual approximation, and

[0411] The determination of the corresponding type and range of the speech pronunciation noise event in the input audio signal based on the obtained feature parameters includes:

[0412] The obtained first kurtosis metric is compared with a first predefined kurtosis threshold; and

[0413] If the first kurtosis metric exceeds the first predefined kurtosis threshold, the speech frame is determined to include a speech click event, and the start and end boundaries of the speech click event are determined based on the corresponding positions where the first kurtosis metric rises above the first predefined kurtosis threshold and falls below the first predefined kurtosis threshold.

[0414] EEE 99. The method according to EEE 98, wherein the residual without speech harmonic components is approximately a second-order waveform difference.

[0415] EEE 100. The method according to EEE 98 or 99 further includes:

[0416] The second kurtosis metric is obtained from the residual sample amplitude of the speech frame;

[0417] The type and extent of the speech pronunciation noise event are determined based on the second kurtosis metric relative to the first kurtosis metric.

[0418] EEE 101. The method according to any one of EEE 98 to 100 further comprises:

[0419] The scope of the identified speech click event is refined through the following operations:

[0420] Locate the sample position with the largest second-order difference within the determined range of the speech click event; and

[0421] The fine-grained range of the speech click event is determined by applying a predefined duration of speech click events around the located sample location.

[0422] EEE 102. The method according to any one of EEE 98 to 101, further comprising:

[0423] The range of the speech click event is further determined based on the minimum / maximum rate of change calculated from the local minimum and local maximum values ​​in the speech frame.

[0424] EEE 103. The method according to any one of EEE 93 to 102, wherein obtaining at least one feature parameter from the audio frame comprises:

[0425] For each nonverbal frame, the corresponding third kurtosis metric of the temporal sample amplitude in the nonverbal frame is obtained, and

[0426] The determination of the corresponding type and range of the speech pronunciation noise event in the input audio signal based on the obtained feature parameters includes:

[0427] The obtained third kurtosis metric is compared with the second predefined kurtosis threshold; and

[0428] If the third kurtosis metric exceeds the second predefined kurtosis threshold, the nonverbal frame is determined to include a nonverbal click event; and the start and end boundaries of the nonverbal click event are determined based on the corresponding positions where the third kurtosis metric rises above and falls below the second predefined kurtosis threshold.

[0429] EEE 104. The method according to EEE 103 further includes:

[0430] If two adjacent non-verbal click events are within a predefined gap threshold, then the two adjacent non-verbal click events are merged into a single verbal click event.

[0431] EEE 105. The method according to EEE 103 or 104, wherein,

[0432] For the nonverbal clicking event identified in the nonverbal frame immediately preceding the verbal frame:

[0433] The high / low frequency band peak ratio is calculated as the amplitude ratio between the maximum peak value above a predefined frequency and the maximum peak value below the predefined frequency; and

[0434] If the calculated high / low frequency peak ratio is higher than a predefined ratio threshold, the non-verbal clicking sound event is determined to be a lip-smacking event.

[0435] EEE 106. The method according to EEE 105, wherein the high / low frequency band peak ratio is calculated as the amplitude ratio between the maximum peak value above a predefined frequency and the maximum peak value below the predefined frequency but above another predefined low frequency.

[0436] EEE 107. The method according to EEE 105 or 106 further includes:

[0437] The range of the determined smacking event is refined based on the high / low frequency band peak ratio, spectral slope, and energy envelope.

[0438] EEE 108. The method according to EEE 107, wherein the scope of the lip-smacking event determined by refinement includes:

[0439] The end position of the smacking event, determined by the third kurtosis metric, is extended if the following conditions are met: the high / low frequency band peak ratio is higher than the predefined ratio threshold, the spectral slope is lower than the predefined slope threshold, and the energy in the energy envelope is reduced.

[0440] EEE 109. The method according to any one of EEE 93 to 102, further comprising:

[0441] The speech articulation noise event is further determined based on the centroid COG calculated for the speech frame according to another predefined threshold, in order to distinguish the mouth click event from the speech transient.

[0442] EEE 110. The method according to any one of EEE 98 to 109, further comprising:

[0443] The determined one or more mouth click events are attenuated based on a corresponding spectral gain, which is obtained from the spectral envelope of the audio frame containing the detected mouth click events and a target envelope calculated based on a corresponding reference frame.

[0444] EEE 111. The method according to EEE 110, wherein, for each detected mouth click event, the reference frame includes audio frames before and after the audio frame containing the detected mouth click event; and wherein the target envelope is calculated by interpolating the spectral envelope of the reference frame.

[0445] EEE 112. The method according to EEE 110 or 111, wherein the attenuation is applied for frequency bands above a predefined high-frequency threshold.

[0446] EEE 113. The method according to any one of EEE 98 to 109, further comprising:

[0447] Replace one or more identified mouth click events based on the corresponding adjacent audio frames.

[0448] EEE 114. The method according to EEE 91, wherein the speech articulation noise event includes at least one speech plosive event; and wherein obtaining at least one feature parameter from the audio frame includes:

[0449] For each of the audio frames, a corresponding low-frequency energy (LFE) metric is obtained to identify outliers.

[0450] EEE 115. The method according to EEE 114, wherein the LFE metric is calculated in the time domain or in the spectral domain.

[0451] EEE 116. The method according to EEE 114 or 115 further includes:

[0452] The range of the speech plosive events is determined based on the outlier identified from the LFE metric and a threshold calculated based on the LFE metric, or based on the LFE ratio calculated from previous and current audio frames.

[0453] EEE 117. The method according to EEE 116 further includes:

[0454] For each of the audio frames, a corresponding zero-crossing maximum value (ZCM) metric is obtained to refine the range of the speech plosive events already determined based on the LFE metric.

[0455] The ZCM metric indicates the length of the maximum interval between consecutive zero-crossings within the audio frame.

[0456] EEE 118. The method according to EEE 116 or 117 further includes:

[0457] The attenuation is determined by the speech plosive event, wherein the attenuation is performed in the time domain or the spectral domain.

[0458] EEE 119. The method according to EEE 118, wherein temporal attenuation is performed by applying a high-pass filter, wherein the cutoff frequency of the filter is determined based on the ZCM metric of the audio frame within the determined range of speech plosive events; and wherein the order of the filter is determined based on the LFE metric of the audio frame within the determined range of speech plosive events.

[0459] EEE 120. The method according to EEE 118, wherein spectral domain attenuation is performed by using an overlapped summative short-time Fourier transform (STFT) with adaptive spectral slope and frequency.

[0460] EEE 121. The method according to EEE 118 or 120, wherein the spectral domain attenuation involves processing the audio frame with a Fast Fourier Transform (FFT), applying an attenuation gain with an adaptive slope and frequency, applying an inverse FFT, windowing, and overlapping addition to produce an attenuated output audio signal; wherein the frequency is determined based on a ZCM metric of the audio frame within the determined range of speech plosive events; and wherein the slope is determined based on an LFE metric of the audio frame within the determined range of speech plosive events.

[0461] EEE 122. The method according to EEE 121 further includes:

[0462] Noise spectrum estimation is applied to limit the attenuation gain to prevent oversuppression.

[0463] EEE 123. The method according to any one of EEE 114 to 122, further comprising:

[0464] A content classifier is applied to the audio frames to distinguish between speech frames and non-speech frames in order to determine the speech plosive events.

[0465] EEE 124. The method according to EEE 118, wherein the spectral domain attenuation involves:

[0466] By using an analytical filter bank, multiple approximately equivalent rectangular bandwidth ERB spacing bands below a predefined frequency threshold and multiple bands above the predefined frequency threshold, which is within the frequency range of the determined speech plosive events;

[0467] Multiple attenuation gains are applied to the audio signal in each frequency band of the frequency band, wherein the attenuation gains are calculated based on the energy calculated for the frequency band; and

[0468] The attenuated audio samples are fed into a synthesis filter bank to generate the output audio signal.

[0469] EEE 125. The method according to EEE 124, wherein the attenuation gain in each frequency band is further constrained to prevent the energy of the frequency band from falling below the estimated noise floor in the frequency band.

[0470] EEE 126. The method according to EEE 125 further includes:

[0471] Calculate time-smoothed low-frequency energy estimates for audio samples that are above the estimated noise floor to distinguish speech plosive events from higher-frequency content in the input audio signal.

[0472] EEE 127. The method according to EEE 126 further includes:

[0473] Calculate the speech harmonic protection metric in the spectrum of the input audio signal; and

[0474] The attenuation gain is calculated based on the speech harmonic protection metric and the time-smoothed low-frequency energy estimate.

[0475] EEE 128. The method according to EEE 127, wherein the speech harmonic protection metric is a periodic metric or a tonality metric.

[0476] EEE 129. The method according to EEE 128, wherein the periodicity measure in the spectrum is calculated from the cepstrum of the audio sample before the final frequency band calculation of the analytical filter bank.

[0477] EEE 130. The method according to EEE 128, wherein the tonality measure in the spectrum is calculated based on the main lobe of the spectral peak compared with the main lobe of the sine peak before the final frequency band calculation of the analytical filter bank.

[0478] EEE 131. The method according to any one of EEE 127 to 130, further comprising:

[0479] The calculated attenuation gain is further constrained by the lower frequency band of the immediate vicinity.

[0480] EEE 132. A method for performing automatic audio enhancement on an input audio signal to detect and / or attenuate at least one speech articulation noise event contained therein, the speech articulation noise event including at least one speech plosive event, the method comprising:

[0481] By using an analytical filter bank, multiple approximately equivalent rectangular bandwidth ERB interval frequency bands below a predefined frequency threshold and multiple frequency bands above the predefined frequency threshold, the predefined frequency threshold being within the frequency range of the speech plosive event;

[0482] Multiple attenuation gains are applied to the audio signal in each frequency band of the frequency band, wherein the attenuation gains are calculated based on the energy calculated for the frequency band; and

[0483] The attenuated audio samples are fed into a synthesis filter bank to generate the output audio signal.

[0484] EEE 133. The method according to EEE 132, wherein the attenuation gain in each frequency band is further constrained to prevent the energy of the frequency band from falling below the estimated noise floor in the frequency band.

[0485] EEE 134. The method according to EEE 133 further includes:

[0486] Calculate time-smoothed low-frequency energy estimates for audio samples that are above the estimated noise floor to distinguish speech plosive events from higher-frequency content in the input audio signal.

[0487] EEE 135. The method according to EEE 132 or EEE 134 further includes:

[0488] Calculate the speech harmonic protection metric in the spectrum of the input audio signal; and

[0489] The attenuation gain is calculated based on the speech harmonic protection metric and the time-smoothed low-frequency energy estimate.

[0490] EEE 136. The method according to EEE 135, wherein the speech harmonic protection metric is a periodic metric or a tonality metric.

[0491] EEE 137. The method according to EEE 136, wherein the periodicity measure in the spectrum is calculated from the cepstrum of the audio sample before the final frequency band calculation of the analytical filter bank.

[0492] EEE 138. The method according to EEE 136, wherein the tonality measure in the spectrum is calculated based on the main lobe of the spectral peak compared with the main lobe of the sine peak before the final frequency band calculation of the analytical filter bank.

[0493] EEE 139. The method according to EEE 132 to 138 further includes:

[0494] The calculated attenuation gain is further constrained by the lower frequency band of the immediate vicinity.

[0495] EEE 140. The method according to any one of EEE 132 to 139, wherein the input audio signal is processed continuously with a predefined look-ahead frame size.

[0496] EEE 141. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to perform a method according to any one of the preceding EEE claims.

[0497] EEE 142. A program comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of EEE 88 to 140.

[0498] EEE 143. A computer-readable storage medium storing a program according to EEE 142.

Claims

1. A method of performing automatic audio enhancement on an input audio signal comprising at least one speech articulatory noise event, the method comprising: segmenting the input audio signal into a plurality of audio frames; obtaining, from the audio frames, at least one feature parameter associated with a respective speech articulatory noise event; determining, based at least partly on the obtained feature parameter, a respective type of the speech articulatory noise event and a respective time-frequency range associated with the speech articulatory noise event within the input audio signal; and attenuating the speech articulatory noise event according to the determined type and range of the speech articulatory noise event, wherein the speech articulatory noise event comprises at least one of: a mouth click event or a speech pop event; and wherein obtaining, from the audio frames, at least one feature parameter comprises: for each audio frame, obtaining at least one kurtosis measure based on time-domain sample amplitudes of the audio frame, and wherein determining, based on the obtained feature parameter, a respective type of the speech articulatory noise event and a respective range of the speech articulatory noise event in the input audio signal comprises: comparing the obtained kurtosis measure to a predefined kurtosis threshold; and if the kurtosis measure exceeds the predefined kurtosis threshold, determining that the audio frame comprises a mouth click event, and determining a start boundary and an end boundary of the mouth click event based on respective locations where the kurtosis measure rises above and falls below the predefined kurtosis threshold. the determined range comprises at least one boundary of the determined speech articulatory noise event in a time domain and / or a spectral domain.

2. The method of claim 1, wherein, the speech articulatory noise event comprises one or more mouth click events; and wherein the one or more mouth click events comprise at least one of: a non-speech click event, a speech click event, or a smacking event.

3. The method of claim 1 or 2, wherein, after segmenting the input audio signal into a plurality of audio frames, the method further comprises:

4. The method of claim 3, wherein, classifying the audio frames as speech frames or non-speech frames. segmenting the input audio signal into the speech frames and the non-speech frames by using a voice activity detector, VAD.

5. The method of claim 4, wherein, performing the segmentation by using two different window sizes, one of which is shorter than the other.

6. The method of claim 4, wherein, the shorter window size is used to detect speech click events in the speech frames, and the longer window size is used to detect non-speech click events in the non-speech frames.

7. The method of claim 6, wherein, obtaining, from the audio frames, at least one feature parameter comprises:

8. The method of claim 4, wherein, for each speech frame, obtaining a respective residual approximation having no speech harmonic components and a respective first kurtosis measure of sample amplitudes of the residual approximation, and wherein determining, based on the obtained feature parameter, a respective type of the speech articulatory noise event and a respective range of the speech articulatory noise event in the input audio signal comprises: comparing the obtained first kurtosis measure to a first predefined kurtosis threshold; and if the first kurtosis measure exceeds the first predefined kurtosis threshold, determining that the speech frame comprises a speech click event, and determining a start boundary and an end boundary of the speech click event based on respective locations where the first kurtosis measure rises above and falls below the first predefined kurtosis threshold. If the first kurtosis measure exceeds the first predefined kurtosis threshold, it is determined that the speech frame comprises a speech click event, and a start boundary and an end boundary of the speech click event are determined based on respective positions where the first kurtosis measure rises above and falls below the first predefined kurtosis threshold.

9. The method of claim 8, wherein, The residual approximation without speech harmonic components is a second order waveform difference.

10. The method of claim 8, further comprising: obtaining a second kurtosis measure from residual sample amplitudes of the speech frame; wherein a type and a range of the speech articulatory noise event are determined based on the second kurtosis measure with respect to the first kurtosis measure.

11. The method of claim 8, further comprising: refining the determined range of the speech click event by: locating a sample position with a maximum second order difference within the determined range of the speech click event; and and determining a refined range of the speech click event by applying a predefined speech click event duration around the located sample position.

12. The method of claim 8, further comprising: further determining the range of the speech click event based on a minimum / maximum change rate computed from local minima and local maxima in the speech frame.

13. The method of claim 4, wherein, obtaining at least one feature parameter from the audio frame comprises: for each non-speech frame, obtaining a respective third kurtosis measure of time domain sample amplitudes in the non-speech frame, and wherein determining respective types of the speech articulatory noise events in the input audio signal and respective ranges of the speech articulatory noise events based on the obtained feature parameters comprises: comparing the obtained third kurtosis measure with a second predefined kurtosis threshold; and if the third kurtosis measure exceeds the second predefined kurtosis threshold, determining that the non-speech frame comprises a non-speech click event; and determining a start boundary and an end boundary of the non-speech click event based on respective positions where the third kurtosis measure rises above and falls below the second predefined kurtosis threshold.

14. The method of claim 13, further comprising: if two adjacent non-speech click events are within a predefined gap threshold, merging the two adjacent non-speech click events into a single speech click event.

15. The method of claim 13, wherein, for a non-speech click event determined in a non-speech frame immediately preceding a speech frame: calculating a high / low band peak ratio as an amplitude ratio between a maximum peak above a predefined frequency and a maximum peak below the predefined frequency; and if the calculated high / low band peak ratio is higher than a predefined ratio threshold, determining the non-speech click event as a smacking event.

16. The method of claim 15, wherein, calculating the high / low band peak ratio as an amplitude ratio between a maximum peak above a predefined frequency and a maximum peak below the predefined frequency but above another predefined low frequency.

17. The method of claim 15, further comprising: refining the determined extent of the smacking event based on the high / low band peak ratio, the spectral slope, and the energy envelope.

18. The method of claim 17, wherein, refining the determined extent of the smacking event includes: extending the end position of the smacking event determined using the third kurtosis measure as long as the high / low band peak ratio is above the predefined ratio threshold, the spectral slope is below a predefined slope threshold, and the energy in the energy envelope decreases.

19. The method according to claim 4, further comprising: determining the speech articulatory noise event further based on a center of gravity, COG, computed for the speech frame according to another predefined threshold to distinguish a mouth click event from a speech transient.

20. The method according to claim 8, further comprising: attenuating the determined one or more mouth click events based on a respective spectral gain derived from a spectral envelope of the audio frame containing the detected mouth click event and a target envelope computed based on a respective reference frame.

21. The method of claim 20, wherein, for each detected mouth click event, the reference frame comprises audio frames before and after the audio frame containing the detected mouth click event; and wherein the target envelope is computed by interpolating spectral envelopes of the reference frames.

22. The method of claim 20, wherein, applying the attenuation for frequency bands above a predefined high frequency threshold.

23. The method according to claim 8, further comprising: replacing the determined one or more mouth click events based on respective adjacent audio frames.

24. The method of claim 2, wherein, the speech articulatory noise event comprises at least one speech burst event; and wherein obtaining at least one feature parameter from the audio frames comprises: obtaining for each of the audio frames a respective low frequency energy, LFE, measure to identify outliers thereof.

25. The method of claim 24, wherein, computing the LFE measure in the time domain or in the spectral domain.

26. The method according to claim 24, further comprising: determining an extent of the speech burst event according to the outliers identified from the LFE measure and a threshold computed based on the LFE measure or according to LFE ratios computed from preceding and current audio frames.

27. The method according to claim 26, further comprising: obtaining for each of the audio frames a respective zero-crossing maximum, ZCM, measure to refine the extent of the speech burst event already determined based on the LFE measure, wherein the ZCM measure is indicative of a length of a maximum interval of consecutive zero-crossings within the audio frame.

28. The method according to claim 26, further comprising: attenuating the determined speech burst event, wherein the attenuation is performed in the time domain or in the spectral domain.

29. The method of claim 28, wherein, performing time domain attenuation by applying a high-pass filter, wherein a cut-off frequency of the filter is determined based on ZCM measures of the audio frames within the determined extent of the speech burst event; and wherein an order of the filter is determined based on LFE measures of the audio frames within the determined extent of the speech burst event.

30. The method of claim 28, wherein, The spectral domain attenuation involves processing the audio frames with a Fast Fourier Transform, FFT, applying attenuation gains with an adaptive slope and frequency, applying an inverse FFT, windowing and overlap-add in order to produce an output audio signal after attenuation; wherein the frequency is determined based on a ZCM measure of audio frames within a range of a determined speech burst event; and wherein the slope is determined based on an LFE measure of audio frames within a range of a determined speech burst event.

31. The method of claim 28, wherein, 32. The method of claim 31, further comprising: applying a noise spectrum estimate to limit the attenuation gains to prevent over-inhibition.

33. The method of claim 24, further comprising: applying a content classifier to the audio frames to distinguish speech frames from non-speech frames in order to determine the speech burst event. The spectral domain attenuation involves:

34. The method of claim 28, wherein, producing a plurality of approximately Equivalent Rectangular Bandwidth, ERB, spaced frequency bands below a predefined frequency threshold and a plurality of frequency bands above the predefined frequency threshold, the predefined frequency threshold being within a frequency range of a determined speech burst event, by using an analysis filter bank; applying a plurality of attenuation gains to the audio signal in each of the frequency bands, respectively, wherein the attenuation gains are computed based on an energy computed for the frequency band; and feeding the attenuated audio samples to a synthesis filter bank to generate an output audio signal. The attenuation gains in each frequency band are further constrained to not reduce the energy of the frequency band below an estimated floor noise in the frequency band.

35. The method of claim 34, wherein, 36. The method of claim 35, further comprising: computing a time-smoothed low frequency energy estimate of audio samples above an estimated floor noise to distinguish speech burst events from higher frequency content in the input audio signal.

37. The method of claim 36, further comprising: computing a speech harmonic protection measure in a spectrum of the input audio signal; and computing the attenuation gains from the speech harmonic protection measure and the time-smoothed low frequency energy estimate. The speech harmonic protection measure is a periodicity measure or a tonality measure.

38. The method of claim 37, wherein, The periodicity measure in the spectrum is computed from a cepstrum of the audio samples prior to final frequency band computation of the analysis filter bank.

39. The method of claim 38, wherein, The tonality measure in the spectrum is computed based on a main lobe of a spectral peak compared to a main lobe of a sinusoidal peak prior to final frequency band computation of the analysis filter bank.

40. The method of claim 38, wherein, 41. The method of claim 37, further comprising: further constraining the computed attenuation gains based on a frequency band immediately lower in frequency.

42. A method of performing automatic audio enhancement on an input audio signal to attenuate at least one speech burst event contained therein in a spectral domain, the speech burst event comprising at least one speech burst event, the method comprising: ​ a plurality of approximately equivalent rectangular bandwidth, ERB, spaced frequency bands below a predefined frequency threshold and a plurality of frequency bands above the predefined frequency threshold, the predefined frequency threshold being within a frequency range of the speech plosive event, by using an analysis filter bank; applying a plurality of attenuation gains to the audio signal in each of the frequency bands, respectively, wherein the attenuation gains are computed based on an energy computed for the frequency band; and feeding the attenuated audio samples to a synthesis filter bank to generate an output audio signal.

43. The method of claim 42, wherein, The attenuation gain in each frequency band is further constrained to not reduce the energy of the frequency band below an estimated floor noise in the frequency band.

44. The method according to claim 43, further comprising: computing a time-smoothed low-frequency energy estimate of the audio samples above an estimated floor noise to distinguish speech plosive events from higher frequency content in the input audio signal.

45. The method according to claim 44, further comprising: computing a speech harmonic protection measure in a spectrum of the input audio signal; and computing the attenuation gains from the speech harmonic protection measure and the time-smoothed low-frequency energy estimate.

46. The method of claim 45, wherein, The speech harmonic protection measure is a periodicity measure or a tonality measure.

47. The method of claim 46, wherein, The periodicity measure in the spectrum is computed from a cepstrum of the audio input samples before the final frequency band computation of the analysis filter bank.

48. The method of claim 46, wherein, The tonality measure in the spectrum is computed based on a main lobe of a spectral peak compared to a main lobe of a sinusoidal peak before the final frequency band computation of the analysis filter bank.

49. The method according to any one of claims 42 to 44, further comprising: further constraining the computed attenuation gains based on a frequency band immediately lower in frequency.

50. The method of any one of claims 42 to 44, wherein, processing the input audio signal continuously with a predefined lookahead frame size.

51. An apparatus comprising a processor and a memory coupled to the processor, wherein, The processor is adapted to cause the apparatus to perform the method according to any one of claims 1 to 50.

52. A program comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 50.

53. A computer-readable storage medium storing a program according to claim 52.