Adjusting sibilance detection based on detecting specific sounds in an audio signal

By detecting short-term and long-term characteristics in the audio signal, adjusting the tooth sound detection parameters, and using machine learning classifiers for tooth sound detection and suppression, the problem of degradation of audio signal quality in the prior art is solved, and is especially suitable for low-fidelity equipment.

CN114127848BActive Publication Date: 2025-06-24DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080051216.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-08
Filing Date
2020-07-16
Publication Date
2025-06-24
Estimated Expiration
2040-07-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and adjust the tooth sound in the audio signal, resulting in a degradation of the audio signal quality, especially in low-fidelity devices.

Method used

By detecting short-term and long-term features in the audio signal, adjusting the tooth sound detection parameters, performing tooth sound detection using a classifier based on supervised or unsupervised machine learning, and suppressing tooth sound through a multi-band compressor.

Benefits of technology

It realizes effective detection and suppression of tooth sound while maintaining the quality of the audio signal, and is especially suitable for low-fidelity equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114127848B_ABST
    Figure CN114127848B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a method for adjusting parameters of a sibilance detector. Time-frequency features are extracted from an audio signal being received. Based on these time-frequency features, it is determined whether the audio signal includes short-term features or long-term features. According to the determination that the audio signal includes short-term features or long-term features, one or more parameters of the sibilance detector for detecting sibilance in the audio signal are adjusted. The sibilance detector with one or more adjusted parameters is used to detect sibilance in the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 884,320, filed on August 8, 2019, and International Application No. PCT / CN2019 / 096399, filed on July 17, 2019, each of which is hereby incorporated by reference in its entirety. Technical field

[0003] Embodiments of the present disclosure generally relate to audio signal processing, and more particularly to the adjustment of sibilance detection. Background art

[0004] In phonetics, sibilance refers to speech with strongly pronounced fricative consonants (e.g., s, sh, ch, z, v, and f). These consonants are produced when the air passing through the vocal tract is restricted by the position of the tongue and lips. Sibilance in an audio signal typically lies in the frequency range from 4 kHz ("kilohertz") to 12 kHz, depending on the speaker. If the energy of the sibilance is high, the speech will have an unnatural harshness, which will degrade the quality of the audio signal and annoy the listener. Summary of the invention

[0005] The disclosed embodiments detect short - term and long - term features in an audio signal and adjust sibilance detection to avoid misinterpreting features as excessive sibilance in the audio signal. The advantage of the disclosed systems and methods is that the quality of the audio signal is maintained by not suppressing short - term or long - term features that may be an expected part of the audio content. The disclosed systems and methods are particularly useful for low - fidelity devices, such as low - quality headsets with poor microphone frequency response at high frequencies or mobile devices with low - quality speakers.

[0006] In some aspects, the present disclosure describes a method for adjusting sibilance parameters and using the adjusted sibilance parameters in sibilance detection. The system receives an audio signal (e.g., movie score, music, user-generated audio, or podcast) and extracts a plurality of time-frequency features (e.g., energy data of multiple frequency bands) from the audio signal. The time-frequency features include short-term features (such as plosives and / or fricatives (e.g., the sound of the letter "f")) and / or long-term features (such as smoothed spectral balance features). Based on determining that the input signal includes short-term features and / or long-term features, the system adjusts one or more parameters of a sibilance detector for detecting sibilance in the audio signal. Using the sibilance detector with one or more adjusted parameters, the system continues to detect sibilance in the audio signal and suppresses the sibilance using a multi-band compressor, or uses the detected sibilance for any other desired application. In an embodiment, the sibilance detector is implemented using a supervised or unsupervised machine learning-based classifier (e.g., a neural network), which is trained on audio samples having one or more short-term features and / or long-term features.

[0007] These and other aspects, features, and embodiments can be expressed as methods, apparatuses, systems, components, program products, constructs, or steps for performing functions, and can be expressed in other ways.

[0008] These and other aspects, features, and embodiments will become apparent from the following description, including the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the drawings, for ease of description, a particular arrangement or ordering of schematic elements is shown, such as those representing devices, modules, instruction blocks, and data elements. However, those skilled in the art should understand that the particular ordering or arrangement of schematic elements in the drawings does not imply a particular processing order or sequence, or separation of processes. Further, including schematic elements in the drawings does not mean that such elements are required in all embodiments, or that in some embodiments, the features represented by such elements may not be included in other elements or combined with other elements.

[0010] Further, in the drawings that use connecting elements, such as solid lines, dashed lines, or arrows, to illustrate the connection, relationship, or association between two or more other schematic elements, the absence of any such connecting element does not mean that a connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, in the case where the connecting element represents the communication of signals, data, or instructions, those skilled in the art should understand that such an element represents one or more signal paths that may be required to effect the communication.

[0011] Figure 1A is a block diagram of a system for adjusting parameters for fricative detection according to some embodiments of the present disclosure.

[0012] Figure 1B is a block diagram of a system for adjusting parameters for fricative detection according to some embodiments of the present disclosure, the system including an plosive detector and a flat fricative detector.

[0013] Figure 2 illustrates operations for adjusting parameters used in fricative detection according to some embodiments of the present disclosure.

[0014] Figure 3 illustrates operations performed by a fricative detection module for detecting plosives according to some embodiments of the present disclosure.

[0015] Figure 4 illustrates operations performed by a fricative detection module for detecting flat fricatives according to some embodiments of the present disclosure.

[0016] Figure 5 illustrates operations for further determining whether a fricative is present according to some embodiments of the present disclosure.

[0017] Figure 6 illustrates a fricative suppression curve that can be used in fricative suppression according to some embodiments of the present disclosure.

[0018] Figure 7 is a block diagram for implementing fricative detection according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0019] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it is apparent that the present disclosure may be practiced without these specific details.

[0020] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to those of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described below, each of which may be used independently of one another or in any combination with any of the other features.

[0021] As used herein, the term "comprising" and variations thereof will be understood to be open-ended terms meaning "including but not limited to". Unless the context clearly dictates otherwise, the term "or" will be understood to mean "and / or". The term "based on" will be understood to mean "at least partially based on".

[0022] Figure 1A and Figure 1B is a block diagram of a system 100 for adjusting parameters for detecting sibilants according to some embodiments of the present disclosure. The system 100 includes a transformation module 110, a banding module 120, a sibilant detection module 130, a multi-band compressor 140, and an inverse transformation module 150. Figure 1A includes a short-term feature detector 131 for detecting short-term features in an audio signal. In some embodiments, the short-term features include detection of transient sounds such as percussion strikes of percussion instruments. These sounds typically have a short duration, sometimes about five milliseconds. Figure 1B includes two examples of short-term feature detectors. An impact sound detector 132 is used to detect impact sounds, such as percussion strikes of percussion instruments like cymbals, while a flat fricative voice detector 136 is used to detect flat fricative sounds (e.g., the letter v sound, the letter t sound, the letter f sound, or the "th" sound). In some embodiments, the impact sound detector 132 and the flat fricative voice detector 136 are combined into a single detector module.

[0023] The transformation module 110 is configured to receive an audio signal and transform the audio signal into a desired transform domain. In some embodiments, the audio signal includes speech and non-speech sounds. To perform sibilant parameter adjustment, the transformation module 110 performs a transformation operation (e.g., using a filter bank) on frames of the audio signal to transform the audio signal into multiple spectral feature bands in the frequency domain. For example, the transformation module 110 may perform a fast Fourier transform (FFT), a modified discrete cosine transform (MDCT), an orthogonal mirror filter (QMF), or another transformation algorithm to transform the audio signal from the time domain to the frequency domain or the time-frequency domain. In some embodiments, the transformation module outputs multiple equally spaced frequency bins.

[0024] The banding module 120 performs a banding operation that groups or aggregates the output of the transformation module 110 (e.g., the frequency bins generated by the transformation module 110) into multiple frequency bands (e.g., equivalent rectangular bandwidth ("ERB") bands). In some embodiments, a 1 / 3 octave filter bank is used in the banding module. The frequency bands include a sibilant frequency band (e.g., from about 4 kHz to about 12 kHz) and a non-sibilant frequency band (e.g., below 4 kHz and from about 12 kHz to about 16 kHz). In an embodiment, the sibilant detection module 130 includes a short-term feature detector 131, a short-term sibilant detector 134, and a long-term sibilant detector 136, as Figure 1AAs shown. The sibilance detection module 130 and its components will be discussed in further detail in this disclosure. The multi-band compressor 140 modifies the gain applied to the sibilant band and / or the non-sibilant band based on the output of the sibilance detection module 130. In some embodiments, the gain on a particular band is mapped to the gain to be applied to a subset of the frequency bins output by the transform module (110). After applying the gain, the bands are input to the inverse transform module 150, in which the bands are transformed back to the time domain. The time-domain audio signal is then sent to one or more output devices (e.g., a speaker system, a storage device).

[0025] The actions performed in this disclosure will be described as being performed by the sibilance detection module. It should be noted that the sibilance detection module may include software, hardware, or a combination of both. Example embodiments of the hardware that may be used to implement system 100 are described with respect to Figure 7 Further description. Although the example embodiments described below include plosive detection and fricative detection that respectively provide short-term features, embodiments may use any detected short-term features.

[0026] Figure 1B is a block diagram of a system for adjusting sibilance detection parameters according to some embodiments of the present disclosure, the system including a plosive detector and a fricative detector.

[0027] Figure 2 Illustrates actions for adjusting parameters used in sibilance detection. At 202, the sibilance detection module 130 receives an audio signal. The audio signal is received and processed by the transform module 110 and the banding module 120. As discussed above, the transform module 110 transforms the audio signal from the time domain to the frequency domain, and the banding module 120 groups or aggregates the output of the transform module 110 into multiple bands, including a sibilant audio band and a non-sibilant audio band.

[0028] At 204, the sibilance detection module 130 extracts a plurality of time-frequency features from the audio signal. These features include the energy level of each band in the sibilant audio band of a particular frame of the audio signal. At 206, the sibilance detection module 130 uses the plurality of time-frequency features to determine whether the audio signal includes a plosive or a fricative. The sibilance detection module 130 is configured to detect plosives and fricatives in parallel or serially depending on the resources available to the module.

[0029] In embodiments including a plosive detector 132, the plosive detector 132 determines whether the audio signal includes a plosive. The plosive detector 132 may include both software components and hardware components. In some embodiments, short-term time-frequency features (e.g., about 5 milliseconds) are used to detect plosives.

[0030] Figure 3Illustrates the actions performed by the sibilance detection module 130 to detect impact sounds. At 302, for a first time interval in the audio signal, the sibilance detection module 130 calculates a first total power in one or more sibilant frequency bands and a second total power in one or more non-sibilant frequency bands. In an embodiment, the sibilance detection module 120 uses Equation 1 (below) to perform the calculation of the sibilant frequency bands:

[0031]

[0032] where b is the number of sibilant frequency bands, P b is the power in sibilant frequency band b, and n is the first time interval (e.g., the current frame or current time period). In an embodiment, the sibilance detection module 130 uses Equation 2 (below) to perform the calculation for the non-sibilant frequency bands:

[0033]

[0034] where b is the number of non-sibilant frequency bands, P b is the power in non-sibilant frequency band b, and n is the first time interval (e.g., the current frame or current time period). As discussed above, the sibilant frequency bands include frequencies between approximately 4 kHz and approximately 12 kHz, and the non-sibilant frequency bands include frequencies below approximately 4 kHz and between approximately 12 kHz and approximately 16 kHz.

[0035] At 304, for a second time interval (e.g., an earlier time interval), the sibilance detection module 130 determines a third total power in one or more sibilant frequency bands and a fourth total power in one or more non-sibilant frequency bands. For example, in an embodiment, the sibilance detection module 130 uses Equation 3 (below) to perform the calculation for the sibilant frequency bands of a previous time interval (e.g., a previous frame):

[0036]

[0037] where b is the number of sibilant frequency bands, P b is the power in sibilant frequency band b, n is the first time interval (e.g., the current frame or time period), and k is an integer such that [n - k] is a previous time interval (e.g., a previous frame). In some embodiments, k is an integer in the range of one to three.

[0038] In an embodiment, the sibilance detection module 130 uses Equation 4 (below) to perform the calculation for the non-sibilant frequency bands of a previous time interval (e.g., a previous frame):

[0039]

[0040] where b is the number of non-sibilant frequency bands, P bP is the power in the non-sibilant audio band, n is the first time interval (e.g., the current frame or time period), and k is an integer that makes [n - k] the previous time interval (e.g., the previous frame or time period). In some embodiments, k is an integer in the range of one to three.

[0041] At 306, the sibilance detection module 130 determines a first flux value based on the difference between the first total power and the third total power, and determines a second flux value based on the difference between the second total power and the fourth total power. For example, in an embodiment, the sibilance detection module 130 uses Equation 5 (below) to calculate the first flux value:

[0042]

[0043] where P sib_bands [n] is the total power of the sibilant audio band for the time interval n (e.g., the current time interval or current frame), and is the total power of the sibilant audio band for the previous time interval [n - k], where k can be an integer between one and three. In some embodiments, k can be a larger integer.

[0044] In an embodiment, the sibilance detection module 130 uses Equation 6 (below) to calculate the second flux value:

[0045] S non_sib_bands [n] = P non_sib_bands [n] - P non_sib_bands [n - k] Equation 6

[0046] where P non_sib_bands [n] is the total power of the non-sibilant audio band for the time interval n (e.g., the current time interval or current frame), and P non_sib_bands [n - k] is the total power of the non-sibilant audio band for the previous time interval [n - k], where k is an integer between one and three. In some embodiments, k can be a larger integer.

[0047] At 308, the sibilance detection module 130 determines whether the first flux value meets a first threshold and whether the second flux value meets a second threshold. If both the first flux value and the second flux value meet their respective thresholds, the process 300 moves to 310, where the sibilance detection module 130 determines that a plosive sound exists. If either the first flux value or the second flux value does not meet its respective threshold, the process 300 moves to 312, where the sibilance detection module 130 determines that no plosive sound exists. The logic of Equation 7 (below) illustrates the determination regarding the existence of a plosive sound:

[0048]

[0049] where Ssib_bands is the flux value of the sibilant audio band at time interval n (e.g., the current frame), Th sib_band is the threshold of the sibilant audio band, S non_sib_bands [n] is the flux value of the non-sibilant audio band, and Th non_sib_band is the threshold of the non-sibilant audio band. In some embodiments, the threshold is ten decibels (“dB”). In some embodiments, if I[n]=1, the sibilance detection module 130 determines that a plosive sound exists. If I[n]=0, the sibilance detection module 130 determines that no plosive sound exists.

[0050] In some embodiments, before outputting a decision on whether a plosive sound is detected, the sibilance detection module 130 applies smoothing to the value output by Equation 7. The logic of Equation 8 (below) illustrates the smoothing operation:

[0051]

[0052] where α A is the attack time constant, which has a value of 0 seconds in some embodiments, and α R is the release time constant, which has a value of one second in some embodiments. Thus, I smooth [n] is the output of the plosive sound detector 132 (i.e., R ISD [n]=I smooth [n]).

[0053] In some embodiments, the attack time constant and the release time constant are adjusted based on the type of plosive sound. For example, one type of plosive sound may be longer than another type of plosive sound. In this case, the release time constant can be increased. In another example, one type of plosive sound has lower energy at the start of the sound (e.g., below the threshold), and thus, the attack time constant is increased.

[0054] In some embodiments, the sibilance detection module 130 identifies the type of plosive sound based on time-frequency characteristics. In some embodiments, the sibilance detection module 120 can access known plosive sounds and corresponding energy and / or flux levels. That is, a given sound can have a set of specific energy and / or flux levels in both the sibilant audio band and the non-sibilant audio band. In some embodiments, these energy levels and / or flux levels are stored and compared with the energy levels and / or flux levels of the detected plosive sound. The comparison is repeated for all known plosive sounds to identify the received plosive sound.

[0055] In some embodiments, the sibilance detection module 130 identifies the type of impact sound based on the fluxes in the sibilant frequency band and the non-sibilant frequency band, using different thresholds for the sibilant frequency band and the non-sibilant frequency band. For example, each known impact sound can be associated with a specific sibilant threshold and a specific non-sibilant threshold. Thus, the sibilant threshold for impact sound type A can be 15 dB and the non-sibilant threshold can be 8 dB. The sibilant frequency band threshold for impact sound B can be 20 dB and the non-sibilant frequency band threshold can be 15 dB. Thus, when calculating the flux values for both the sibilant frequency band and the non-sibilant frequency band, these flux values are compared with the flux values of each known impact sound to determine which impact sound it is. For example, the closest sibilant and non-sibilant threshold matches can be used to determine the type of impact sound. The logic of Equation 9 (below) illustrates impact sound detection.

[0056]

[0057] Where S sib_bands [n] is the flux value of the sibilant frequency band at time interval n (e.g., the current frame), Th sib_bandA is the sibilant frequency band threshold for impact sound type A, S non_sib_bands [n] is the flux value of the non-sibilant frequency band, and Th non_sib_bandA is the threshold of the non-sibilant frequency band. Additionally, Th sib_bandB is the sibilant frequency band threshold for impact sound type B, and Th non_sib_bandB is the non-sibilant frequency band threshold for impact sound type B.

[0058] In some embodiments, the sibilance detection module 130 uses a counter to generate an output from the impact sound detector 132. The logic of Equation 10 (below) illustrates generating an output from the impact sound detector 132 using a counter:

[0059]

[0060] Where N countdown is a preset reciprocal value, and n is the current time period (e.g., the current frame). In some embodiments, the value depends on the sampling rate and the frame size. In some embodiments, the reciprocal duration is equal to one second. The logic of Equation 11 (below) illustrates using the reciprocal for the output from the impact sound detector 132:

[0061]

[0062] Where I count [n] is the output of the counter in Equation 10.

[0063] In some embodiments, the fricative detection module 130 uses a flat fricative voice detector 136 to determine whether the audio signal includes a flat fricative. In some embodiments, the flat fricative voice detector 136 includes both software components and hardware components. In some embodiments, short-term time-frequency features (e.g., about 5 milliseconds) are used to detect flat fricatives. Generally, flat fricative speech has a flat spectrum compared to fricative sounds (e.g., those with excessive or harsh fricatives). In some embodiments, the fricative spectral flatness is calculated by dividing the geometric mean of the power spectrum by the arithmetic mean of the power spectrum. Thus, flat fricatives can be detected based on the fricative spectral flatness metric ("SSFM"). In some embodiments, the fricative detection module 130 uses Equation 12 (below) to calculate the SSFM:

[0064]

[0065] where X(k) is the fricative speech band spectrum at frequency band index k, and K is the number of frequency bands. In some embodiments, the fricative detection module 120 uses the variance and / or standard deviation of the power in adjacent fricative frequency bands to determine whether a flat fricative is present. In some embodiments, the fricative detection module 120 uses the peak-to-average ratio or peak-to-median ratio of the power in the fricative frequency bands to determine whether a flat fricative is present. In still other embodiments, the fricative detection module 120 uses the spectral entropy of the power in the fricative frequency bands to determine whether a flat fricative is present. The logic of Equation 13 (below) illustrates the output of the flat fricative voice detector 136:

[0066]

[0067] where Th SSFM is the threshold for detection. Thus, if the output of the SSFM is greater than the threshold, the fricative detection module 130 determines that a flat fricative is present.

[0068] Figure 4 Illustrates the actions performed by the fricative detection module 130 to detect flat fricatives. At 402, the fricative detection module 130 calculates the fricative spectral flatness metric based on the fricative speech band spectrum and the number of frequency bands. In some embodiments, the fricative detection module 130 uses Equation 12 to perform the calculation. At 404, the fricative detection module 130 obtains (e.g., from regarding Figure 7(in the memory under discussion) sibilant audio spectral flatness threshold. At 406, the sibilant detection module 130 compares the sibilant audio spectral flatness metric with the sibilant audio spectral flatness threshold. At 408, the sibilant detection module 130 determines whether the sibilant audio spectral flatness metric meets the sibilant audio spectral flatness threshold. If the sibilant audio spectral flatness metric meets the sibilant audio spectral flatness threshold, the process 400 moves to 410, where the sibilant detection module 130 determines that there is a flat fricative. If the sibilant audio spectral flatness metric does not meet the sibilant audio spectral flatness threshold, the process 400 moves to 412, where the sibilant detection module 130 determines that there is no flat fricative.

[0069] Return Figure 2 In process 200, at 208, based on determining that the input signal includes an impulse sound or a flat fricative, the sibilant detection module 130 adjusts one or more parameters of the sibilant detection for detecting sibilants in the audio signal. In some embodiments, at 208, the sibilant detection module adjusts one or more parameters of the sibilant detection for detecting sibilants in the audio signal based on the output from the short-term feature detector 131. For example, the short-term feature detector may include one or more detectors (e.g., impulse sound detector, flat fricative detector, and other suitable detectors). The output of the short-term feature detector 131 is input into the short-term sibilant detector 134. In some embodiments, the sibilant detection module 130 adjusts the sibilant detection threshold based on the output value generated by determining whether an impulse sound is detected and the output value generated by determining whether a flat fricative is detected. In still other embodiments, the sibilant detection module 130 adjusts the sibilant detection threshold based on the output of any suitable feature of the short-term feature detector 131. The sibilant detection module 130 uses the sibilant detection threshold in the short-term sibilant detection operation. Thus, at 210, the sibilant detection module 130 uses the sibilant detection with one or more adjusted parameters to detect sibilants in the audio signal.

[0070] As discussed above, the sibilance detection module includes a short-term sibilance detector 134. In some embodiments, the actions described above are performed by the short-term sibilance detector 134. In those embodiments, the short-term sibilance detector 134 uses the output from any other components of the plosive detector 132, the fricative speech detector 136, and / or the short-term feature detector 131 to determine whether there is a type of sibilance that needs to be suppressed. The short-term sibilance detector 134 can be software, hardware, or a combination of software and hardware. In some embodiments, the sibilance detection module 130 (e.g., using the short-term sibilance detector 134) calculates a spectral balance feature and compares the spectral balance feature with a threshold (e.g., a threshold based on the output of the short-term feature detector including the plosive detector 132, the fricative speech detector 136, and / or any other suitable detector) to determine whether there is sibilance in the audio signal.

[0071] As used herein, the term "spectral balance" refers to the balanced property of signal energy over the speech frequency band. In some cases, the spectral balance characterizes the degree of balance of signal energy over the entire speech frequency band. The term "speech frequency band" as used herein means the frequency band in which the speech signal is located and which, for example, ranges from approximately 0 kHz to approximately 16 kHz. Since sibilance has special spectral distribution characteristics (i.e., sibilant speech is typically concentrated in a certain frequency band), the spectral balance feature is useful for distinguishing non-sibilant speech and sibilant speech.

[0072] In some embodiments, the spectral balance feature is obtained based on the signal energy in the sibilant frequency band and the signal energy in the entire speech frequency band. Specifically, the spectral balance feature can be calculated as the ratio of the signal energy in the sibilant frequency band to the signal energy in the entire speech frequency band. That is, the spectral balance feature can be expressed as the ratio of the sum of the signal energy over all sibilant frequency bands to the sum of the signal energy over the entire speech frequency band.

[0073] In some embodiments, the spectral balance feature is calculated based on the signal energy in the sibilant frequency band and the signal energy in the non-sibilant frequency band. In this case, the speech frequency band is divided into two parts: the sibilant frequency band and the non-sibilant frequency band. That is, the frequency band is divided into two groups of frequency bands, one group that can contain the signal energy of sibilance and another group that does not contain or contains very little signal energy of sibilance. Therefore, the spectral balance feature is calculated as the ratio of the signal energy over the two frequency bands.

[0074] In some embodiments of the present disclosure, the spectral balance feature is determined based on the signal-to-noise ratio (SNR) in the sibilant frequency band and the non-sibilant frequency band. Specifically, the spectral balance feature is determined as the ratio of the two SNRs.

[0075] In some embodiments, the sibilant detection module 130 calculates a threshold for comparison with the spectral balance feature using the output of the short-term detector 131 (e.g., the plosive detector 132 and / or the flat friction voice detector 136). In some embodiments, the sibilant detection module 130 uses the higher value among the output of the plosive detector 132 and the output of the flat friction voice detector 136. For example, if a plosive is detected and the output from the plosive detector 132 is one, but no flat friction voice is detected and the output from the flat friction voice detector is zero, then the sibilant detection module 130 uses the value one as the input to the short-term sibilant detector 134. Thus, in an embodiment, the sibilant detection module 130 uses Equation 14 (below) to determine the threshold:

[0076] Th STSD [n]=Th normal +f(R FFVD [n],R ISD [n])·Th delta Equation 14

[0077] where Th normal is the normal threshold used when no plosive or flat friction voice is detected. In some embodiments, the threshold is -5 dB. Th delta is the difference between the normal threshold Th normal and the strict threshold Th tight , where the value of Th tight can be -1 dB. Additionally, f(R FFVD [n],R ISD [n]) can be max(R FFVD [n],R ISD [n]), where R FFFD [n] represents the output value from the flat friction voice detector 136, and R ISD [n] represents the output value from the plosive detector 132. That is, the max function is used to select the higher value. Although Equation 14 determines the maximum value of the outputs of the plosive detector 132 and the flat friction voice detector 136, in some embodiments, the sibilant detection module determines the maximum value of the outputs of any short-term feature detection.

[0078] In some embodiments, the function is more complex. For example, weights can be given to each output of the short-term detector 131 (e.g., alternatively or additionally to the fricative detector 136 and the plosive detector 132). If a particular output of the short-term feature detector 131 is related to speech and speech is detected in the portion of the audio signal being processed, a greater weight is given to that output. If a particular output of the short-term feature detector 131 is related to non-speech and speech is detected in the portion of the audio signal being processed, a smaller weight is given to that output. In some embodiments, f(R FFVD [n],R ISD [n] is more complex. For example, if the audio type is speech, more weight is given to the fricative detector 136, and if the content is non-speech (e.g., music, sound effects, or other suitable sounds), more weight is given to the plosive detector 132. In an embodiment, the fricative detection module 130 uses Equation 15 (below) to determine the value to be added to Equation 14:

[0079] f(R FFVD [n],R ISD [n]) = w FFVD [n]·R FFVD [n] + w ISD [n]·R ISD [n] Equation 15

[0080] where w FFVD [n] and w ISD [n] are the weights corresponding to the output of the fricative detector 136 and the output of the plosive detector 132, respectively. In some embodiments, the weights are determined based on the output from a content type classifier (e.g., a neural network). Although Equation 15 uses the weights of the outputs of the plosive detector 132 and the fricative detector 136, in some embodiments, the fricative detection module can assign / use weights for the output of any short-term feature detection. Thus, in some embodiments, Equation 15 can include the results from other short-term feature detectors with associated weights.

[0081] In some embodiments, when determining the threshold, the fricative detection module 130 uses the threshold to determine whether fricatives are present. In an embodiment, the fricative detection module 130 uses the logic of Equation 16 (below) to make the determination.

[0082]

[0083] where SPD[n] is the spectral balance feature and Th STSD [n] is the threshold determined by, for example, Equation 14.

[0084] In some embodiments, the sibilance detection module 130 transmits the results of the short-term sibilance detector 134 to the multi-band compressor 140. In some embodiments, the sibilance detection module 130 uses the results of the short-term sibilance detector 134 to perform long-term sibilance detection (e.g., by using the long-term sibilance detector 138). In some embodiments, long-term sibilance detection is performed on a longer portion of the audio signal (e.g., approximately 200 milliseconds). In some embodiments, the sibilance detection module 130 uses the reference Figure 5 actions described to further determine whether sibilance is present. These actions are only examples of long-term sibilance detection. In some embodiments, a classifier (e.g., a neural network) is used to perform long-term sibilance detection. For example, any detected short-term features and an appropriate portion of the audio signal can be used as inputs to the classifier (e.g., the classifier can be configured to receive short-term features and a portion of the audio signal), and the output of the classifier is a determination as to whether sibilance is present.

[0085] At 502, the sibilance detection module 130 accesses the output of the short-term sibilance detector 134. For example, the short-term sibilance detector 134 can be a function that outputs a value (e.g., one or zero) indicating whether sibilance is detected, and can also output the spectral balance feature discussed above. At 504, the sibilance detection module 130 selects a time constant based on whether the short-term sibilance detector 134 detects sibilance. In some embodiments, if sibilance is detected in the short-term sibilance detector 134, the constant is 0.2 seconds, and if sibilance is not detected in the short-term sibilance detector 134, the constant is one second.

[0086] At 504, the sibilance detection module 130 uses the selected time constant to calculate a smoothed version of the spectral balance feature.

[0087] In an embodiment, the sibilance detection module 130 uses the logic of Equation 17 for the calculation:

[0088]

[0089] where α S is the time constant used when the short-term sibilance detector 134 detects sibilance, i.e., R STSD [n]=1, and α ns is the time constant used when sibilance is not detected.

[0090] In some embodiments, the result of the non-sibilance smoothed spectral balance feature is given by Equation 18 (below):

[0091] R NSSSPD [n]=f(SPD smooth [n]) Equation 18

[0092] where f(·) is a comparison with a threshold. In an embodiment, the sibilance detection module 130 performs calculations using the logic of Equation 19 (below):

[0093]

[0094] where Th NSSSPD is the threshold (e.g., -12 dB).

[0095] In some embodiments, f(·) is a more complex function, as shown in Equation 20 (below):

[0096]

[0097] where Th NSSSPD1 and Th NSSSPD2 are thresholds (e.g., values of -15 dB and -12 dB, respectively) and SPD smooth [n] is a smoothed version of the spectral balance feature.

[0098] To continue Figure 5 process 500, the sibilance detection module 130 determines whether the smoothed version of the spectral balance feature meets the threshold. In some embodiments, as described with respect to Equation 20, the sibilance detection module 130 determines whether the spectral balance feature meets multiple thresholds. If the smoothed version of the spectral balance feature meets the threshold, process 500 moves to 510, where the sibilance detection module 130 determines that a sibilance exists. If the smoothed version of the spectral balance feature does not meet the threshold, process 500 moves to 512, where the sibilance detection module 130 determines that no sibilance exists.

[0099] In some embodiments, the output of the long-term sibilance detector 138 includes the results of both short-term sibilance detection and long-term sibilance detection. In some embodiments, the sibilance detection module 130 uses a function to determine the output of the long-term sibilance detector 138. In an embodiment, the output is as shown in Equation 21:

[0100] R LTSD [n] = f(R STSD [n], R NSSSPD [n]) Equation 21

[0101] where R STSD [n] and R NSSSPD [n] are the outputs from the short-term sibilance detector 134 and the long-term sibilance detector 138, respectively. For example, in Equation 21, f(·) is the product of R STSD [n] and R NSSSPD [n].[[]END]]

[0102] In some embodiments, the output of the short-term sibilance detection, the long-term sibilance detection, or both the short-term sibilance detection and the long-term sibilance detection is used for sibilance suppression. However, those skilled in the art will understand that sibilance suppression is just an example of using the detected sibilance. For example, the sibilance detection module 130 can use the output to control the multi-band compressor 140. Thus, the threshold of the multi-band compressor 140 is dynamically adjusted to suppress sibilance in the audio signal. In some embodiments, Equation 21 (below) is used for sibilance suppression:

[0103] Th k [n]=Th_static k +a k R LTSD [n] Equation 21

[0104] where k is in the sibilant frequency band of the multi-band compressor 140 (e.g., 4 kHz to 10 kHz), Th_static k is the static threshold of frequency band k, and a k is the dynamic adjustment value of frequency band k. In some embodiments, the dynamic adjustment is the same across all sibilant frequency bands. In some embodiments, for some sibilant frequency bands, the dynamic adjustment is different. The dynamic adjustment includes a preset value, an adjustable parameter, or other suitable dynamic adjustment. The adjustable parameter can be used to adapt to various characteristics of the device (e.g., a mobile device).

[0105] In some embodiments, the sibilance detection module 130 adjusts one or more parameters of the sibilance detector based on a combination of short-term features and long-term features. The sibilance detection module 130 determines one or more short-term features (e.g., plosives, flat fricatives, or other suitable features). The sibilance detection module 130 determines one or more long-term features based on the one or more short-term features. For example, the sibilance detection module 130 obtains the output of the short-term feature detector and uses the output as the input to the long-term feature detector, as described above. The sibilance detection module then adjusts one or more sibilance parameters based on the combination of the short-term features and the long-term features. For example, as described above, the sibilance detection module 130 changes the sibilance threshold based on the long-term sibilance features determined as using the output of the short-term sibilance features or the output of the transform module 110 and / or the banding module 120.

[0106] In some embodiments, the sibilant detection module uses a machine learning-based classifier (e.g., a neural network) to determine the presence of sibilants. In these embodiments, the sibilant detection module 130 uses a combination of any outputs of a short-term feature detector 131 (including an impulse sound detector 132, a flat friction voice detector 136, and / or any other short-term feature detector), a short-term sibilant detector 134, and a long-term sibilant detector 138 as inputs to the machine learning-based classifier. The machine learning-based classifier can be trained to output a determination regarding the presence of sibilants based on the information.

[0107] Figure 6 Illustrated is a sibilant suppression curve that can be used in sibilant suppression. The sibilant suppression curve includes three parts C1, C2, and C3. In part C1, the sibilant level is below the low threshold TH_low, and thus the attenuation gain of sibilant suppression will be 0 dB, which means that no suppression processing will be performed on non-sibilant sounds and non-sibilant sounds. In part C2, the sibilant level falls within the range between TH_Low and TH_high, and thus linear suppression can be triggered. In part C3, the sibilant level is above the high threshold TH_high, and the attenuation gain of sibilant suppression is set to G1, which is the maximum sibilant suppression depth of the system.

[0108] Figure 7 A block diagram of an example system 700 suitable for implementing example embodiments of the present disclosure is shown. As shown, the system 700 includes a central processing unit (CPU) 701 that can execute various processes according to a program stored in, for example, a read-only memory (ROM) 702, or a program loaded into a random access memory (RAM) 703 from, for example, a storage unit 708. In the RAM 703, data required for the CPU 701 to execute various processes is also stored as needed. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0109] The following components are connected to the I / O interface 705: an input unit 706, which may include a keyboard, a mouse, etc.; an output unit 707, which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 708, which includes a hard disk or other suitable storage device; and a communication unit 709, which includes a network interface card such as a network card (e.g., wired or wireless). The communication unit 709 is configured to communicate with other devices (e.g., via a network). Optionally, a driver 710 is also connected to the I / O interface 705. Optionally, a removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or other suitable removable medium is installed on the driver 710 so that a computer program read therefrom is installed into the storage unit 708. Those skilled in the art will understand that although the system 700 is described as including the components described above, in practical applications, some of these components may be added, removed, and / or replaced, and all such modifications or changes fall within the scope of the present disclosure.

[0110] According to an exemplary embodiment of the present disclosure, the processes described above may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing a method. In such an embodiment, the computer program may be downloaded and installed from a network via the communication unit 709, and / or installed from the removable medium 711.

[0111] Generally, the various exemplary embodiments of the present disclosure may be implemented in hardware or a dedicated circuit (e.g., a control circuit), software, logic, or any combination thereof. For example, the sibilance detection module 130 may be executed by a control circuit (e.g., a CPU combined with Figure 7 other components of), and thus, the control circuit may perform the actions described in the present disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, a microprocessor, or other computing devices (e.g., a control circuit). Although the various aspects of the exemplary embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, devices, systems, techniques, or methods described herein, as non-limiting examples, may be implemented in hardware, software, firmware, a dedicated circuit or logic, general hardware or a controller, or other computing devices, or some combination thereof.

[0112] Additionally, each block shown in the flowchart can be regarded as a method step, and / or an operation resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform the associated (multiple) functions. For example, embodiments of the present disclosure include a computer program product that includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code configured to perform the methods described above.

[0113] In the context of the present disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0114] The computer program code for performing the methods of the present disclosure can be written in any combination of one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device having control circuitry, such that the program code, when executed by the processor of the computer or other programmable data processing device, implements the functions / operations specified in the flowchart and / or block diagram. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

Claims

1. A method, comprising: Receiving an audio signal; Extracting a plurality of time-frequency features from the audio signal, the plurality of time-frequency features including one or more short-term features; Adjusting one or more thresholds of a sibilance detector for detecting sibilance in the audio signal according to the extracted short-term features; and Using the sibilance detector with one or more adjusted thresholds to detect sibilance in the audio signal; Wherein, detecting plosives from the one or more short-term features; and Wherein, performing the adjustment of the one or more thresholds of the sibilance detector for detecting sibilance in the audio signal according to an output value obtained by determining whether the plosives are detected.

2. The method according to claim 1, wherein, Detecting flat fricatives from the one or more short-term features; The flat fricative is a fricative with a flat spectrum; And The method further comprises: adjusting one or more thresholds of the sibilance detector for detecting sibilance in the audio signal according to an output value obtained by determining whether the flat fricatives are detected.

3. The method according to claim 1 or 2, wherein, The plurality of time-frequency features include long-term features; And The method further comprises: adjusting one or more thresholds of the sibilance detector for detecting sibilance in the audio signal according to the extracted long-term features.

4. The method according to claim 3, wherein, The long-term features include smoothed audio spectrum balance features.

5. The method according to claim 1 or 2, wherein, Adjusting the one or more thresholds of the sibilance detector includes generating a control signal, the control signal including a value generated by the detection of short-term features.

6. The method according to claim 1 or 2, wherein Adjusting the one or more thresholds of the sibilance detector includes: Determining the one or more short-term features; Determining one or more long-term features included in the plurality of time-frequency features; and Adjusting the one or more thresholds based on a combination of the one or more short-term features and the one or more long-term features.

7. The method according to claim 1 or 2, wherein Detecting the plosives from the one or more short-term features includes: For a first time interval in the audio signal, calculating a first total power in one or more sibilant frequency bands and a second total power in one or more non-sibilant frequency bands; For a second time interval in the audio signal, calculating a third total power in one or more sibilant frequency bands and a fourth total power in one or more non-sibilant frequency bands; Determining a first flux value based on a difference between the first total power and the third total power, and determining a second flux value based on a difference between the second total power and the fourth total power; and Determining whether there is the plosive based on whether the first flux value satisfies a first threshold and whether the second flux value satisfies a second threshold.

8. The method according to claim 7, further comprising, in response to determining that there is the plosive: Generating an output value; and Applying a smoothing algorithm to the output value.

9. The method according to claim 8, wherein Applying the smoothing algorithm to the output value includes using an attack time constant and a release time constant.

10. The method according to claim 9, further comprising adjusting the attack time constant or the release time constant based on the type of the plosive.

11. The method according to claim 1 or 2, further comprising determining the type of the plosive based on the plurality of time-frequency features.

12. The method according to claim 11, wherein, Determining the type of impact sound includes: Comparing the data of one or more dental frequency bands and one or more non-dental frequency bands with the corresponding frequency band data of multiple known impact sounds; and Identifying the impact sound based on the comparison.

13. The method according to claim 2, wherein Using the one or more short-term features to determine whether the audio signal includes the fricative sound includes: Calculating a dental spectral flatness metric based on the dental speech frequency band spectrum and the number of frequency bands.

14. The method according to claim 2, wherein Using the one or more short-term features to determine whether the audio signal includes the fricative sound includes: Calculating the variance of the power of adjacent dental frequency bands.

15. The method according to claim 2, wherein, Using the one or more short-term features to determine whether the audio signal includes the fricative sound includes: Calculating the peak-mean ratio or peak-median ratio of the power in the dental frequency band.

16. The method according to claim 2, wherein, Using the one or more short-term features to determine whether the audio signal includes the fricative sound includes: Calculating a spectral entropy metric in the dental frequency band.

17. The method according to claim 2, wherein Adjusting one or more thresholds of a dental sound detector for detecting dental sounds in the audio signal includes: adjusting the dental sound detection threshold based on the output value generated by determining whether the impact sound is detected and the output value generated by determining whether the fricative sound is detected.

18. The method according to claim 17, wherein, Adjusting one or more thresholds of the dental sound detector includes: Determining whether the current portion of the audio signal includes speech; In response to determining that the current portion of the audio signal includes speech, adding a first weight to the output value generated by determining whether the impact sound is detected, and adding a second weight higher than the first weight to the output value generated by determining whether the fricative sound is detected; and In response to determining that the current portion of the audio signal includes non-speech, adding a first weight to the output value generated by determining whether the impact sound is detected, and adding a second weight lower than the first weight to the output value generated by determining whether the fricative sound is detected.

19. The method according to claim 1 or 2, further comprising: Accessing the output of the dental sound detector and the spectral balance value; Selecting a time constant based on whether the dental sound detector detects a dental sound; Using the selected time constant to calculate a smoothed version of the spectral balance value; Comparing the smoothed version of the spectral balance with a threshold; Determining whether there is a dental sound based on comparing the smoothed version of the spectral balance with the threshold.

20. A system, comprising: One or more computer processors; And One or more non-transitory storage media storing instructions that, when executed by the one or more computer processors, cause the execution of the method according to any one of claims 1 to 19.

21. A computer program product, comprising a computer program that, when executed by one or more processors, causes the execution of the method according to any one of claims 1 to 19.

Citation Information

Patent Citations

  • Sibilance detection and mitigation

    EP3261089A1

  • Method for abnormal sound source detection and apparatus for performing same

    WO2018097620A1