Audio processing method and device
The audio processing method efficiently reduces self-noise in PSAPs and hearing-aids by dividing signals into frequency bands, estimating noise, and applying a smoothed gain, addressing the issue of degraded sound quality and latency in existing technologies.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-02
AI Technical Summary
Personal sound amplification products (PSAPs) and hearing-aids amplify self-noise, degrading sound quality and disrupting the user's hearing experience, especially when large gains are applied, and existing noise reduction methods introduce latency or require high computation.
An audio processing method that divides audio signals into frequency bands, estimates self-noise, calculates signal-to-self-noise ratio (SSNR) and noise reduction gain, and applies a smoothed gain to reduce self-noise in the time domain with minimal latency and light computation.
Effectively reduces self-noise without distorting audio signals, providing a better hearing experience with no additional latency and low computational load, suitable for PSAPs and hearing-aids.
Smart Images

Figure CN2024121698_02042026_PF_FP_ABST
Abstract
Description
AUDIO PROCESSING METHOD AND DEVICETECHNICAL FIELD
[0001] The present disclosure relates to a field of audio processing, and in particular, to an audio processing method, an audio processing apparatus and device, a computer-readable storage medium, and a computer program product.BACKGROUND
[0002] With the aging of the global population, more and more people are suffering from hearing loss. A personal sound amplification product (PSAP) or a hearing-aid aims to provide a better auditory experience or compensate for hearing loss of a user. Specifically, a PSAP or a hearing-aid can amplify acoustic signals received by an in-air microphone or output by a microphone array beamformer, then playback the amplified signals via a speaker plugged into or close to the user’s ear.
[0003] There is always some noise generated by the microphone itself even without any environmental sound present, which may be referred to as self-noise or noise floor. A spectrum of the self-noise is often nearly white and normally not too unpleasant to hear for a user. However, PSAPs or hearing-aids with very large gains may make the self-noise beyond a certain loudness. In these cases, the self-noise may degrade sound quality, especially its cleanness and clarity, and thus disrupt the user’s hearing experience. As PSAP functionality becomes increasingly popular in headphone products, such as true wireless stereo (TWS) headphones, it is necessary and highly beneficial to develop an approach to reduce the self-noise sent to a speaker in PSAP or hearing-aid products.
[0004] SUMMARY OF THE DISCLOSURE
[0005] The present disclosure proposes an audio processing method, an audio processing apparatus and device, a computer-readable storage medium, and a computer program product, which can efficiently reduce self-noise in the time domain without introducing additional significant latency during the processing and with only light computation load.
[0006] According to one or more aspects of the present disclosure, there is provided an audio processing method, comprising: acquiring an audio signal; dividing the audio signal into a plurality of audio signal components in a plurality of frequency bands; estimating self-noise in each of the plurality of audio signal components of the audio signal; calculating a signal to self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components; smoothing the noise reduction gain, based at least on the SSNR and the noise reduction gain, to generate a smoothed gain; and applying the smoothed gain to the audio signal to generate a processed audio signal with the self-noise reduced.
[0007] According to one or more aspects of the present disclosure, there is provided an audio processing apparatus, comprising: an audio acquisition unit configured to acquire an audio signal; a frequency dividing unit configured to divide the audio signal into a plurality of audio signal components in a plurality of frequency bands; a self-noise estimation unit configured to estimate self-noise in each of the plurality of audio signal components of the audio signal; a gain calculation unit configured to calculate a signal to self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components; a smoothing unit configured to smooth the noise reduction gain, based at least on the SSNR and the noise reduction gain, to generate a smoothed gain; and a noise reduction unit configured to apply the smoothed gain to the audio signal to generate a processed audio signal with the self-noise reduced.
[0008] According to another aspect of the present disclosure, there is provided an audio processing device, comprising: at least one microphone configured to capture an audio signal; at least one speaker configured to output a processed audio signal; and one or more processors coupled to the at least one microphone and the at least one speaker, configured to: divide the audio signal into a plurality of audio signal components in a plurality of frequency bands; estimate self-noise in each of the plurality of audio signal components of the audio signal; calculate a signal to self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise; smooth the noise reduction gain, based at least on the SSNR and the noise reduction gain of the audio signal, to generate a smoothed gain of the audio signal; and apply the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced.
[0009] According to one or more aspects of the present disclosure, there is provided an audio processing device, comprising: one or more processors; and one or more memories having stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to perform the method described in the aforementioned aspects.
[0010] According to one or more aspects of the present disclosure, there is provided a computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to perform the method described in the aforementioned aspects.
[0011] According to one or more aspects of the present disclosure, there is provided a computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method described in the aforementioned aspects.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other objects, features and advantages of embodiments of the present disclosure will become obvious from the following detailed description of embodiments of the present disclosure taken in conjunction with accompanying drawings. The accompanying drawings are used to provide further understanding of the embodiments of the present disclosure, constitute a part of the specification, explain the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation of the present disclosure. In the drawings, like reference numerals generally represent like components or steps.
[0013] FIG. 1 illustrates a flow diagram of an audio processing method in accordance with one or more embodiments of the present disclosure;
[0014] FIG. 2 illustrates an example process flow of the audio processing method in accordance with one or more embodiments of the present disclosure;
[0015] FIG. 3 illustrates an example of sound captured by a microphone and its spectrum in accordance with one or more embodiments of the present disclosure;
[0016] FIG. 4 illustrates the sound of FIG. 3 amplified by a fixed gain and its spectrum in accordance with one or more embodiments of the present disclosure;
[0017] FIG. 5 illustrates the sound of FIG. 3 amplified by a PSAP device without utilizing the proposed audio processing method and its spectrum in accordance with one or more embodiments of the present disclosure;
[0018] FIG. 6 illustrates the sound of FIG. 3 amplified by a PSAP device utilizing the proposed audio processing method and its spectrum in accordance with one or more embodiments of the present disclosure;
[0019] FIG. 7 illustrates a schematic structural diagram of an audio processing apparatus in accordance with one or more embodiments of the disclosure; and
[0020] FIG. 8 illustrates a schematic diagram of an architecture of an exemplary audio device in accordance with one or more embodiments of the present disclosure.
[0021] DESCRIPTION OF THE EMBODIMENTS
[0022] In order to make objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and thoroughly with reference to the accompanying drawings. Obviously, these described embodiments are only a part of the present disclosure, not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without paying creative efforts fall into the protection scope of the present disclosure.
[0023] As used herein and in the claims, the words “a, ” “an, ” “an, ” and / or “the” do not refer to the singular, but may include the plural unless the context clearly dictates otherwise. In general, the terms “comprise” and “comprising” only imply the inclusion of steps and elements specifically identified, these steps and elements do not constitute an exclusive list and a method or apparatus may also contain other steps or elements.
[0024] Flowcharts are used herein to illustrate steps of a method according to one or more embodiments of the present disclosure. It should be understood that preceding or subsequent steps do not have to be performed exactly in order. Rather, various steps may be processed in reverse order or simultaneously, as desired. Meanwhile, other steps may also be added to the method, or certain step or steps may be removed from the method.
[0025] As used herein and in the claims, audio signals and self-noises may refer to the signals themselves or statistical measures such as power, amplitudes, root mean square values, and the like of the signals. In other words, in the present disclosure, audio signals and statistical measures of the audio signals sometimes may be used interchangeably, and self-noises and statistical measures of the self-noises sometimes may be used interchangeably, unless otherwise stated explicitly.
[0026] In sounds captured by a microphone, there are always self-noises due to circuits or even thermal noises of the microphone. Self-noise is normally white noise and near to flat in a relatively wide band in the spectrum, and thus usually not too unpleasant to hear. However, when a sound captured by the microphone is amplified without any pre-processing, for example, in a PSAP or hearing-aid product, the self-noise will also be amplified. Loud self-noise may interrupt a user’s hearing experience and become annoying. A sound with a higher microphone signal-to-self-noise ratio (SSNR) , that is, having a given loudness but with less self-noise, is clearer and more comfortable to hear.
[0027] Some audio products choose microphones with a high signal-to-noise ratio (SNR) to provide high-quality sounds while applying small gains. In this way, microphones with high SNR can perform better in terms of self-noise, at least to reach the same loudness for a sound, because self-noise will not be amplified too much by small gains. However, high-SNR microphones are expensive, and for a large amplification gain, self-noise would still be audible and cause an unpleasant auditory experience.
[0028] Some audio products apply a noise gate method to purify sounds captured by microphones. Noise gate is a versatile tool used in audio production to eliminate unwanted sounds below a predetermined threshold and allow wanted sounds above the threshold to pass through. The noise gate method can attenuate noise in a quiet environment, but when other sounds are present, it is difficult to separate self-noise from the sounds.
[0029] Some audio products adopt an environmental noise cancellation (ENC) module to reduce self-noise and background noise together. However, ENC is normally frame-based spectral processing and requires overlapped adding / saving operations in the frequency domain to avoid artifacts in junctions between frames of audio signals, which would introduce additional latency in the process. However, latency is a key parameter to be controlled in audio devices to achieve transparency and clarity requirements of audio signals. Besides, ENC usually requires a large amount of computation, thereby imposing high demands on device performance.
[0030] To solve the above problems, the present disclosure proposes an audio processing method that can efficiently reduce or substantially remove self-noise in the time domain without introducing additional latency and with only a light computation load.
[0031] FIG. 1 illustrates a flow diagram of an audio processing method 100 in accordance with one or more embodiments of the present disclosure. The audio processing method 100 may be performed by stand-alone audio devices such as headphones, headsets, PSAPs, hearing-aid devices, or any other audio devices having PSAP or similar functionality, or may be performed by devices incorporating audio processing capability such as smartphones, tablets, desktop computers, smart wearable devices, televisions, and the like, which is not specifically limited by the embodiments of the present disclosure.
[0032] As shown in FIG. 1, in step S102, an audio signal is acquired. In one or more embodiments of the present disclosure, the audio signal is a signal to be processed, which is acquired by an audio acquisition device such as a microphone, or acquired by the microphone and further processed by other audio processing modules such as a microphone array beamformer. The microphone used to acquire audio signals may be, for example, a microphone in a personal sound amplification product (PSAP) or a hearing-aid, which is not specifically limited by the embodiments of the present disclosure. For example, a k-th frame of audio signal may be denoted as x (k) , where k is an integer greater than or equal to 1. A specific origin of the audio signal to be processed is not specifically limited by the embodiments of the present disclosure.
[0033] In step S104, the audio signal may be divided into a plurality of audio signal components in a plurality of frequency bands, for example, by using some frequency band division methods, to facilitate subsequent audio processing. For example, a crossover network or a filter bank may be used to perform the band dividing. These processing units are often existing blocks in PSAP or hearing-aid devices. It should be noted that overlap is allowed between these bands, and centers or crossover frequencies and widths of the bands may be variable, which are not specifically limited by the present disclosure. Specifically, a current frame of audio signal may be referred to as a k-th frame of audio signal, x (k) , and may be divided into L frequency bands, x (l, k) , l=1, 2, …L, where L is an integer greater than 1 and l denotes an index of the l-th frequency band. Accordingly, an audio signal component in the l-th frequency band may be referred to as the l-th audio signal component.
[0034] In step S106, self-noise in each of the plurality of audio signal components of the audio signal is estimated. Usually, self-noise of a microphone will be measured in the product developing stage in an experimental environment, and then the self-noise among different mockup samples should be similar, assuming proper quality control in the factory. However, some precision error in manufacturing is allowed, also the microphone self-noise may change during consumer use, for example, due to subtle changes in the structure of the microphone. The audio processing method 100 of the present disclosure may estimate and update the self-noise of the microphone, thereby providing a better noise reduction effect. For example, a minimum tracking method may be used to estimate and update the self-noise. Minimum tracking is a noise estimation algorithm based on short-term stationarity of audio signals and randomness of noise, which estimates a noise level of a frequency band by tracking minimum power in the frequency band. The minimum tracking method mainly includes minimum statistics, minimum search, minimum tracking of continuous spectrums, and the like.
[0035] Specifically, for each audio signal component of the plurality of audio signal components, minimum power of the audio signal component in its corresponding frequency band may be estimated, and then the self-noise of the audio signal component may be determined based on the minimum power. For example, based on the minimum power or a statistical value of the minimum power, an amplitude, power, sound pressure, sound level or the like of the self-noise may be determined. In particular, the minimum power or the statistical value of the minimum power may be determined as the power of the self-noise in the audio signal component. The determined self-noise may be denoted as n (l, k) .
[0036] In step S108, SSNR and a noise reduction gain of each of the plurality of audio signal components may be calculated based on the audio signal and the self-noise. In the present disclosure, each frame of audio signal may include a plurality of sampling time points, for example, M sampling time points where M is an integer greater than 1. For a particular audio signal component in the l-th band of the current k-th frame of audio signal, SSNR may be calculated by using a statistical value of signal or self-noise in the band. For example, SSNR may be calculated as:
[0037] where r (k, j) is SSNR of the l-th audio signal component of the k-th frame of audio signal; x (l, k) is a statistical value of the l-th audio signal component of the k-th frame of audio signal; and n (l, k) is a statistical value of the self-noise in the l-th audio signal component of the k-th frame of audio signal. The statistical value may be, for example, a sum, an average value, a peak value, a root mean square value, or the like of power of the audio signal component or the self-noise within the l-th frequency band.
[0038] For the plurality of sampling time points contained in each audio signal component of the plurality of audio signal components, a statistical value of signal power and a statistical value of self-noise power of the audio signal component at the plurality of sampling time points may be determined, and then the SSNR and the noise reduction gain of the audio signal component may be calculated based at least on the statistical value of the signal power and the statistical value of the self-noise power. For example, the statistical value may be a sum, an average value, a peak value, a root mean square value, or the like of power at the plurality of sampling time points. Hereinafter, the statistical value being the sum of power at the plurality of sampling time points may be taken as an example for description, which should not be considered as limiting the present disclosure.
[0039] Considering that there will be power fluctuation and a possible over-estimation issue due to statistical reasons, Equation (1) may be modified as:
[0040] where R (l, k) is SSNR of the l-th audio signal component of the k-th frame of audio signal, and γ (l) is an over-estimation coefficient for the l-th audio signal component in the l-th frequency band. where x (l, k) m is a power of the l-th audio signal component of the k-th frame of audio signal at the sampling time point m; and n (l, j) m is a power of the self-noise in the l-th audio signal component of the k-th frame of audio signal at the sampling time point m.
[0041] The over-estimation coefficient for each audio signal component may be determined based on a frequency distribution of the audio signal component. For example, the value of the over-estimation coefficient may be different in different frequency bands, for example, may be larger in a frequency band with lower frequencies and may be smaller in a frequency band with higher frequencies. Specifically, the over-estimation coefficient may be determined depending on a frequency range to which human ears are sensitive, or determined according to empirical parameters, which is not specifically limited by the embodiments of the present disclosure.
[0042] Similarly, the noise reduction gain for the l-th audio signal component of the k-th frame of audio signal may be calculated as:
[0043] where g (l, k) is the noise reduction gain for the l-th audio signal component of the k-th frame of audio signal.
[0044] In step S110, the noise reduction gain may be smoothed based at least on the SSNR and the noise reduction gain of the audio signal to generate a smoothed gain of the audio signal. The smoothing operation is used to reduce fluctuation among frames and among frequency bands, thereby providing a processed audio signal with a better hearing experience.
[0045] In one or more embodiments of the present disclosure, the smoothing operation may include, for each audio signal component of the plurality of audio signal components, performing, based at least on the SSNR and the noise reduction gain of the audio signal component and between a previous frame of audio signal component of the audio signal component and the current audio signal component, a first smoothing operation on the noise reduction gain of the audio signal component, to generate a first smoothed gain of the audio signal component.
[0046] Specifically, the first smoothing operation may include: determining a first weight coefficient for the audio signal component based at least on the SSNR and the noise reduction gain of the audio signal component; and performing, by using the first weight coefficient for the audio signal component, a weighted summation of a first smoothed gain of the previous frame of audio signal component in a corresponding frequency band and the noise reduction gain of the audio signal component, to generate the first smoothed gain of the audio signal component. For example, the first smoothed gain of the l-th audio signal component of the k-th frame of audio signal may be calculated as: G1 (l, k) =αG1 (l, k-1) + (1-α) g (l, k) , (4)
[0047] where G1 (l, j) is the first smoothed gain of the l-th audio signal component of the k-th frame of audio signal; G1 (l, k-1) is the first smoothed gain of the l-th audio signal component of the (k-1) -th frame of audio signal, and an initial value of G1 (l, 0) may be any value in a range of (0, 1] , for example, 1; g (l, k) is the noise reduction gain of the l-th audio signal component of the k-th frame of audio signal; and α is the first weight coefficient for the l-th audio signal component of the k-th frame of audio signal, where 0<α <1.
[0048] The value of the first weight coefficient for each audio signal component may be determined by SSNR for each audio signal component, a weighted sum of SSNRs for different audio signal components, and the noise reduction gain for the audio signal component. Specifically, for each audio signal component of the plurality of audio signal components, a weighted sum of SSNRs of the plurality of audio signal components may calculated, and then the first weight coefficient for the audio signal component may be determined based on the weighted sum of the SSNRs and the SSNR and the noise reduction gain of the audio signal component. The first weight coefficient for the l-th audio signal component of the k-th frame of audio signal may be denoted as:
[0049] where R (l, k) is SSNR of the l-th audio signal component of the k-th frame of audio signal; is the weighted sum of the SSNRs of the plurality of audio signal components of the k-th frame of audio signal in L frequency bands, and herein the sum of weights for each frequency band is one; g (l, k) is the noise reduction gain of the l-th audio signal component of the k-th frame of audio signal; and the operator f (·) represents a function of the SSNR, the weighted sum of SSNRs and the noise reduction gain, which may be determined, for example, according to practical requirements.
[0050] Then, the first smoothed gain of each audio signal component of the audio signal may be applied to the audio signal to generate a processed audio signal with the self-noise reduced or removed in step S112. The first smoothing operation can reduce fluctuation among frames of audio signals, thereby providing a processed audio signal with a better hearing experience.
[0051] In one or more embodiments of the present disclosure, the smoothing operation may further include, for each audio signal component of the plurality of audio signal components, performing, based at least on the SSNR and the noise reduction gain of the audio signal component and among a frequency band corresponding to the audio signal component, a first adjacent frequency band with a smaller central frequency than the frequency band and a second adjacent frequency band with a larger central frequency than the frequency band, a second smoothing operation on the first smoothed gain of the audio signal component, to generate a second smoothed gain of the audio signal component.
[0052] Specifically, the second smoothing operation may include performing, by using a plurality of second weight coefficients, a weighted summation of the first smoothed gain of the audio signal component, a first smoothed gain of a first adjacent audio signal component in the first adjacent frequency band, and a first smoothed gain of a second adjacent audio signal component in the second adjacent frequency band, to generate the second smoothed gain of the audio signal component. For example, the second smoothed gain of the l-th audio signal component of the k-th frame of audio signal may be calculated as: G2 (l, k) =β1G1 (l-1, k) +β2H1 (l, k) +β3G1 (l+1, k) , (6)
[0053] where G2 (l, k) is the second smoothed gain of the l-th audio signal component of the k-th frame of audio signal; G1 (l-1, k) is the first smoothed gain of the (l-1) -th audio signal component of the k-th frame of audio signal, and an initial value of G1 (0, k) may be any value in the range of (0, 1] , for example, 1; G1 (l, k) is the first smoothed gain of the l-th audio signal component of the k-th frame of audio signal; G1 (l+1, k) is the first smoothed gain of the (l+1) -th audio signal component of the k-th frame of audio signal; β1, β2, and β3 are the second weight coefficients for the respective first smoothed gains, and β1+β2+β3=1.
[0054] The second weight coefficients may be determined, for example, according to practical requirements, which are not specifically limited by the embodiments of the present disclosure. For example, among the plurality of second weight coefficients, a second weight coefficient for the audio signal component may be greater than that for the first adjacent audio signal component in the first adjacent band with a smaller central frequency and the second adjacent audio signal component in the second adjacent band with a larger central frequency, to ensure that the first smoothed gain for the current audio signal component contributes the most to the final second smoothed gain.
[0055] Then, the second smoothed gain of each audio signal component of the audio signal may be applied to the audio signal to generate a processed audio signal with the self-noise reduced or removed in step S112. The second smoothing operation can reduce fluctuation among frequency bands of audio signals, thereby providing a processed audio signal with a better hearing experience.
[0056] In one or more embodiments of the present disclosure, only the first smoothing operation may be performed for the noise reduction gain, and the first smoothed gain may be the final noise reduction gain to be applied to the audio signal in step S112. In one or more embodiments of the present disclosure, both the first and second smoothing operations may be performed for the noise reduction gain, and the second smoothed gain may be the final noise reduction gain to be applied to the audio signal in step S112.
[0057] In step S112, the smoothed gain, such as the first smoothed gain or the second smoothed gain, may be applied to the audio signal to generate a processed audio signal with the self-noise reduced or removed. In one or more embodiments of the present disclosure, the smoothed gain may be multiplied directly with the audio signal to generate the processed audio signal. For example, the processed audio signal component of the l-th audio signal component of the k-th frame of audio signal may be calculated as: s (l, k) = G (l, k) ·x (l, k) , (7)
[0058] where s (l, k) is the processed audio signal component of the l-th audio signal component of the k-th frame of audio signal; G (l, k) is the smoothed gain of the l-th audio signal component of the k-th frame of audio signal; and x (l, k) is the l-th audio signal component of the k-th frame of audio signal.
[0059] In one or more embodiments of the present disclosure, the smoothed gain, such as the first smoothed gain or the second smoothed gain, may be applied to a wide dynamic range compression (WDRC) processing module to generate the processed audio signal. WDRC is a technology that dynamically adjusts a wide range of sound levels to better adapt to a user’s hearing level and environmental noise levels. In this case, the processed audio signal component of the l-th audio signal component of the k-th frame of audio signal may be obtained as: s (l, k) =WDRC (G (l, k) , x (l, k) ) , (8)
[0060] Then, the processed audio signal component in different frequency bands may be composited to obtain the processed audio signal of the k-th frame as:
[0061] where y (k) is the processed audio signal of the k-th frame, and s (l, k) is the processed audio signal component of the l-th audio signal component of the k-th frame of audio signal.
[0062] Specific steps of the audio processing method 100 to reduce self-noise from the audio signal are described above and denoted in Equations (1) ~ (8) . In order to provide a better understanding of the principle of the audio processing method 100, FIG. 2 is provided to illustrate an example process flow of the audio processing method 100.
[0063] As shown in FIG. 2, an audio signal to be processed, for example, an audio signal acquired by a microphone, or acquired by the microphone and further processed by other audio processing modules, is first input to a frequency divider 202. The frequency divider 202 may divide the audio signal into a plurality of audio signal components in a plurality of frequency bands, for example, by using a crossover network or a filter bank. Then, self-noise in each of the plurality of audio signal components of the audio signal is estimated by a self-noise estimation module 204. An SSNR and gain calculation module 206 then calculates SSNR and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components. A smoother 208 may perform a first smoothing operation and a second smoothing operation to smooth the noise reduction gain among frames and among frequency bands. The smoother 208 may consider SNR, SSNR, differences between new and old gains, as well as differences between gains of adjacent frequency bands, thereby providing an adaptive smoothing speed and noise reduction gain. At last, the smoothed gain may be input to a WDRC module 210, which will apply the smoothed gain to the input audio signals to generate a processed audio signal from which the self-noise is reduced or substantially removed.
[0064] With the audio processing method 100, self-noises in audio signals captured by audio acquisition devices, such as microphones, may be effectively reduced or substantially removed without distorting the audio signals. It is also noted that the audio processing method 100 of the present disclosure is performed in the time domain rather than the conventional frequency domain noise reduction, which will result in continuous noise reduction and no additional latency will be introduced. Furthermore, the computation load of the audio processing method 100 is light and can be effectively performed in audio devices such as headphones, earphones, hearing-aids, smartphones, public address systems, and the like. The audio processing method 100 of the present disclosure is particularly suitable for reducing self-noises in PSAPs and hearing-aids where large amplification is needed.
[0065] The self-noise reduction effect of the audio processing method 100 may be obvious from the comparison of FIGS. 3 to 6. FIG. 3 illustrates an example of sound captured by a microphone and its spectrum in accordance with one or more embodiments of the present disclosure. FIG. 4 illustrates the sound of FIG. 3 amplified by a fixed gain and its spectrum in accordance with one or more embodiments of the present disclosure. FIG. 5 illustrates the sound of FIG. 3 amplified by a PSAP device without utilizing the proposed audio processing method 100 and its spectrum in accordance with one or more embodiments of the present disclosure. FIG. 6 illustrates the sound of FIG. 3 amplified by a PSAP device utilizing the proposed audio processing method 100 and its spectrum in accordance with one or more embodiments of the present disclosure.
[0066] In FIGS. 3 to 6, the upper figure shows the waves of the sound, in which the horizontal axis represents sampling time points (i.e., samples) and the vertical axis represents an amplitude of the sound as a digitized value; the lower figure shows the spectrum of the sound, in which the horizontal axis represents sampling time points, the left vertical axis represents normalized frequencies, the right vertical axis represents power of the sound, and the darker the color of the spectrum, the stronger the sound intensity.
[0067] In the upper figure of FIG. 3, there are nearly flat sound waves near the horizontal axis in the middle of the sound waves, which represent self-noise of the microphone recorded in a quiet environment, and the two larger sound waves in the left and right sides are recorded speech signals. Correspondingly, in the lower figure of FIG. 3, the self-noise spreads over the whole frequency band and appears like snowflakes.
[0068] In FIG. 4, amplification of the sound by a fixed gain (e.g., tenfold) makes the self-noise larger, which becomes darker in the spectrum. In FIG. 5, the sound is amplified by a PSAP device, for example, some around tenfold but different gains in different frequency bands, to compare with FIG. 3. It can be seen that the sound waves corresponding to the quiet moment are in small values, meaning that the self-noise is not amplified as much as the signal of speech, although the proposed audio processing method 100 has not been utilized. This is because other modules in the PSAP device may also help reduce SNR.
[0069] Then in FIG. 6, the sound is amplified by a PSAP device utilizing the proposed audio processing method 100, for example, also some around tenfold but different gains in different frequency bands, exactly the same as that in FIG. 5. From the spectrum of FIG. 6, it is obvious that the self-noise is significantly reduced compared with the spectrum of FIG. 5. This fully demonstrates the effectiveness of the audio processing method 100 proposed by the present disclosure in reducing self-noises of microphones.
[0070] An audio processing apparatus according to one or more embodiments of the present disclosure will be described below with reference to FIG. 7. FIG. 7 illustrates a schematic structural diagram of an audio processing apparatus 700 in accordance with one or more embodiments of the disclosure. As shown in FIG. 7, the audio processing apparatus 700 may include an audio acquisition unit 702, a frequency dividing unit 704, a self-noise estimation unit 706, a gain calculation unit 708, a smoothing unit 710 and a noise reduction unit 712. In addition to these units, the audio processing apparatus 700 may further include other components, but since these components are not relevant to the present disclosure, a detailed description thereof is omitted herein. In addition, since details of a part of the functions of the audio processing apparatus 700 are similar to details of the steps of the audio processing method 100 as described with reference to FIG. 1, repeated descriptions of some content are omitted herein for brevity. The audio processing apparatus 700 may be a stand-alone audio device such as a headphone, a headset, a PSAP, a hearing-aid device, or any other audio device having PSAP or similar functionality, or may be a device incorporating audio processing capability such as a smartphone, a tablet, a desktop computer, a smart wearable device, a television, and the like, which is not specifically limited by the embodiments of the present disclosure.
[0071] The audio acquisition unit 702 may be configured to acquire an audio signal. The audio acquisition unit 702 may be, for example, a microphone in a personal sound amplification product (PSAP) or a hearing-aid, which is not specifically limited by the embodiments of the present disclosure. The audio signal is a signal to be processed, which may be acquired by the microphone, or acquired by the microphone and further processed by other audio processing modules such as a microphone array beamformer.
[0072] The frequency dividing unit 704 may be configured to divide the audio signal into a plurality of audio signal components in a plurality of frequency bands, for example, by using some frequency band division tools, to facilitate subsequent audio processing. For example, the frequency dividing unit 704 may include a crossover network or a filter bank to perform the band dividing. It should be noted that overlap is allowed between these bands and central frequencies and widths of the bands may be variable, which is not specifically limited by the present disclosure.
[0073] The self-noise estimation unit 706 may be configured to estimate self-noise in each of the plurality of audio signal components of the audio signal. For example, the self- noise estimation unit 706 may use a minimum tracking method to estimate and update the self-noise. Specifically, for each audio signal component of the plurality of audio signal components, the self-noise estimation unit 706 may estimate minimum power of the audio signal component in its corresponding frequency band, and then determine the self-noise of the audio signal component based on the minimum power. For example, based on the minimum power or a statistical value of the minimum power, the self-noise estimation unit 706 may determine an amplitude, power, sound pressure, sound level, or the like of the self-noise. In particular, the minimum power or the statistical value of the minimum power may be determined as the power of the self-noise in the audio signal component.
[0074] The gain calculation unit 708 may be configured to calculate SSNR and a noise reduction gain of each of the plurality of audio signal components based on the audio signal and the self-noise. In the present disclosure, each frame of audio signal may include a plurality of sampling time points, and the gain calculation unit 708 may determine a statistical value of signal power and a statistical value of self-noise power of the audio signal component at the plurality of sampling time points, and then calculate the SSNR and the noise reduction gain of the audio signal component based at least on the statistical value of the signal power and the statistical value of the self-noise power. For example, the statistical value may be a sum, an average value, a peak value, a root mean square value, or the like, of power at the plurality of sampling time points. For example, the gain calculation unit 708 may calculate the SSNR and the noise reduction gain of the audio signal component according to Equations (2) and (3) , respectively.
[0075] The smoothing unit 710 may be configured to smooth the noise reduction gain based at least on the SSNR and the noise reduction gain of the audio signal to generate a smoothed gain of the audio signal. The smoothing operation is used to reduce fluctuation among frames and among frequency bands, thereby providing a processed audio signal with a better hearing experience.
[0076] In one or more embodiments of the present disclosure, the smoothing unit 710 may be configured to, for each audio signal component of the plurality of audio signal components, perform, based at least on the SSNR and the noise reduction gain of the audio signal component and between a previous frame of audio signal component of the audio signal component and the audio signal component, a first smoothing operation on the noise reduction gain of the audio signal component, to generate a first smoothed gain of the audio signal component.
[0077] Specifically, the first smoothing operation may include: determining a first weight coefficient for the audio signal component based at least on the SSNR and the noise reduction gain of the audio signal component; and performing, by using the first weight coefficient for the audio signal component, a weighted summation of a first smoothed gain of the previous frame of audio signal component in a corresponding frequency band and the noise reduction gain of the audio signal component, to generate the first smoothed gain of the audio signal component. For example, the smoothing unit 710 may calculate the first smoothed gain according to the above Equation (4) .
[0078] The smoothing unit 710 may determine the value of the first weight coefficient by SSNR for each audio signal component, a weighted sum of SSNRs for different audio signal components, and the noise reduction gain. Specifically, for each audio signal component of the plurality of audio signal components, the smoothing unit 710 may calculate a weighted sum of SSNRs of the plurality of audio signal components, and then determine the first weight coefficient for the audio signal component based on the weighted sum of the SSNRs and the SSNR and the noise reduction gain of the audio signal component.
[0079] Then, the first smoothed gain of each audio signal component of the audio signal may be applied, by the noise reduction unit 712, to the audio signal to generate a processed audio signal with the self-noise reduced or removed. The first smoothing operation can reduce fluctuation among frames of audio signals, thereby providing a processed audio signal with a better hearing experience.
[0080] In one or more embodiments of the present disclosure, the smoothing unit 710 may be further configured to, for each audio signal component of the plurality of audio signal components, perform, based at least on the SSNR and the noise reduction gain of the audio signal component and among a frequency band corresponding to the audio signal component, a first adjacent frequency band with a smaller central frequency than the frequency band and a second adjacent frequency band with a larger central frequency than the frequency band, a second smoothing operation on the first smoothed gain of the audio signal component, to generate a second smoothed gain of the audio signal component.
[0081] Specifically, the second smoothing operation may include performing, by using a plurality of second weight coefficients, a weighted summation of the first smoothed gain of the audio signal component, a first smoothed gain of a first adjacent audio signal component in the first adjacent frequency band, and a first smoothed gain of a second adjacent audio signal component in the second adjacent frequency band, to generate the second smoothed gain of the audio signal component. For example, the smoothing unit 710 may calculate the second smoothed gain according to the above Equation (6) .
[0082] The smoothing unit 710 may determine the second weight coefficients, for example, according to practical requirements, which are not specifically limited by the embodiments of the present disclosure. For example, among the plurality of second weight coefficients, a second weight coefficient for the audio signal component may be greater than that for the first adjacent audio signal component in the first adjacent band with a smaller central frequency and the second adjacent audio signal component in the second adjacent band with a larger central frequency, to ensure that the first smoothed gain for the current audio signal component contributes the most to the final second smoothed gain.
[0083] Then, the second smoothed gain of each audio signal component of the audio signal may be applied, by the noise reduction unit 712, to the audio signal to generate a processed audio signal with the self-noise reduced or removed. The second smoothing operation can reduce fluctuation among frequency bands of audio signals, thereby providing a processed audio signal with a better hearing experience.
[0084] In one or more embodiments of the present disclosure, the smoothing unit 710 may perform only the first smoothing operation for the noise reduction gain, and the first smoothed gain may be the final noise reduction gain to be applied to the audio signal by the noise reduction unit 712. In one or more embodiments of the present disclosure, the smoothing unit 710 may perform both the first and second smoothing operations for the noise reduction gain, and the second smoothed gain may be the final noise reduction gain to be applied to the audio signal by the noise reduction unit 712.
[0085] The noise reduction unit 712 may be configured to apply the smoothed gain, such as the first smoothed gain or the second smoothed gain, to the audio signal to generate a processed audio signal with the self-noise reduced or removed. In one or more embodiments of the present disclosure, noise reduction unit 712 may multiply the smoothed gain directly with the audio signal to generate the processed audio signal, for example, according to the above Equation (7) .
[0086] In one or more embodiments of the present disclosure, the noise reduction unit 712 may be configured to apply the smoothed gain, such as the first smoothed gain or the second smoothed gain, to a WDRC processing module to generate the processed audio signal, for example, according to the above Equation (8) .
[0087] The processed audio signal component in different frequency bands may be composited, for example, by a compositing module (not illustrated) , to obtain the processed audio signal of the current frame of the audio signal.
[0088] With the audio processing apparatus 700, self-noises in audio signals captured by audio acquisition devices, such as microphones, may be effectively reduced or substantially removed without distorting the audio signals in a continuous noise reduction process and no additional latency of the processed audio signals will be introduced.
[0089] In one or more embodiments of the present disclosure, there is further provided an audio processing device comprising one or more processors and one or more memories, where the one or more memories have stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to execute the audio processing method as described above.
[0090] In addition, an audio device according to one or more embodiments of the present disclosure may also be realized by means of an architecture of an exemplary audio device shown in FIG. 8. FIG. 8 illustrates a schematic diagram of an architecture of an exemplary audio device 800 in accordance with one or more embodiments of the present disclosure. As shown in FIG. 8, the audio device 800 may include one or more audio acquisition components 802, a bus 804, one or more processors 806, a Read-Only Memory (ROM) 808, a Random Access Memory (RAM) 810, a communication interface 812 connected to a network, one or more audio playback components 814, and the like.
[0091] The one or more audio acquisition components 802 may be, for example, microphones, which can acquire audio signals. The one or more audio playback components 814 may be, for example, one or more speakers, which may play audio signals received or processed by the audio device 800. The audio device 800 may be connected, via the communication interface 812, to a network such as WiFi, Bluetooth, 4G or 5G wireless network, etc., to receive audio signals or control signaling from the network, or to send audio signals to the network.
[0092] A storage device in the audio device 800, such as the ROM 808, may store various data or files processed by the device and / or used for communication, as well as program instructions to be executed by the processors 806. In some cases, the audio device 800 may also include a user interface (not shown) . It should be appreciated that the architecture shown in FIG. 8 is only exemplary, and one or more components of the audio device 800 shown in FIG. 8 can be omitted according to practical requirements. The audio device 800 according to one or more embodiments of the present disclosure may be configured to perform the audio processing method according to one or more embodiments of the present disclosure, or to implement the audio processing apparatus according to one or more embodiments of the present disclosure.
[0093] One or more embodiments of the present disclosure may also be implemented as a computer-readable storage medium. A computer-readable storage medium according to one or more embodiments of the present disclosure has computer-readable instructions stored thereon, which, when executed by a processor, cause the processor to execute the audio processing method according to one or more embodiments of the present disclosure described with reference to the above drawings. The computer-readable storage medium may include, but is not limited to, volatile memory and / or nonvolatile memory, for example. The volatile memory may include, for example, Random Access Memory (RAM) and / or cache, and the like. The nonvolatile memory may include, for example, a Read-Only Memory (ROM) , a hard disk, a flash memory, and the like.
[0094] According to one or more embodiments of the present disclosure, there is also provided a computer program product or computer program including computer-readable instructions stored in a computer-readable storage medium. A processor of a computer device may read the computer-readable instructions from the computer-readable storage medium, and the processor executes the computer-readable instructions, so that the computer device performs the audio processing method described in one or more embodiments described above.
[0095] The following is a non-limiting list of examples that are in accordance with one or more techniques of this disclosure.
[0096] Example 1. An audio processing method, comprising: acquiring an audio signal; dividing the audio signal into a plurality of audio signal components in a plurality of frequency bands; estimating self-noise in each of the plurality of audio signal components of the audio signal; calculating a signal to self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components; smoothing the noise reduction gain, based at least on the SSNR and the noise reduction gain, to generate a smoothed gain; and applying the smoothed gain to the audio signal to generate a processed audio signal with the self-noise reduced.
[0097] Example 2. The method of Example 1, wherein estimating the self-noise in each of the plurality of audio signal components of the audio signal comprises: estimating minimum power of each of the plurality of audio signal components; and determining the self-noise in each of the plurality of audio signal components based on the minimum power.
[0098] Example 3. The method of any one of Examples 1-2, wherein the audio signal is sampled at a plurality of sampling time points, and wherein calculating the SSNR and the noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components comprises, for each audio signal component of the plurality of audio signal components: determining a statistical value of signal power and a statistical value of self-noise power of the audio signal component at the plurality of sampling time points; and calculating the SSNR and the noise reduction gain of the audio signal component based at least on the statistical value of the signal power and the statistical value of the self-noise power.
[0099] Example 4. The method of any one of Examples 1-3, wherein calculating the SSNR and the noise reduction gain of the audio signal component based at least on the statistical value of the signal power and the statistical value of the self-noise power comprises: determining an over-estimation coefficient for the audio signal component based on a frequency distribution of the audio signal component; and calculating the SSNR and the noise reduction gain of the audio signal component based on the over-estimated coefficient, the statistical value of the signal power and the statistical value of the self-noise power.
[0100] Example 5. The method of any one of Examples 1-4, wherein the statistical value is a sum, an average value, a peak value, or a root mean square value of power at the plurality of sampling time points.
[0101] Example 6. The method of any one of Examples 1-5, wherein smoothing the noise reduction gain based at least on the SSNR and the noise reduction gain of the audio signal to generate the smoothed gain of the audio signal comprises, for each audio signal component of the plurality of audio signal components: performing, based at least on the SSNR and the noise reduction gain of the audio signal component and between the audio signal component and a previous frame of audio signal component of the audio signal component, a first smoothing operation on the noise reduction gain of the audio signal component, to generate a first smoothed gain of the audio signal component.
[0102] Example 7. The method of any one of Examples 1-6, wherein the first smoothing operation comprises: determining a first weight coefficient for the audio signal component based at least on the SSNR and the noise reduction gain of the audio signal component; and performing, by using the first weight coefficient for the audio signal component, a weighted summation of a first smoothed gain of the previous frame of audio signal component in a corresponding frequency band and the noise reduction gain of the audio signal component, to generate the first smoothed gain of the audio signal component.
[0103] Example 8. The method of any one of Examples 1-7, wherein determining the first weight coefficient for the audio signal component based at least on the SSNR and the noise reduction gain of the audio signal component comprises: calculating a weighted sum of SSNRs of the plurality of audio signal components; and determining the first weight coefficient for the audio signal component based on the weighted sum of the SSNRs and the SSNR and the noise reduction gain of the audio signal component.
[0104] Example 9. The method of any one of Examples 1-8, wherein smoothing the noise reduction gain based at least on the SSNR and the noise reduction gain of the audio signal to generate the smoothed gain of the audio signal further comprises, for each audio signal component of the plurality of audio signal components: performing, based at least on the SSNR and the noise reduction gain of the audio signal component and among a frequency band corresponding to the audio signal component, a first adjacent frequency band with a smaller central frequency than the frequency band and a second adjacent frequency band with a larger central frequency than, a second smoothing operation on the first smoothed gain of the audio signal component, to generate a second smoothed gain of the audio signal component.
[0105] Example 10. The method of any one of Examples 1-9, wherein the second smoothing operation comprises: performing, by using a plurality of second weight coefficients, a weighted summation of the first smoothed gain of the audio signal component, a first smoothed gain of a first adjacent audio signal component in the first adjacent frequency band, and a first smoothed gain of a second adjacent audio signal component in the second adjacent frequency band, to generate the second smoothed gain of the audio signal component.
[0106] Example 11. The method of any one of Examples 1-10, wherein, among the plurality of second weight coefficients, a second weight coefficient for the audio signal component is greater than that for the first adjacent audio signal component and the second adjacent audio signal component.
[0107] Example 12. The method of any one of Examples 1-11, wherein applying the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced comprises: multiplying the smoothed gain with the audio signal to generate the processed audio signal.
[0108] Example 13. The method of any one of Examples 1-12, wherein applying the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced comprises: applying the smoothed gain to a wide dynamic range compression (WDRC) processing module to generate the processed audio signal.
[0109] Example 14. The method of any one of Examples 1-13, wherein the audio signal is a signal captured by a microphone, or a signal captured by the microphone and processed by other audio processing modules.
[0110] Example 15. The method of any one of Examples 1-14, wherein the microphone is in a personal sound amplification product (PSAP) or a hearing-aid.
[0111] Example 16. An audio processing apparatus, comprising: an audio acquisition unit configured to acquire an audio signal; a frequency dividing unit configured to divide the audio signal into a plurality of audio signal components in a plurality of frequency bands; a self-noise estimation unit configured to estimate self-noise in each of the plurality of audio signal components of the audio signal; a gain calculation unit configured to calculate a signal to self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components; a smoothing unit configured to smooth the noise reduction gain, based at least on the SSNR and the noise reduction gain, to generate a smoothed gain; and a noise reduction unit configured to apply the smoothed gain to the audio signal to generate a processed audio signal with the self-noise reduced.
[0112] Example 17. An audio processing device, comprising: at least one microphone configured to capture an audio signal; at least one speaker configured to output a processed audio signal; and one or more processors coupled to the at least one microphone and the at least one speaker, configured to: divide the audio signal into a plurality of audio signal components in a plurality of frequency bands; estimate self-noise in each of the plurality of audio signal components of the audio signal; calculate a signal to self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components; smooth the noise reduction gain, based at least on the SSNR and the noise reduction gain of the audio signal, to generate a smoothed gain of the audio signal; and apply the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced.
[0113] Example 18. An audio processing device, comprising: one or more processors; and one or more memories having stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to perform the method of any one of any one of Examples 1-15.
[0114] Example 19. A computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to perform the method of any of Examples 1-15.
[0115] Example 20. A computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method of any of Examples 1-15.
[0116] It is to be recognized that depending on the examples, certain acts or events of any of the techniques described herein can be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the techniques) . Moreover, in certain examples, acts or events may be performed concurrently, e.g., through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.
[0117] Program portions of the technology may be considered to be “product” or “article” that exists in the form of executable codes and / or related data, which are embodied or implemented by a computer-readable medium. A tangible, permanent storage medium may include an internal memory, or a storage used by computers, processors, or similar devices or associated modules. For example, various semiconductor memories, tape drivers, disk drivers, or any similar devices capable of providing storage functionality for software.
[0118] All software or parts of it may sometimes communicate over a network, such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. For example, loading from one server or host computer to a hardware environment of one computer environment, or other computer environment implementing the system, or a system having a similar function associated with providing information needed for the communication method. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, electric waves, electromagnetic waves, etc., to be propagated through cables, optical cables, or air. A physical medium used for carrying the waves such as cables, wireless connections, or fiber optic cables may also be considered as a medium for carrying the software. In usage herein, unless a tangible “storage” medium is defined, other terms referring to a computer or machine “readable medium” mean a medium that participates in execution of any instruction by the processor.
[0119] The present application uses specific words to describe embodiments of the present disclosure. Reference to “an embodiment, ” “one or more embodiments, ” and / or “some embodiments” means a feature, structure, or characteristic in connection with at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that two or more references to “an embodiment, ” “one embodiment, ” or “an alternative embodiment” in various places throughout this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics may be combined as suitable in one or more embodiments of the application.
[0120] Moreover, one skilled in the art will appreciate that aspects of the present disclosure may be illustrated and described in terms of a number of patentable categories or instances, including any new and useful process, machine, manufacture, or combination of matter, or any new and useful improvement thereof. Accordingly, aspects of the present disclosure may be performed entirely by hardware, entirely by software (including firmware, resident software, micro-code, etc. ) , or by a combination of hardware and software. The above hardware or software may each be referred to as a “data block, ” “module, ” “engine, ” “unit, ” “component, ” or “system. ” Furthermore, aspects of the present disclosure may be embodied as a computer product embodied in one or more computer-readable media including computer-readable program code.
[0121] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or extremely formal sense unless expressly so defined herein.
[0122] While various embodiments of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more embodiments and implementations are possible that are within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
1.An audio processing method, comprising:acquiring an audio signal;dividing the audio signal into a plurality of audio signal components in a plurality of frequency bands;estimating self-noise in each of the plurality of audio signal components of the audio signal;calculating a signal-to-self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components;smoothing the noise reduction gain, based at least on the SSNR and the noise reduction gain, to generate a smoothed gain; andapplying the smoothed gain to the audio signal to generate a processed audio signal with the self-noise reduced.2.The method of claim 1, wherein estimating the self-noise in each of the plurality of audio signal components of the audio signal comprises:estimating minimum power of each of the plurality of audio signal components; anddetermining the self-noise in each of the plurality of audio signal components based on the minimum power.3.The method of claim 1, wherein the audio signal is sampled at a plurality of sampling time points, and wherein calculating the SSNR and the noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components comprises, for each audio signal component of the plurality of audio signal components:determining a statistical value of signal power and a statistical value of self-noise power of the audio signal component at the plurality of sampling time points; andcalculating the SSNR and the noise reduction gain of the audio signal component based at least on the statistical value of the signal power and the statistical value of the self-noise power.4.The method of claim 3, wherein calculating the SSNR and the noise reduction gain of the audio signal component based at least on the statistical value of the signal power and the statistical value of the self-noise power comprises:determining an over-estimation coefficient for the audio signal component based on a frequency distribution of the audio signal component; andcalculating the SSNR and the noise reduction gain of the audio signal component based on the over-estimated coefficient, the statistical value of the signal power, and the statistical value of the self-noise power.5.The method of claim 1, wherein the statistical value is a sum, an average value, a peak value, or a root mean square value of power at the plurality of sampling time points.6.The method of claim 1, wherein smoothing the noise reduction gain based at least on the SSNR and the noise reduction gain of the audio signal to generate the smoothed gain of the audio signal comprises, for each audio signal component of the plurality of audio signal components:performing, based at least on the SSNR and the noise reduction gain of the audio signal component and between the audio signal component and a previous frame of audio signal component of the audio signal component, a first smoothing operation on the noise reduction gain of the audio signal component, to generate a first smoothed gain of the audio signal component.7.The method of claim 6, wherein the first smoothing operation comprises:determining a first weight coefficient for the audio signal component based at least on the SSNR and the noise reduction gain of the audio signal component; andperforming, by using the first weight coefficient for the audio signal component, a weighted summation of a first smoothed gain of the previous frame of audio signal component in a corresponding frequency band and the noise reduction gain of the audio signal component, to generate the first smoothed gain of the audio signal component.8.The method of claim 7, wherein determining the first weight coefficient for the audio signal component based at least on the SSNR and the noise reduction gain of the audio signal component comprises:calculating a weighted sum of SSNRs of the plurality of audio signal components; anddetermining the first weight coefficient for the audio signal component based on the weighted sum of the SSNRs, and the SSNR and the noise reduction gain of the audio signal component.9.The method of claim 6, wherein smoothing the noise reduction gain based at least on the SSNR and the noise reduction gain of the audio signal to generate the smoothed gain of the audio signal further comprises, for each audio signal component of the plurality of audio signal components:performing, based at least on the SSNR and the noise reduction gain of the audio signal component and among a frequency band corresponding to the audio signal component, a first adjacent frequency band with a smaller central frequency than the frequency band and a second adjacent frequency band with a larger central frequency than the frequency band, a second smoothing operation on the first smoothed gain of the audio signal component, to generate a second smoothed gain of the audio signal component.10.The method of claim 9, wherein the second smoothing operation comprises:performing, by using a plurality of second weight coefficients, a weighted summation of the first smoothed gain of the audio signal component, a first smoothed gain of a first adjacent audio signal component in the first adjacent frequency band, and a first smoothed gain of a second adjacent audio signal component in the second adjacent frequency band, to generate the second smoothed gain of the audio signal component.11.The method of claim 10, wherein, among the plurality of second weight coefficients, a second weight coefficient for the audio signal component is greater than that for the first adjacent audio signal component and the second adjacent audio signal component.12.The method of claim 1, wherein applying the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced comprises:multiplying the smoothed gain with the audio signal to generate the processed audio signal.13.The method of claim 1, wherein applying the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced comprises:applying the smoothed gain to a wide dynamic range compression (WDRC) processing module to generate the processed audio signal.14.The method of claim 1, wherein the audio signal is a signal captured by a microphone, or a signal captured by the microphone and processed by other audio processing modules.15.The method of claim 14, wherein the microphone is in a personal sound amplification product (PSAP) or a hearing-aid.16.An audio processing apparatus, comprising:an audio acquisition unit configured to acquire an audio signal;a frequency dividing unit configured to divide the audio signal into a plurality of audio signal components in a plurality of frequency bands;a self-noise estimation unit configured to estimate self-noise in each of the plurality of audio signal components of the audio signal;a gain calculation unit configured to calculate a signal-to-self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components;a smoothing unit configured to smooth the noise reduction gain, based at least on the SSNR and the noise reduction gain, to generate a smoothed gain; anda noise reduction unit configured to apply the smoothed gain to the audio signal to generate a processed audio signal with the self-noise reduced.17.An audio processing device, comprising:at least one microphone configured to capture an audio signal;at least one speaker configured to output a processed audio signal; andone or more processors coupled to the at least one microphone and the at least one speaker, configured to:divide the audio signal into a plurality of audio signal components in a plurality of frequency bands;estimate self-noise in each of the plurality of audio signal components of the audio signal;calculate a signal-to-self-noise ratio (SSNR) and a noise reduction gain of the audio signal based on the audio signal and the self-noise in each of the plurality of audio signal components;smooth the noise reduction gain, based at least on the SSNR and the noise reduction gain of the audio signal, to generate a smoothed gain of the audio signal; andapply the smoothed gain to the audio signal to generate the processed audio signal with the self-noise reduced.18.An audio processing device, comprising:one or more processors; andone or more memories having stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 1-15.19.A computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to perform the method of any of claims 1-15.20.A computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method of any of claims 1-15.
Citation Information
Patent Citations
Microphone channel self-noise silencing
US20240223948A1