Apparatus and method for own voice suppression

TWI938926BActive Publication Date: 2026-09-11REALTEK SEMICON CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
TW114112670
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2026-09-11
Estimated Expiration
2045-03-31

Smart Images

  • Figure TWG2TB001910494_001
    Figure TWG2TB001910494_001
  • Figure TWG2TB001910494_002
    Figure TWG2TB001910494_002
  • Figure TWG2TB001910494_003
    Figure TWG2TB001910494_003
Patent Text Reader

Abstract

The self-speech suppression method includes: receiving a first audio signal and a second audio signal from a first microphone and a second microphone, respectively; receiving the first audio signal through an optimized filter and generating a third audio signal based on the optimal filter coefficients; subtracting the second audio signal from the third audio signal to obtain an error signal; converting the error signal and the omnidirectional signal into an error frequency domain signal and an omnidirectional frequency domain signal, respectively, wherein the omnidirectional signal is the first audio signal, the second audio signal, or a linear combination thereof; calculating the amplitude ratio of the error frequency domain signal and the omnidirectional frequency domain signal in each frequency band; and determining the compression ratio of each frequency band based on the amplitude ratio of each frequency band.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A self-speech suppression device, comprising: A first microphone for receiving audio and outputting a first audio signal accordingly; A second microphone for receiving audio and correspondingly outputting a second audio signal, wherein each of the first and second microphones is an external microphone; an optimization filter, communicatively connected to the first microphone to receive the first audio signal and generate a third audio signal according to an optimal filter coefficient; a subtractor, communicatively connected to the second microphone and the optimization filter to subtract the second audio signal from the third audio signal to obtain an error signal; a frequency domain converter, communicatively connected to the subtractor to convert the error signal and an omnidirectional signal into an error frequency domain signal and an omnidirectional frequency domain signal, respectively, wherein the omnidirectional signal is the first audio signal, the second audio signal, or a linear combination thereof; and a spectrum comparator, communicatively connected to the frequency domain converter to calculate an amplitude ratio of the error frequency domain signal and the omnidirectional frequency domain signal in each frequency band. A similarity comparator, communicatively connected to the second microphone and the optimization filter, compares the similarity between the second audio signal and the third audio signal to generate a self-speech flag, wherein the self-speech flag indicates whether the first microphone and the second microphone have recorded self-speech; and a compression ratio determiner, communicatively connected to the spectrum comparator, determines a compression ratio for each frequency band based on the self-speech flag and the amplitude ratio of each frequency band, wherein the compression ratio determiner performs a smoothing process on the amplitude ratio of each frequency band based on the self-speech flag to obtain the compression ratio of each frequency band, wherein the smoothing process includes: when the self-speech flag indicates that the first microphone and the second microphone have recorded self-speech, causing the compression ratio to move closer to the amplitude ratio; and when the self-speech flag indicates that the first microphone and the second microphone have not recorded self-speech, causing the compression ratio to move closer to 1.

2. The self-speech suppression device as described in claim 1, wherein the formula for bringing the compression ratio toward the amplitude ratio is as follows: ; wherein the formula for bringing the compression ratio toward 1 is as follows: ; wherein is the amplitude ratio, is the compression ratio, and α is a value less than and close to 1.

3. The self-speech suppression device as described in claim 1 further includes: A multiplier, communicatively connected to the compression ratio determiner, multiplies a frequency domain signal to be suppressed by the compression ratio to generate a suppressed frequency domain signal.

4. The self-speech suppression device as described in claim 3, wherein the frequency domain signal to be suppressed is obtained by selectively subjecting the first audio signal and the second audio signal to beamforming or not, and selectively subjecting them to noise suppression or not, and then converting them from the time domain to the frequency domain.

5. The self-speech suppression device as claimed in claim 1, wherein the optimized filter is trained in an environment where the ambient sound pressure level does not exceed 50 dB, and the optimal filter coefficients are such that the error signal has a minimum value.

6. The self-speech suppression device as described in claim 1, wherein the similarity is a cosine similarity or a correlation coefficient, wherein when the similarity is greater than a threshold, the self-speech flag indicates that the first microphone and the second microphone have recorded self-speech.

7. A self-speech suppression method, comprising: A first microphone picks up sound and outputs a first audio signal accordingly, and a second microphone picks up sound and outputs a second audio signal accordingly, wherein each of the first and second microphones is an external microphone; the first audio signal is received by an optimized filter and a third audio signal is generated according to an optimal filter coefficient; the second audio signal is subtracted from the third audio signal to obtain an error signal; the error signal and an omnidirectional signal are respectively converted into an error frequency domain signal and an omnidirectional frequency domain signal, wherein the omnidirectional signal is the first audio signal, the second audio signal, or a linear combination thereof; the amplitude ratio of the error frequency domain signal and the omnidirectional frequency domain signal in each frequency band is calculated; A self-voice flag is generated by comparing the similarity between the second audio message and the third audio message, wherein the self-voice flag is used to indicate whether the first microphone and the second microphone have recorded self-voice; and a compression ratio for each frequency band is determined based on the self-voice flag and the amplitude ratio of each frequency band, wherein the compression ratio of each frequency band is obtained by smoothing the amplitude ratio of each frequency band based on the self-voice flag, wherein the smoothing process includes: when the self-voice flag indicates that the first microphone and the second microphone have recorded self-voice, the compression ratio is made to move closer to the amplitude ratio; and when the self-voice flag indicates that the first microphone and the second microphone have not recorded self-voice, the compression ratio is made to move closer to 1.

Citation Information

Patent Citations

  • Methods and apparatus for monitoring consciousness

    CN101401724A

  • A hearing device comprising a dynamic compressive amplification system and a method of operating a hearing device

    CN108235211A

  • Systems, methods, apparatus, and computer-readable media for phase-based processing of multichannel signal

    TW201132138A

  • Modem with Voice Processing Capability

    US20110200048A1

  • Hearing assistance system with own voice detection

    US20150043765A1