A noise reduction method applied to three microphones

By constructing a filter through a three-microphone system and utilizing frequency response ratio and post-filter adjustment, the problems of noise filtering and voice restoration in open offices are solved, achieving higher wearer voice recognition accuracy and noise suppression effect.

CN115720316BActive Publication Date: 2025-10-03YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211460635.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-10-03
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively filtering out individual angle noise and interfering human voices in open office scenarios, and the degree of speech restoration is greatly affected by the way the headphones are worn. The existing dual-microphone correlation and energy difference methods have a high false detection rate for speech noise.

Method used

A three-microphone system is used to form beams in different directions through the main microphone and two auxiliary microphones. The frequency response ratio is calculated and filtered. Combined with the post-filter and the adjustment of the long-short-time energy difference ratio, a filter is constructed to reduce noise.

Benefits of technology

It improves the wearer's call quality, accurately recognizes the wearer's voice, effectively filters out interfering noise in open offices, especially interfering human voices, and solves the high-frequency loss and wearing angle dependence problems of traditional beamforming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115720316B_ABST
    Figure CN115720316B_ABST
Patent Text Reader

Abstract

The present invention discloses a noise reduction method applied to three microphones. The method includes collecting sound signals through the main microphone to form a target beam, collecting two sound signals through the right feedforward auxiliary microphone and the left feedforward auxiliary microphone to form dual-microphone beams in different directions; calculating the frequency response of the first dual-microphone beam, the second dual-microphone beam and the third dual-microphone beam respectively, calculating the frequency response ratio, and performing a first filtering process on the target beam according to the frequency response ratio to obtain a first target signal; constructing a post-filter according to the prior signal-to-noise ratio of the current frame, and performing a second filtering process on the first target signal according to the post-filter; adjusting the post-filter according to the long-time filtering energy difference ratio and the short-time filtering energy difference ratio of the sound signal collected by the main microphone. The technical solution of the present invention reduces the external noisy background noise and interfering human voices in the microphone voice signal, and improves the call quality of the wearer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound signal noise reduction, and in particular to a noise reduction method applied to three microphones. Background Art

[0002] In a typical open office setting, external noise interference, such as keyboard sounds, tapping, and talking, can affect the quality of office calls. This is especially true when there are other interfering human voices around the wearer. Existing technology uses array microphones and directional algorithms for noise reduction. However, this technology has the disadvantage of not filtering out noise from certain angles, and the degree of noise reduction and speech reproduction is significantly affected by the angle at which the sound enters the earphones, that is, by how the earphones are worn. Existing technology also uses dual-microphone correlation and energy difference for speech detection and then noise reduction. However, this method has the disadvantage of being difficult to filter out noise between voices, and the false detection rate for speech noise is high and difficult to avoid. Summary of the Invention

[0003] The present invention provides a noise reduction method applied to three microphones, which reduces external noisy background noise and interfering human voices in the microphone voice signal and improves the call quality of the wearer.

[0004] The present invention provides a noise reduction method applied to three microphones, comprising the following steps:

[0005] The main microphone collects sound signals to form a target beam, and the right and left feedforward sub-microphones collect two sound signals to form dual-microphone beams in different directions, including a first dual-microphone beam in the direction of the right feedforward sub-microphone, a second dual-microphone beam in the direction of the left feedforward sub-microphone, and a third dual-microphone beam in the direction of the main microphone;

[0006] respectively calculating the frequency responses of the first dual-microphone beam, the second dual-microphone beam, and the third dual-microphone beam to obtain a first frequency response, a second frequency response, and a third frequency response, calculating a frequency response ratio according to the first frequency response, the second frequency response, and the third frequency response, and performing a first filtering process on the target beam according to the frequency response ratio to obtain a first target signal;

[0007] constructing a post-filter according to a priori signal-to-noise ratio of the current frame, and performing a second filtering process on the first target signal according to the post-filter;

[0008] The post-filter is adjusted according to the long-time filtering energy difference ratio and the short-time filtering energy difference ratio of the sound signal collected by the main microphone; the long-time filtering energy difference ratio refers to the energy difference ratio updated once every first preset number of frames, and the short-time filtering energy difference ratio refers to the energy difference ratio updated once every second preset number of frames, the first preset number of frames is greater than the second preset number of frames, and the energy difference ratio is calculated based on the energy cumulative value of the sound signal collected by the main microphone before the first filtering process and the energy cumulative value after the second filtering process.

[0009] Furthermore, the frequency response ratio is calculated based on the first frequency response, the second frequency response, and the third frequency response, specifically:

[0010] .

[0011] Furthermore, after the frequency response ratio value is subjected to nonlinear processing, the third dual-microphone beam is filtered according to the frequency response ratio value to obtain a first target signal.

[0012] Furthermore, the short-term filtering energy difference and the long-term filtering energy difference of the current frame are calculated according to the following formula:

[0013] P(t)=delta*[(OrignalSum(t)-AfterfilterSum(t)) / OrignalSum(t)]+(1-delta)*H(t)

[0014] When calculating the short-time filtered energy difference ratio of the current frame, P(t) represents the short-time filtered energy difference ratio of the current frame, t represents the second preset frame number, OrignalSum(t) represents the cumulative energy value of the sound signal collected by the main microphone t frames before the first filtering process, AfterfilterSum(t) represents the cumulative energy value of the sound signal collected by the main microphone t frames after the second filtering process, delta represents the second forgetting factor, and H(t) represents the short-time filtered energy difference ratio t frames before;

[0015] When calculating the long-term filtered energy difference ratio of the current frame, P(t) represents the long-term filtered energy difference ratio of the current frame, t represents the first preset frame number, OrignalSum(t) represents the cumulative energy value of the sound signal collected by the main microphone t frames before the first filtering process, AfterfilterSum(t) represents the cumulative energy value of the sound signal collected by the main microphone t frames after the second filtering process, delta represents the first forgetting factor, and H(t) represents the long-term filtered energy difference ratio t frames ago.

[0016] Furthermore, adjusting the post filter according to the long-time filtering energy difference ratio and the short-time filtering energy difference ratio comprises the following steps:

[0017] When the long-term filtered energy difference ratio is less than or equal to a first preset threshold, reducing the filtering degree of the post-filter;

[0018] When the long-time filtering energy difference ratio is greater than the first preset threshold and the short-time filtering energy difference ratio is less than or equal to the second preset threshold, reducing the filtering degree of the post-filter;

[0019] When the long-time filtering energy difference ratio is greater than the first preset threshold and the short-time filtering energy difference ratio is greater than the second preset threshold, the filtering degree of the post-filter is increased.

[0020] Furthermore, the a priori signal-to-noise ratio of the current frame is calculated based on the first target signal, the signal-to-noise ratio of the previous frame, and the dual-microphone beam frequency response, where the dual-microphone beam frequency response is the larger frequency response of the first frequency response and the second frequency response.

[0021] Furthermore, the a priori signal-to-noise ratio of the current frame is calculated based on the first target signal, the signal-to-noise ratio of the previous frame, and the dual-microphone beam frequency response, specifically:

[0022] snr =alpha* (y – n) / n + (1-alpha)snr_old

[0023] Wherein, y represents the first target signal, n represents the dual-microphone beam frequency response, snr_old represents the signal-to-noise ratio of the previous frame, and alpha represents the second forgetting factor.

[0024] Furthermore, a post-filter is constructed based on the prior signal-to-noise ratio of the current frame, specifically:

[0025] filterpost = (snr) / (snr + 1)

[0026] Wherein, filterpost represents the post filter, and snr represents the prior signal-to-noise ratio of the current frame.

[0027] The embodiments of the present invention have the following beneficial effects:

[0028] This invention provides a noise reduction method for a three-microphone system. This method utilizes multiple beamforming frequency responses at a target direction and at a maximum angle from the target direction to construct a filter for noise reduction. Compared to existing technologies, this method achieves higher accuracy in speaker recognition and better noise cancellation for ambient noise, particularly interfering human voices in open office environments. It addresses the high-frequency loss and headphone angle dependency issues of traditional beamforming. This method uses the beamforming frequency response at the maximum angle from the target direction as a noise estimate for the postfilter and modifies the postfilter using the difference in long- and short-time energy before and after filtering. Compared to traditional postfilters, this method more accurately recognizes the wearer's voice, avoids target signal loss, and improves noise filtering. Compared to existing noise reduction algorithms, this method effectively filters out interfering noise in open office environments, particularly ambient human voices (which are difficult for existing speech noise reduction algorithms to filter out), while also improving the quality of the wearer's speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 1 is a flow chart of a noise reduction method applied to three microphones provided by an embodiment of the present invention;

[0030] Figure 2 1 is a schematic diagram of the position relationship of three microphones in a noise reduction method for three microphones provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0032] like Figure 1 As shown, an embodiment of the present invention provides a noise reduction method applied to three microphones, including the following steps:

[0033] Step S101: The main microphone collects sound signals to form a target beam, and the right and left feedforward sub-microphones collect two sound signals and form dual-microphone beams in different directions, including a first dual-microphone beam in the direction of the right feedforward sub-microphone, a second dual-microphone beam in the direction of the left feedforward sub-microphone, and a third dual-microphone beam in the direction of the main microphone. Figure 2 As shown, the left feedforward sub-microphone FFL_mic is set at 0°, the right feedforward sub-microphone FFR_mic is set at 180°, the main microphone Main_mic is set at 90°, and the noise sources include ambient noise and interfering human voices.

[0034] In one embodiment, the two-channel sound signal is collected according to the MVDR model to form a dual-microphone beamformation. Specifically, the two-channel noisy sound signal x = [x1, x2] collected by the right and left feedforward auxiliary microphones is subjected to a short-time Fourier transform to obtain X = [X1, X2]. The noise reduction signal Y = W*X is calculated based on the MVDR model weight W, where W = (Φ-1*ds) / (ds*Φ-1* ds). The principle of the MVDR model is to minimize the total noise output power W*Φ*W without generating speech distortion W*ds = 1, where Φ is the noise cross-power spectrum matrix of the two microphones, Φ = E{X*X}, W is the weight value of the beamforming output, and ds is the transfer function between the sound and the microphones. ds is related to the microphone spacing d and the relative angle θ between the microphones and the human mouth.

[0035] Step S102: Calculate the frequency responses of the first dual-microphone beam, the second dual-microphone beam, and the third dual-microphone beam to obtain a first frequency response, a second frequency response, and a third frequency response, calculate a frequency response ratio based on the first frequency response, the second frequency response, and the third frequency response, and perform a first filtering process on the target beam based on the frequency response ratio to obtain a first target signal.

[0036] As one embodiment, the frequency response ratio value is calculated according to the following formula:

[0037]

[0038] If the sound picked up by the two secondary microphones is the wearer's voice, the frequency response of the 90° beamforming is greater, while the frequency responses of the 0° and +180° beamforming are smaller, resulting in a larger frequency response ratio. If the sound picked up by the two secondary microphones is external noise or interfering human voices, the frequency response of the 90° beamforming is smaller, while the frequency responses of the 0° and +180° beamforming are greater, resulting in a smaller frequency response ratio. Existing beamforming methods only focus on the frequency response in the target direction (i.e., enhancing sounds from the target direction and suppressing sounds from non-target directions). However, unidirectional beamforming relies heavily on transfer function estimation, which is affected by the headphone wearing angle, and jitter in the wearing angle can affect the noise reduction results. Furthermore, unidirectional beamforming suffers from limited noise reduction energy and a loss of high-frequency energy. The present invention, by simultaneously addressing frequency responses in all three directions, effectively addresses these existing issues.

[0039] In one embodiment, after the frequency response ratio value is nonlinearized, the third dual-microphone beam is filtered based on the frequency response ratio value to obtain a first target signal. Specifically, the calculated frequency response ratio value ratio is nonlinearized within the range of [0, 1], including using a sigmoid function for nonlinearization. Nonlinearizing the frequency response ratio value can enhance the distinction between the target speech and interfering noise, while also accelerating the response speed for determining the difference between the target speech and the interfering noise. Then, the frequency response ratio value is used as a filter to perform a first filtering process on the target beam to obtain the first target signal.

[0040] Step S103: constructing a post filter according to the priori signal-to-noise ratio of the current frame, and performing a second filtering process on the first target signal according to the post filter.

[0041] In one embodiment, the a priori signal-to-noise ratio of the current frame is calculated based on the first target signal, the signal-to-noise ratio of the previous frame, and the dual-microphone beam frequency response, where the dual-microphone beam frequency response is the larger of the first frequency response and the second frequency response calculated in step S102. Specifically, the a priori signal-to-noise ratio of the current frame is calculated according to the following formula:

[0042] snr =alpha* (y – n) / n + (1-alpha)snr_old

[0043] Wherein, y represents the first target signal, n represents the dual-microphone beam frequency response, snr_old represents the signal-to-noise ratio of the previous frame, and alpha represents the third forgetting factor. Preferably, alpha is 0.02.

[0044] The post-filter is constructed based on the prior signal-to-noise ratio of the current frame, specifically:

[0045] filterpost = (snr) / (snr + 1)

[0046] Where filterpost represents the postfilter, and snr represents the a priori signal-to-noise ratio of the current frame. The present invention uses a postfilter to perform a second filtering process on the main microphone signal (i.e., the first target signal), further improving the noise reduction effect.

[0047] The main microphone of this invention is a directional microphone. Its characteristic is that it primarily picks up sound from the direction of the wearer's speech while also suppressing external noise to a certain extent. The main microphone signal (i.e., the first target signal) that undergoes the initial filtering process in step S102 is used as the noisy signal in the post-filter. The other two auxiliary microphones of this invention are omnidirectional, meaning they can pick up sound from all directions. In step S102, the beamforming frequency responses (i.e., the first and second frequency responses) are calculated for the 0° and +180° directions, respectively. The larger of the first and second frequency responses is used as the noise estimate in the post-filter.

[0048] Step S104: Adjust the post-filter according to the long-time filtering energy difference ratio and the short-time filtering energy difference ratio of the sound signal collected by the main microphone; the long-time filtering energy difference ratio refers to the energy difference ratio updated once every first preset number of frames, and the short-time filtering energy difference ratio refers to the energy difference ratio updated once every second preset number of frames, the first preset number of frames is greater than the second preset number of frames, and the energy difference ratio is calculated based on the energy cumulative value of the sound signal collected by the main microphone before the first filtering process and the energy cumulative value after the second filtering process.

[0049] As one embodiment, adjusting the post filter according to the long-time filtering energy difference ratio and the short-time filtering energy difference ratio includes the following steps:

[0050] When the long-term filtered energy difference ratio is less than or equal to a first preset threshold, reducing the filtering degree of the post-filter;

[0051] When the long-time filtering energy difference ratio is greater than the first preset threshold and the short-time filtering energy difference ratio is less than or equal to the second preset threshold, reducing the filtering degree of the post-filter;

[0052] When the long-time filtering energy difference ratio is greater than the first preset threshold and the short-time filtering energy difference ratio is greater than the second preset threshold, the filtering degree of the post-filter is increased.

[0053] Because the noise is largely suppressed after the two filtering processes in steps S102 and S103, the energy difference before and after filtering is large. Meanwhile, the wearer's voice is largely preserved, with a small energy difference before and after filtering. Therefore, the present invention utilizes this characteristic to further modify the filtering level of the post-filter. Specifically, if the energy difference before and after filtering is large, it is determined to be a noise segment, and the filtering level of the post-filter can be appropriately increased. Conversely, it is determined to be the wearer's voice, and the filtering level of the post-filter can be appropriately reduced.

[0054] As one embodiment, the short-term filtering energy difference and the long-term filtering energy difference of the current frame are calculated according to the following formula:

[0055] P(t)=delta*[(OrignalSum(t)-AfterfilterSum(t)) / OrignalSum(t)]+(1-delta)*H(t)

[0056] When calculating the short-time filtering energy difference ratio of the current frame, t represents the second preset number of frames, which is 10 frames; P(10) represents the short-time filtering energy difference ratio of the current frame, OrignalSum(10) represents the energy accumulation value of the sound signal collected by the main microphone 10 frames before the first filtering process, AfterfilterSum(10) represents the energy accumulation value of the sound signal collected by the main microphone 10 frames after the second filtering process, delta represents the first forgetting factor, which is 0.8; H(10) represents the short-time filtering energy difference ratio 10 frames ago.

[0057] When calculating the long-time filtering energy difference ratio of the current frame, t represents the first preset number of frames, which is 50 frames; P(50) represents the long-time filtering energy difference ratio of the current frame, OrignalSum(50) represents the energy accumulation value of the sound signal collected by the main microphone 50 frames before the first filtering process, AfterfilterSum(50) represents the energy accumulation value of the sound signal collected by the main microphone 50 frames after the second filtering process, delta represents the second forgetting factor, which is 0.2; H(50) represents the long-time filtering energy difference ratio 50 frames ago.

[0058] The present invention calculates the long-term and short-term filter energy difference ratios to perform long- and short-term front-to-back energy difference tracking. Long-term front-to-back energy difference tracking is used to preserve the wearer's speech information, while short-term front-to-back energy difference tracking is used to further eliminate short, rapid noise. Long-term and short-term refer to different rates at which the front-to-back energy difference ratio is updated. For example, the front-to-back energy difference ratio is updated every 10 frames, which is the short-term filter energy difference ratio; the front-to-back energy difference ratio is updated every 50 frames, which is the long-term filter energy difference ratio.

[0059] The present invention utilizes multiple beamforming frequency responses in the target direction and at the maximum angle from the target direction to calculate the frequency response ratio, and after performing nonlinear processing on the frequency response ratio, constructs a filter based on the frequency response ratio to perform noise reduction. Compared with the existing technology, the present invention has a higher accuracy rate in recognizing the wearer's speech, better noise cancellation of ambient noise, especially interfering human voices in open offices, and can solve problems such as the lack of high frequency in traditional beamforming and dependence on the wearer's headphone wearing angle. The present invention utilizes the beamforming frequency response at the maximum angle from the target direction as the noise estimate of the post-filter, and uses the long-time and short-time energy difference before and after filtering to correct the post-filter; compared with the traditional post-filter, it can more accurately recognize the wearer's voice information, avoid the loss of target signals, and improve the noise filtering effect.

[0060] Compared with existing noise reduction algorithms, this invention can effectively filter out interfering noise in open office environments, especially interfering human voices around (existing voice noise reduction algorithms have difficulty filtering out interfering human voices), while also improving the wearer's voice quality.

[0061] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

[0062] Those skilled in the art will appreciate that all or part of the processes in the above embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

Claims

1. A noise reduction method applied to three microphones, characterized in that: The following steps are involved: The main microphone collects sound signals to form a target beam, and the right and left feedforward sub-microphones collect two sound signals to form dual-microphone beams in different directions. The dual-microphone beams include a first dual-microphone beam in the direction of the right feedforward sub-microphone, a second dual-microphone beam in the direction of the left feedforward sub-microphone, and a third dual-microphone beam in the direction of the main microphone. respectively calculating the frequency responses of the first dual-microphone beam, the second dual-microphone beam, and the third dual-microphone beam to obtain a first frequency response, a second frequency response, and a third frequency response, calculating a frequency response ratio according to the first frequency response, the second frequency response, and the third frequency response, and performing a first filtering process on the target beam according to the frequency response ratio to obtain a first target signal; constructing a post-filter according to a priori signal-to-noise ratio of the current frame, and performing a second filtering process on the first target signal according to the post-filter; Adjusting the post filter according to the long-time filter energy difference ratio and the short-time filter energy difference ratio of the sound signal collected by the main microphone; The long-time filtering energy difference ratio refers to an energy difference ratio updated once every first preset number of frames, and the short-time filtering energy difference ratio refers to an energy difference ratio updated once every second preset number of frames. The first preset number of frames is greater than the second preset number of frames. The energy difference ratio is calculated based on the energy cumulative value of the sound signal collected by the main microphone before the first filtering processing and the energy cumulative value after the second filtering processing.

2. The noise reduction method for three microphones according to claim 1, characterized in that: The frequency response ratio is calculated according to the first frequency response, the second frequency response and the third frequency response, specifically:

3. The noise reduction method for three microphones according to claim 2, characterized in that: After the frequency response ratio value is subjected to nonlinear processing, the third dual-microphone beam is filtered according to the frequency response ratio value to obtain a first target signal.

4. The noise reduction method for three microphones according to claim 3, characterized in that: The short-term filter energy difference and long-term filter energy difference of the current frame are calculated according to the following formula: P(t)=delta*[(OrignalSum(t)-AfterfilterSum(t)) / OrignalSum(t)]+(1-delta)*H(t) When calculating the short-time filtering energy difference ratio of the current frame, P(t) represents the short-time filtering energy difference ratio of the current frame, t represents the second preset frame number, OrignalSum(t) represents the energy accumulation value of the sound signal collected by the main microphone t frames before the first filtering process, AfterfilterSum(t) represents the energy accumulation value of the sound signal collected by the main microphone t frames after the second filtering process, delta represents the second forgetting factor, and H(t) represents the short-time filtering energy difference ratio t frames before; When calculating the long-term filtered energy difference ratio of the current frame, P(t) represents the long-term filtered energy difference ratio of the current frame, t represents the first preset frame number, OrignalSum(t) represents the cumulative energy value of the sound signal collected by the main microphone t frames before the first filtering process, AfterfilterSum(t) represents the cumulative energy value of the sound signal collected by the main microphone t frames after the second filtering process, delta represents the first forgetting factor, and H(t) represents the long-term filtered energy difference ratio t frames before.

5. The noise reduction method for three microphones according to claim 4, characterized in that: Adjusting the post filter according to the long-time filtering energy difference ratio and the short-time filtering energy difference ratio comprises the following steps: When the long-term filtered energy difference ratio is less than or equal to a first preset threshold, reducing the filtering degree of the post-filter; When the long-time filtering energy difference ratio is greater than the first preset threshold and the short-time filtering energy difference ratio is less than or equal to the second preset threshold, reducing the filtering degree of the post-filter; When the long-time filtering energy difference ratio is greater than the first preset threshold and the short-time filtering energy difference ratio is greater than the second preset threshold, the filtering degree of the post-filter is increased.

6. The noise reduction method for three microphones according to claim 5, characterized in that: The priori signal-to-noise ratio of the current frame is calculated based on the first target signal, the signal-to-noise ratio of the previous frame, and the dual-microphone beam frequency response, where the dual-microphone beam frequency response is the larger frequency response of the first frequency response and the second frequency response.

7. The noise reduction method applied to three microphones according to claim 6, characterized in that: The priori signal-to-noise ratio of the current frame is calculated according to the first target signal, the signal-to-noise ratio of the previous frame, and the dual-microphone beam frequency response, specifically: snr=alpha*(y–n) / n+(1-alpha)snr_old Wherein, y represents the first target signal, n represents the dual-microphone beam frequency response, snr_old represents the signal-to-noise ratio of the previous frame, and alpha represents the second forgetting factor.

8. The noise reduction method applied to three microphones according to any one of claims 1 to 7, characterized in that: The post-filter is constructed based on the prior signal-to-noise ratio of the current frame, specifically: filterpost=(snr) / (snr+1) Wherein, filterpost represents the post filter, and snr represents the prior signal-to-noise ratio of the current frame.

Citation Information

Patent Citations

  • Dual-microphone noise reduction method with adjustable expected sound source direction

    CN114724574A

  • Sound collection device

    JP2016046769A