A speech noise reduction method, system, device and medium based on differential beamformer
By using the theoretical beam pattern and Chebyshev polynomials based on the differential beamformer, combined with the zero-point constraint method and the combined delay-sum beamformer, the problem of difficult control of the beam mainlobe width and sidelobe amplitude is solved, and flexible beam steering and effective speech noise reduction effects are achieved.
Patent Information
- Application Number
- CN202510109472.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Existing differential beamforming technology has difficulty in accurately controlling the main lobe width and side lobe amplitude of the beam in speech signal processing, and cannot perform arbitrary angle adjustment without beam pattern distortion, and cannot effectively suppress background noise.
Based on the theoretical beam pattern of the differential beamformer, the relationship between the mainlobe width and sidelobe amplitude is derived using Chebyshev polynomials to construct an ideal beam pattern that can be precisely controlled. Different beamformers are designed using the zero-point constraint method, diagonal loading method and combined delay-sum beamformer to achieve flexible steering and robust noise reduction of the microphone array.
It achieves precise control of the main lobe width and side lobe amplitude of the beam, enhances the flexibility and adaptability of the system, and can adjust the direction at any angle without beam pattern distortion. It is suitable for microphone arrays of any formation and can effectively reduce voice noise.
Smart Images

Figure CN119943077B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound signal processing, and in particular to a speech noise reduction method, system, device and medium based on a differential beamformer. Background Art
[0002] A beamformer is essentially a spatial filter that leverages the spatial information contained in signals collected by multiple sensors to extract signal sources in a specific direction while suppressing noise and interference sources in other directions. The earliest beamforming technologies originated in radar, sonar, antennas, and other fields, primarily processing narrowband signals. However, given that speech signals are typically wideband, traditional beamforming techniques cannot be directly applied to microphone arrays for speech processing. To address this, researchers have proposed a variety of beamforming techniques suitable for speech signal processing, including nested microphone array-based beamforming, model-based beamforming, super-directional beamforming, and differential beamforming. Among these beamforming techniques, differential beamforming has attracted widespread attention and application due to its advantages: 1. a frequency-invariant beam pattern; 2. high directivity; and 3. compact size, making it easily integrated into portable devices.
[0003] Despite the aforementioned advantages, differential beamforming technology still presents some challenges. For example, when the order is low, the main lobe width of the beam pattern is relatively wide. However, in some applications, a narrower main lobe is required to better capture the sound source signal. Another example is that current differential array design methods cannot precisely control the amplitude of the beamformer's side lobes. However, in some speech noise reduction applications, it may be undesirable to introduce excessive distortion to the background noise field. In this case, it is necessary to control the amplitude of each side lobe to ensure consistency.
[0004] In the paper "On the Robustness of the Superdirective Beamformer," Chen Xi et al. proposed a differential beamformer that can precisely control the robustness of the beamformer. Luo Xueqin proposed a differential beamformer capable of full spatial steerability in the paper "Design of fully steerable broadband beamformers with concentric circular superarrays." However, neither of these two beamformers can precisely control the mainlobe width and sidelobe amplitude of the beamformer. Summary of the Invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a speech noise reduction method, system, device and medium based on a differential beamformer. The method is based on the theoretical beam pattern of the differential beamformer and uses Chebyshev polynomials to derive the relationship between the mainlobe width, sidelobe amplitude and the order of the Chebyshev polynomials, thereby constructing an ideal beam pattern that can accurately control the mainlobe width and sidelobe amplitude; based on the constructed ideal beam pattern, different beamformers are constructed using a zero-point constraint method, a diagonal loading method and a combined delay-sum beamformer; these different beamformers can all accurately control the mainlobe width and sidelobe amplitude of the beam, can be steered at any angle without beam pattern distortion, are applicable to microphone arrays of any formation, are relatively robust to microphone self-noise, and can effectively perform speech noise reduction; and have good applicability and robustness.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A speech noise reduction method based on a differential beamformer comprises the following steps:
[0008] Step 1: Collect a noisy speech signal and convert it into a short-time Fourier transform domain to obtain the spectrum of the noisy speech signal;
[0009] Step 2: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector for the microphone array;
[0010] Step 3: Determine the first ideal beam pattern based on the actual noise reduction level required and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern using Chebyshev polynomials; adjust the angle of the second ideal beam pattern to the incident azimuth angle of the sound source in step 2 A third ideal beam pattern is obtained;
[0011] Step 4: Based on the steering vector of the microphone array in step 2, the second ideal beam pattern and the third ideal beam pattern in step 3, the first beamformer and the second beamformer are designed respectively using the zero-point constraint method; based on the second beamformer, the third beamformer is designed using the minimum norm method and diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer with the delay-sum beamformer;
[0012] Step 5: Based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in step 4, combined with the observation signal of the microphone array, determine a frequency domain estimate of the clean speech signal based on the observation signal of the microphone array;
[0013] Step 6: Based on the frequency domain estimation value of the clean speech signal in step 5, the frequency domain estimation value of the clean speech signal is transformed into the time domain using the overlap-add method or the overlap-save method to obtain the speech signal after noise reduction processing.
[0014] Furthermore, in step 2, the incident pitch angle θ of the sound source is s and the incident azimuth of the sound source The steering vector of the microphone array is determined as follows:
[0015] Based on the incident pitch angle θ of the sound source s and the incident azimuth The position of the mth microphone is expressed using Cartesian coordinates:
[0016]
[0017] Among them, θ m represents the elevation angle of the mth microphone; represents the azimuth angle of the mth microphone; r m It represents the distance between the mth microphone and the origin of the coordinate system; (·) T Represents the transpose operation of a vector or matrix; p m is the position of the mth microphone;
[0018] Based on the incident pitch angle θ of the sound source s and the incident azimuth Get the wave number of the sound source:
[0019]
[0020] Where c = represents the speed of sound; ω represents the angular frequency; f represents the frequency; θ s Indicates the pitch angle of the incident direction of the sound source; Indicates the azimuth of the incident direction of the sound source; k s is the wave number of the sound source;
[0021] Based on the position of the mth microphone and the wave number of the sound source, obtain the steering vector of the microphone array:
[0022]
[0023] Among them, k s is the wave number of the sound source; p mis the position of the mth microphone, and M represents the total number of microphones in the microphone array; (·) T Represents the transpose operation of a vector or matrix; = is the steering vector of the microphone array.
[0024] Furthermore, the step 3 is specifically as follows:
[0025] Step 3.1: Determine the order N of the beamformer based on the actual noise reduction requirements;
[0026] Step 3.2: Based on the order N in step 3.1, combined with the theoretical beam pattern of the differential beamformer, we can obtain The first ideal beam pattern at time:
[0027]
[0028] Among them: a N,n (n=0,1,...,N) represents the real coefficient of the beam pattern; N represents the order of the beamformer; n represents the nth coefficient;
[0029] Step 3.3: Based on the order N in step 3.1, define the Nth order Chebyshev polynomial:
[0030]
[0031] Among them: cosh -1 x and coshx are inverse functions of each other; N represents the order of the beamformer;
[0032] Step 3.4: Based on the Nth order Chebyshev polynomial in step 3.3, let the main lobe amplitude correspond to The value of , and x0>1, the ratio of the main lobe amplitude to the side lobe amplitude is defined as:
[0033]
[0034] Where: L m Indicates the main lobe amplitude; L s represents the side lobe amplitude; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0035] Step 3.5: Based on the Nth-order Chebyshev polynomial in step 3.3 and the ratio of the main lobe amplitude to the side lobe amplitude in step 3.4, we obtain:
[0036]
[0037] Where: x0 represents the location of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0038] Step 3.6: Based on the entire interval [-1, x0] of the Nth order Chebyshev polynomial in step 3.3, let
[0039]
[0040] Where: c1 and c2 are constants; Indicates azimuth;
[0041] Step 3.7: Based on the mapping relationship between π and x, by solving c1+c2=x0 and -c1+c2=-1, we can obtain:
[0042]
[0043] Where: c1 and c2 are constants; x0 represents the position of the main lobe of the beam;
[0044] The mapping relationship between π and x is as follows:
[0045]
[0046] Where: π represents 180°; represents the azimuth; c1 and c2 are both constants;
[0047] Step 3.8: Substitute equation (11) into equation (10) to obtain:
[0048]
[0049] Where: x0 represents the location of the main lobe of the beam; - indicates the azimuth;
[0050] Step 3.9: Based on equations (8), (9), (12) and the first ideal beam pattern in step 3.2, we obtain The second ideal beam pattern is:
[0051]
[0052] Where: R represents the ratio of the main lobe amplitude to the side lobe amplitude; x0 represents the position of the main lobe of the beam; Indicates the azimuth; at this time, within the scope of The root is: exist within the scope of The root is:
[0053] Step 3.10: Adjust the angle of the second ideal beam pattern in step 3.9 to the angle in step 2.1. The third ideal beam pattern is obtained:
[0054]
[0055] in: Indicates the incident azimuth of the sound source; - represents the azimuth; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0056] Furthermore, the step 4 is specifically as follows:
[0057] Step 4.1: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraint h H (ω)d(ω,90°,0)=1, using zero point constraint, i.e. constraint The beam pattern in the direction is 0, Beam pattern The root of θ s =90° first linear system:
[0058]
[0059] In formula (23), Defined as:
[0060]
[0061] Where: θ0=[θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; i M Represents the identity matrix I of size M×M M The first column of
[0062] Step 4.2: Based on the first linear system in step 4.1 and the second ideal beam pattern in step 3.9, we get θ s The first beamformer when =90°:
[0063]
[0064] in: Representation matrix The inverse matrix of
[0065] Step 4.3: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraints Set the direction of zero to Using zero point constraints, we get The second linear system:
[0066]
[0067] In formula (17), Defined as:
[0068]
[0069] Where: i M Represents the identity matrix I of size M×M M The first column of
[0070] θ0=[θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array;
[0071] Step 4.4: Based on the second linear system in step 4.3 and the third ideal beam pattern in step 3.10, we get The second beamformer:
[0072]
[0073] in: Representation matrix The inverse matrix of
[0074] Step 4.5: Assume that M > 2N + 1. Based on the second beamformer in step 4.4, use the minimum norm method and diagonal loading to obtain the third beamformer:
[0075]
[0076] in: Representation matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of ; η represents a small scalar (positive number); the value of η can be obtained by solving the following function:
[0077]
[0078] Where: ∈ represents a small scalar (positive number) that can be set freely; h DL (ω) represents the third beamformer; represents the third beamformer h DL The conjugate transpose of (ω);
[0079] Step 4.6: Based on the second beamformer in step 4.4, combine the delay-sum beamformer to obtain beamformer 4:
[0080]
[0081] Where: μ is a freely settable constant, 0≤μ≤1; Representation matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of .
[0082] A speech noise reduction system based on a differential beamformer, comprising:
[0083] Speech signal acquisition and spectrum calculation module: collects noisy speech signals and converts them into the short-time Fourier transform domain to obtain the spectrum of the noisy speech signals;
[0084] Steering vector determination module: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector for the microphone array;
[0085] Ideal beam pattern determination module: Determines the first ideal beam pattern based on the actual noise reduction level and the theoretical beam pattern of the differential beamformer; determines the second ideal beam pattern based on the first ideal beam pattern using Chebyshev polynomials; and adjusts the angle of the second ideal beam pattern to the incident azimuth of the sound source. A third ideal beam pattern is obtained;
[0086] Beamformer Design Module: Based on the microphone array's steering vector, the second ideal beam pattern, and the third ideal beam pattern, the first and second beamformers are designed using the zero-point constraint method. Based on the second beamformer, the third beamformer is designed using the minimum norm method and diagonal loading. Based on the second beamformer, the fourth beamformer is designed using a delay-sum beamformer.
[0087] A frequency domain clean speech signal estimation module is configured to determine a frequency domain estimate of the clean speech signal based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in combination with an observation signal from the microphone array.
[0088] Time-domain clean speech signal calculation module: Based on the frequency-domain estimate of the clean speech signal, the overlap-add or overlap-save method is used to transform the frequency-domain estimate of the clean speech signal into the time domain to obtain the denoised speech signal.
[0089] A speech noise reduction device based on a differential beamformer, comprising:
[0090] Memory: used for storing a computer program to implement the above-mentioned speech noise reduction method based on a differential beamformer;
[0091] Processor: configured to implement the above-mentioned speech noise reduction method based on a differential beamformer when executing the computer program.
[0092] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the speech noise reduction methods based on a differential beamformer.
[0093] Compared with the prior art, the present invention has the following beneficial effects:
[0094] 1. Based on the theoretical beam pattern of a differential beamformer, the present invention utilizes Chebyshev polynomials to construct a second ideal beam pattern and a third ideal beam pattern that can precisely control the mainlobe width and sidelobe amplitude. The second ideal beam pattern and the third ideal beam pattern can precisely control the mainlobe beam width and sidelobe amplitude, thereby enhancing the flexibility and adaptability of the system.
[0095] 2. Based on the constructed ideal beam pattern, the present invention designs different beamformers using a zero-point constraint method, a diagonal loading method, and a combined delay-sum beamformer. The first beamformer, the second beamformer, the third beamformer, and the fourth beamformer can all precisely control the main lobe width and sidelobe amplitude of the beam, can be steered at any angle without distortion, are applicable to microphone arrays of any formation, are relatively robust to the self-noise of the microphone, and can effectively perform speech noise reduction.
[0096] In summary, based on the theoretical beam pattern of the differential beamformer, the present invention utilizes Chebyshev polynomials to derive the relationship between the mainlobe width, sidelobe amplitude, and the order of the Chebyshev polynomials, and then constructs an ideal beam pattern that can accurately control the mainlobe width and sidelobe amplitude. Based on the constructed ideal beam pattern, different beamformers are constructed using a zero-point constraint method, a diagonal loading method, and a combined delay-sum beamformer. These different beamformers can accurately control the mainlobe width and sidelobe amplitude of the beam, can be steered at any angle without beam pattern distortion, are applicable to microphone arrays of any formation, are relatively robust to the self-noise of the microphone, can effectively perform speech noise reduction, and have good applicability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] Figure 1 This is a flow chart of the speech noise reduction method based on the differential beamformer of the present invention.
[0098] Figure 2 Schematic diagram of the stereo microphone array structure of the present invention.
[0099] Figure 3 For the present invention, N=3, θ s =90°, the desired beam pattern corresponding to different ratios of main lobe width to side lobe amplitude R; where, Figure 3 (a) in the equation is N=3, θ s =90°, R = 5 dB; Figure 3 (b) in the equation is N=3, θ s =90°, R = 10 dB; Figure 3 (c) in the equation is N=3, θ s =90°, R = 20 dB; Figure 3 (d) in is N=3, θ s =90°, and the desired beam pattern corresponding to R = 30 dB.
[0100] Figure 4 B of the present invention NN =π / 3, θ s =90°, the desired beam patterns corresponding to different orders N; where, Figure 4 (a) in the equation is B NN =π / 3, θ s =90°, and the desired beam pattern corresponding to N=4; Figure 4 (b) in the equation is B NN =π / 3, θ s =90°, and the desired beam pattern corresponding to N=5; Figure 4 (c) in the equation is B NN =π / 3, θ s =90°, N=6 corresponding to the desired beam pattern; Figure 4 (d) in the equation is B NN =π / 3, θ s =90°, and the desired beam pattern corresponding to N=7.
[0101] Figure 5 For the present invention, N=3, M=7, R=30 decibels, θ s =90°, beam pattern of the third beamformer when the frequencies are different; where, Figure 5 (a) in the equation is N=3, M=7, R=30dB. θ s =90°, beam pattern of the third beamformer when f = 500 Hz; Figure 5 (b) in the equation is N=3, M=7, R=30dB. θ s =90°, beam pattern of the third beamformer when f = 1000 Hz; Figure 5 (c) in the figure is N=3, M=7, R=30dB, r=0.02m, θ s =90°, beam pattern of the third beamformer when f = 2000 Hz; Figure 5 (d) in the equation is N=3, M=7, R=30dB, θ s Beam pattern of the third beamformer when f = 90° and f = 4000 Hz.
[0102] Figure 6 For the present invention, N=3, M=7, R=30dB, f=1000Hz, θ s =90°, incident azimuth The beam pattern corresponding to the second beamformer at different times; wherein, Figure 6 (a) is N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the second beamformer when ; Figure 6 (b) in the figure is N=3, M=7, R=30dB, f=1000Hz, θ s =90°, The beam pattern corresponding to the second beamformer when ; Figure 6 (c) in the figure is N=3, M=7, R=30dB, f=1000Hz, θ s =90°, The beam pattern corresponding to the second beamformer when ; Figure 6 (d) in the figure is N=3, M=7, R=30dB, f=1000Hz, θ s =90°, The beam pattern corresponding to the second beamformer when .
[0103] Figure 7 For the present invention, N=3, M=7, R=30dB, f=1000Hz, θ s =90°, incident azimuth The beam patterns corresponding to the third beamformer at different times; wherein, Figure 7 (a) is N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the third beamformer when ; Figure 7 (b) in the figure is N=3, M=7, R=30dB, f=1000Hz, θ s =90°, The beam pattern corresponding to the third beamformer when ; Figure 7 (c) in the figure is N=3, M=7, R=30dB, f=1000Hz, θ s =90°, The beam pattern corresponding to the third beamformer when ; Figure 7 (d) in the figure is N=3, M=7, R=30dB, f=1000Hz, θ s =90°, The beam pattern corresponding to the third beamformer when . DETAILED DESCRIPTION
[0104] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0105] See also Figure 1 , a speech noise reduction method based on a differential beamformer, comprising the following steps:
[0106] Step 1: Collect a noisy speech signal and convert it into a short-time Fourier transform domain to obtain the spectrum of the noisy speech signal;
[0107] Step 2: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector for the microphone array;
[0108] Figure 2 The schematic diagram of the stereo microphone array structure of this embodiment is as follows. The microphone array consists of M omnidirectional microphones. The incident pitch angle of the sound source is θ s , the incident azimuth of the sound source is The steering vector of the microphone array is obtained by combining the position of each microphone in the microphone array and the wave number of the sound source.
[0109] In step 2, the incident pitch angle θ of the sound source is s and the incident azimuth of the sound source The steering vector of the microphone array is determined as follows:
[0110] Based on the incident pitch angle θ of the sound source s and the incident azimuth The position of the mth microphone is expressed using Cartesian coordinates:
[0111]
[0112] Among them, θ m represents the elevation angle of the mth microphone; represents the azimuth angle of the mth microphone; r m It represents the distance between the mth microphone and the origin of the coordinate system; (·) T Represents the transpose operation of a vector or matrix; p m is the position of the mth microphone;
[0113] Based on the incident pitch angle θ of the sound source s and the incident azimuth Get the wave number of the sound source:
[0114]
[0115] Where c = represents the speed of sound; ω represents the angular frequency; f represents the frequency; θ s Indicates the pitch angle of the incident direction of the sound source; Indicates the azimuth of the incident direction of the sound source; k s is the wave number of the sound source;
[0116] Based on the position of the mth microphone and the wave number of the sound source, obtain the steering vector of the microphone array:
[0117]
[0118] Among them, k s is the wave number of the sound source; p m is the position of the mth microphone, and M represents the total number of microphones in the microphone array; (·) T Represents the transpose operation of a vector or matrix; = is the steering vector of the microphone array.
[0119] Step 3: Determine the first ideal beam pattern based on the actual noise reduction level required and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern using Chebyshev polynomials; adjust the angle of the second ideal beam pattern to the incident azimuth angle of the sound source in step 2 A third ideal beam pattern is obtained;
[0120] The step 3 is specifically as follows:
[0121] Step 3.1: Determine the order N of the beamformer based on the actual noise reduction requirement. The higher the required noise reduction, the higher the corresponding beamformer order N. In this embodiment, the order N of the beamformer can be flexibly adjusted as needed.
[0122] Step 3.2: Based on the order N in step 3.1, combined with the theoretical beam pattern of the differential beamformer, we can obtain The first ideal beam pattern at time:
[0123]
[0124] Among them: a N,n (n=0,1,…,N) represents the real coefficient of the beam pattern, which determines the shape of the beam pattern; N represents the order of the beamformer; n represents the nth coefficient;
[0125] From equation (4), we can see the beam pattern It's about Symmetric N-order polynomial. The main goal of the design process of the differential microphone array in this embodiment is to convert the actual beam pattern Approximating the theoretical beam pattern The actual beam pattern of the beamformer is:
[0126]
[0127] Where: h(ω) represents the beamformer, represents the actual beam pattern, Indicates that the signal comes from the pitch angle θ and azimuth angle The steering vector of the array.
[0128] In order to precisely control the beam main lobe width and side lobe amplitude, and obtain the narrowest beam main lobe width under a given side lobe amplitude, this embodiment uses Chebyshev polynomials to assist in designing the ideal beam pattern. The actual beam pattern of the beamformer h(ω) Approach
[0129] Step 3.3: Based on the order N in step 3.1, define the Nth order Chebyshev polynomial:
[0130]
[0131] Among them: cosh -1 x and coshx are inverse functions of each other; N represents the order of the beamformer;
[0132] Since Chebyshev polynomials are real coefficient polynomials, when x∈[-1,1], There are N real roots. By taking It can be seen that the root is Place, that is:
[0133]
[0134] From formula (6), we can see that the Chebyshev polynomial has an equiripple characteristic in the interval [-1,1] (the amplitude of the polynomial is limited to ±1, that is, And all polynomials pass through the points (1,1) and (-1 , (-1) N ). When |x|>1,
[0135] Step 3.4: Based on the Nth order Chebyshev polynomial in step 3.3, let the main lobe amplitude correspond to The value of , and x0>1, the ratio of the main lobe amplitude to the side lobe amplitude is defined as:
[0136]
[0137] Where: L m Indicates the main lobe amplitude; L s represents the sidelobe amplitude; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the sidelobe amplitude;
[0138] Step 3.5: Based on the Nth-order Chebyshev polynomial in step 3.3 and the ratio of the main lobe amplitude to the side lobe amplitude in step 3.4, we obtain:
[0139]
[0140] Where: x0 represents the location of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0141] Therefore, the curve The point (x0, R) on the graph indicates the position where the main lobe amplitude is the largest.
[0142] Step 3.6: Based on the entire interval [-1, x0] of the Nth order Chebyshev polynomial in step 3.3, let
[0143]
[0144] Where: c1 and c2 are constants; Indicates azimuth;
[0145] Step 3.7: Based on the mapping relationship between π and x, by solving c1+c2=x0 and -c1+c2=-1, we can obtain:
[0146]
[0147] Where: c1 and c2 are constants; x0 represents the location of the main lobe of the beam; x∈[-1,x0] is related to Therefore, in order to control the entire range of the beam pattern, The mapping relationship between π and x is shown in Table 1:
[0148] Table 1 Mapping relationship between θ and x
[0149]
[0150] Where: π represents 180°; represents the azimuth; c1 and c2 are both constants;
[0151] Step 3.8: Substitute equation (11) into equation (10) to obtain:
[0152]
[0153] Where: x0 represents the location of the main lobe of the beam; - indicates the azimuth;
[0154] Step 3.9: Based on equations (8), (9), (12) and the first ideal beam pattern in step 3.2, we obtain The second ideal beam pattern is:
[0155]
[0156] Where: R represents the ratio of the main lobe amplitude to the side lobe amplitude; x0 represents the position of the main lobe of the beam; - indicates the azimuth; at this time, within the scope of The root is: exist within the scope of The root is
[0157] Step 3.10: Adjust the angle of the second ideal beam pattern in step 3.9 to The third ideal beam pattern is obtained:
[0158]
[0159] in: Indicates the incident azimuth of the sound source; represents the azimuth angle; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude.
[0160] Through the following verification, it is shown that when the order N is determined, the third ideal beam pattern can flexibly control the balance between the beam mainlobe width and the sidelobe amplitude.
[0161] From formula (7), we can see that The root is:
[0162]
[0163] Where: N represents the order of the beamformer; k represents the kth root; π represents 180°; express The angle where the kth root of is located;
[0164] Therefore, the roots in x-space (x∈[-1,x0]) are:
[0165]
[0166] Where: x0 represents the location of the main lobe of the beam;
[0167] The corresponding The root is:
[0168]
[0169] in: express The location of the root; = means The angle where the kth root of is located;
[0170] From formula (17), we can know that the main lobe width of the beam is:
[0171]
[0172] in: express x0 represents the location of the main lobe of the beam;
[0173] In formula (18), B NN represents the main lobe width of the beamformer (the range between the first zero points on both sides of the main lobe), and equation (18) can be written as:
[0174]
[0175] When x0>1 and N≥1, we get:
[0176]
[0177] Where: x0 represents the location of the main lobe of the beam; N represents the order of the ideal beam pattern;
[0178] It can be observed that the beam main lobe width B NNIt increases with the increase of R (increase of x0) and decreases with the increase of N. Therefore, the balance between the main lobe width and the side lobe amplitude can be flexibly controlled under the condition of a certain order N. When the order N is determined, the main lobe width of the beam can be changed by adjusting R. When R approaches 1 (i.e., x0→1), the beam width is minimized, which means that the main lobe width and the side lobe amplitude are the same, that is:
[0179]
[0180] Among them: B NN represents the main lobe width of the beamformer; N represents the order of the ideal beam pattern;
[0181] Therefore, among the different first-order differential microphone array beam patterns, the first-order dipole beam pattern with the same main lobe width and side lobe amplitude and a zero point at π / 2 has the smallest main lobe width, that is, B NN =π.
[0182] From formula (19), we can get
[0183]
[0184] When N and B NN When it is determined, R can be obtained by formula (22). However, it should be noted that B NN It will definitely not be narrower than π / N.
[0185] In order to approximate the desired beam pattern, the beamformer is designed using the zero-point constraint method in this embodiment. From equation (4), it can be seen that the desired beam pattern is about However, when designing the beam pattern of the array, it is necessary to consider The entire value range of Therefore, 2N zero points are used in the present invention, namely The first N zeros are The last N zeros are It should be noted that due to Passing point (-1, (-1) N ), so under the condition that the parameter R is constant, there will never be a zero at θ = π. Therefore, there are no high-order (greater than 1) zeros.
[0186] Step 4: Based on the steering vector of the microphone array in step 2, the second ideal beam pattern and the third ideal beam pattern in step 3, the first beamformer and the second beamformer are designed respectively using the zero-point constraint method; based on the second beamformer, the third beamformer is designed using the minimum norm method and diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer with the delay-sum beamformer;
[0187] The step 4 is specifically as follows:
[0188] Step 4.1: Let M = 2N +1 , based on the steering vector of the microphone array in step 2.2.3, the combined distortion-free constraint h H (ω)d(ω,90°,0)=1, using zero point constraint, i.e. constraint The beam pattern in the direction is 0, Beam pattern The root of θ s =90° first linear system:
[0189]
[0190] In formula (23), Defined as:
[0191]
[0192] Where: θ0=[θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; i M Represents the identity matrix I of size M×M M The first column of
[0193] Step 4.2: Based on the first linear system in step 4.1 and the second ideal beam pattern in step 3.9, we get θ s The first beamformer when =90°:
[0194]
[0195] in: Representation matrix The inverse matrix of
[0196] If you want to steer the main lobe of the beam to The direction of the zero point can be set to The undistortion constraint at this time is
[0197] Step 4.3: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraints Set the direction of zero to Using zero point constraints, we get The second linear system:
[0198]
[0199] In formula (17), Defined as:
[0200]
[0201] Where: i M Represents the identity matrix I of size M×M M The first column of θ0 = [θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array;
[0202] Step 4.4: Based on the second linear system in step 4.3 and the third ideal beam pattern in step 3.10, we get The second beamformer:
[0203]
[0204] in: Representation matrix The inverse matrix of
[0205] When there are many microphones, the matrix The condition number of will be relatively high, and it is easy to be ill-conditioned. The inversion will cause a large error, causing the designed beam pattern to deviate too much from the ideal beam pattern. To solve this problem, this embodiment adopts a diagonal loading method.
[0206] Step 4.5: Assume that M > 2N + 1. Based on the second beamformer in step 4.4, use the minimum norm method and diagonal loading to obtain the third beamformer:
[0207]
[0208] in: Representation matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of ; η represents a small scalar (positive number); the value of η can be obtained by solving the following function:
[0209]
[0210] Where: ∈ represents a small scalar (positive number) that can be set freely; h DL (ω) represents the third beamformer; represents the third beamformer h DL The conjugate transpose of (ω);
[0211] Step 4.6: Based on the second beamformer in step 4.4, combine the delay-sum beamformer to obtain beamformer 4:
[0212]
[0213] Where: μ is a freely settable constant, 0≤μ≤1; Representation matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of .
[0214] Step 5: Determine a frequency domain estimate of the clean speech signal based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in step 4 and the observation signal of the microphone array;
[0215] Taking the third beamformer as an example, when the sound source signal is incident at an angle of θ s and the incident azimuth is When the direction of propagation reaches the microphone array, the observation signal of the microphone array is:
[0216]
[0217] Where: Y m (ω) represents the noisy speech signal collected by the mth microphone, that is, the observation signal; y(ω) represents the observation signal (that is, the noisy signal) vector of the microphone array; X(ω) represents the sound source signal; represents the clean speech signal vector received by the microphone array; v(ω) represents the additive noise signal vector received by the microphone array;
[0218] Based on the observation signal of the microphone array, that is, Equation (31), the frequency domain estimate of the clean speech signal when beamformer 3 is applied is obtained:
[0219]
[0220] Where: Z(ω) is the frequency domain estimate of the sound source signal (clean speech signal) X(ω); superscript (·) H Represents the conjugate transpose operation of a vector or matrix;
[0221] In addition, the first beamformer, the second beamformer, and the fourth beamformer can respectively obtain frequency domain estimation values of corresponding clean speech signals according to the operations of the above steps.
[0222] Step 6: Based on the frequency domain estimation value of the clean speech signal in step 5, the frequency domain estimation value of the clean speech signal is transformed into the time domain using the overlap-add method or the overlap-save method to obtain the speech signal after noise reduction processing.
[0223] A speech noise reduction system based on a differential beamformer, comprising:
[0224] Speech signal acquisition and spectrum calculation module: collects noisy speech signals and converts them into the short-time Fourier transform domain to obtain the spectrum of the noisy speech signals;
[0225] Steering vector determination module: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector for the microphone array;
[0226] Ideal beam pattern determination module: Determines the first ideal beam pattern based on the actual noise reduction level and the theoretical beam pattern of the differential beamformer; determines the second ideal beam pattern based on the first ideal beam pattern using Chebyshev polynomials; and adjusts the angle of the second ideal beam pattern to the incident azimuth of the sound source. A third ideal beam pattern is obtained;
[0227] Beamformer Design Module: Based on the microphone array's steering vector, the second ideal beam pattern, and the third ideal beam pattern, the first and second beamformers are designed using the zero-point constraint method. Based on the second beamformer, the third beamformer is designed using the minimum norm method and diagonal loading. Based on the second beamformer, the fourth beamformer is designed using a delay-sum beamformer.
[0228] A frequency domain clean speech signal estimation module is configured to determine a frequency domain estimate of the clean speech signal based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in combination with an observation signal from the microphone array.
[0229] Time-domain clean speech signal calculation module: Based on the frequency-domain estimate of the clean speech signal, the overlap-add or overlap-save method is used to transform the frequency-domain estimate of the clean speech signal into the time domain to obtain the denoised speech signal.
[0230] A speech noise reduction device based on a differential beamformer, comprising:
[0231] Memory: used for storing a computer program to implement the above-mentioned speech noise reduction method based on a differential beamformer;
[0232] Processor: configured to implement the above-mentioned speech noise reduction method based on a differential beamformer when executing the computer program.
[0233] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the speech noise reduction methods based on a differential beamformer.
[0234] The application effect of the present invention is described in detail below in conjunction with simulation experiments.
[0235] In order to better illustrate the effect of the present invention, the following specific examples are given to verify the effectiveness of the method proposed in the present invention (in the simulation of this part, θ s Always set to 90°).
[0236] In the simulation, the order N of the Chebyshev polynomial is set to 3. Figure 3 Figures (a)-(d) show the desired beam pattern variations for different mainlobe width to sidelobe amplitude ratios, R (unit: dB). It can be seen that for different values of R, the sidelobe amplitudes are precisely controlled, and the mainlobe beamwidth increases with increasing R. When R is small, the mainlobe beamwidth is narrower than that of the super-directional beamformer.
[0237] Figure 4 (a)-(d) in the figure give the main lobe beamwidth B NN Beam patterns corresponding to different polynomial orders N at a certain time. It can be observed that the method in the present invention can accurately control the main lobe beamwidth.
[0238] In the simulation, the order N of the Chebyshev polynomial is set to 3, R=30 decibel, M=7 microphones are taken, a uniform circular array is used and the radius is set to r=0.02m. Figure 5 (a)-(d) in FIG1 show the changes in the desired beam pattern corresponding to the beamformer with different frequency zero-point constraints. It can be seen that the beam pattern of the beamformer designed by the method of the present invention is very close to the desired beam pattern, that is, Figure 3 In addition, if Figure 5 As shown in (b), compared with the beam pattern of the maximum front-to-back ratio beamformer, the third beamformer h DL The sidelobe amplitude of (ω) is precisely controlled and the mainlobe is narrower.
[0239] Using the same array and Chebyshev polynomial parameters as in the above simulation, we investigate the second beamformer h NC (ω) and the third beamformer h DL (ω)’s steering capability. Figure 6(a)-(d) in the figure describe the second beamformer h NC (ω) and the third beamformer h LS (ω) Desired signals in different directions at frequency f = 1000 Hz ( 60°, 90° and 120°); Figure 7 (a)-(d) in the figure describe the third beamformer h DL (ω) Desired signals in different directions at frequency f = 1000 Hz ( 60°, 90° and 120°) corresponding beam patterns; as Figure 6 and Figure 7 As shown, the beam patterns of the second beamformer and the third beamformer proposed in the present invention can be steered at any angle without distortion.
[0240] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A speech noise reduction method based on a differential beamformer, characterized by: The following steps are involved: Step 1: Collect a noisy speech signal and convert it into a short-time Fourier transform domain to obtain the spectrum of the noisy speech signal; Step 2: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector for the microphone array; Step 3: Determine the first ideal beam pattern based on the actual noise reduction level required and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern using Chebyshev polynomials; adjust the angle of the second ideal beam pattern to the incident azimuth angle of the sound source in step 2 A third ideal beam pattern is obtained; Step 4: Based on the steering vector of the microphone array in step 2, the second ideal beam pattern and the third ideal beam pattern in step 3, the first beamformer and the second beamformer are designed respectively using the zero-point constraint method; based on the second beamformer, the third beamformer is designed using the minimum norm method and diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer with the delay-sum beamformer; Step 5: Determine a frequency domain estimate of the clean speech signal based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in step 4 and the observation signal of the microphone array; Step 6: Based on the frequency domain estimation value of the clean speech signal in step 5, the frequency domain estimation value of the clean speech signal is transformed into the time domain using the overlap-add method or the overlap-save method to obtain the speech signal after noise reduction processing.
2. The speech noise reduction method based on a differential beamformer according to claim 1, characterized in that: In step 2, the incident pitch angle θ of the sound source is s and the incident azimuth of the sound source The steering vector of the microphone array is determined as follows: Based on the incident pitch angle θ of the sound source s and the incident azimuth The position of the mth microphone is expressed using Cartesian coordinates: Among them, θ m represents the elevation angle of the mth microphone; represents the azimuth angle of the mth microphone; r m It represents the distance between the mth microphone and the origin of the coordinate system; (·) T Represents the transpose operation of a vector or matrix; p m is the position of the mth microphone; Based on the incident pitch angle θ of the sound source s and the incident azimuth Get the wave number of the sound source: Where c = represents the speed of sound; ω represents the angular frequency; f represents the frequency; θ s Indicates the pitch angle of the incident direction of the sound source; Indicates the azimuth of the incident direction of the sound source; k s is the wave number of the sound source; Based on the position of the mth microphone and the wave number of the sound source, obtain the steering vector of the microphone array: Among them, k s is the wave number of the sound source; p m is the position of the mth microphone, and M represents the total number of microphones in the microphone array; (·) T Represents the transpose operation of a vector or matrix; = is the steering vector of the microphone array.
3. The speech noise reduction method based on differential beamforming according to claim 2, characterized in that: The step 3 is specifically as follows: Step 3.1: Determine the order N of the beamformer based on the actual noise reduction requirements; Step 3.2: Based on the order N in step 3.1, combined with the theoretical beam pattern of the differential beamformer, we can obtain The first ideal beam pattern at time: Among them: a N,n (n=0,1,...,N) represents the real coefficient of the beam pattern; N represents the order of the beamformer; n represents the nth coefficient; Step 3.3: Based on the order N in step 3.1, define the Nth order Chebyshev polynomial: Among them: cosh -1 x and coshx are inverse functions of each other; N represents the order of the beamformer; Step 3.4: Based on the Nth order Chebyshev polynomial in step 3.3, let the main lobe amplitude correspond to The value of , and x0>1, the ratio of the main lobe amplitude to the side lobe amplitude is defined as: Where: L m Indicates the main lobe amplitude; L s represents the side lobe amplitude; R represents the ratio of the main lobe amplitude to the side lobe amplitude; Step 3.5: Based on the Nth-order Chebyshev polynomial in step 3.3 and the ratio of the main lobe amplitude to the side lobe amplitude in step 3.4, we obtain: Where: x0 represents the location of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude; Step 3.6: Based on the entire interval [-1, x0] of the Nth order Chebyshev polynomial in step 3.3, let Where: c1 and c2 are constants; Indicates azimuth; Step 3.7: Based on the mapping relationship between π and x, by solving c1+c2=x0 and -c1+c2=-1, we can obtain: Where: c1 and c2 are constants; x0 represents the position of the main lobe of the beam; The mapping relationship between π and x is as follows: Where: π represents 180°; - represents the azimuth; c1 and c2 are both constants; Step 3.8: Substitute equation (11) into equation (10) to obtain: Where: x0 represents the location of the main lobe of the beam; - indicates the azimuth; Step 3.9: Based on equations (8), (9), (12) and the first ideal beam pattern in step 3.2, we obtain The second ideal beam pattern at time: Where: R represents the ratio of the main lobe amplitude to the side lobe amplitude; x0 represents the position of the main lobe of the beam; Indicates the azimuth; at this time, within the scope of The root is: exist within the scope of The root is Step 3.10: Adjust the angle of the second ideal beam pattern in step 3.9 to the angle in step 2.
1. The third ideal beam pattern is obtained: in: Indicates the incident azimuth of the sound source; - represents the azimuth angle; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude.
4. The speech noise reduction method based on a differential beamformer according to claim 3, characterized in that: The step 4 is specifically as follows: Step 4.1: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraint h H (ω)d(ω,90°,0)=1, using zero point constraint, i.e. constraint The beam pattern in the direction is 0, Beam pattern The root of θ s =90° first linear system: In formula (23), Defined as: Where: θ0=[θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; i M Represents the identity matrix I of size M×M M The first column of Step 4.2: Based on the first linear system in step 4.1 and the second ideal beam pattern in step 3.9, we get θ s The first beamformer when =90°: in: Representation matrix The inverse matrix of Step 4.3: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraints Set the direction of zero to Using zero point constraints, we get The second linear system: In formula (17), Defined as: Where: i M Represents the identity matrix I of size M×M M The first column of θ0 = [θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; Step 4.4: Based on the second linear system in step 4.3 and the third ideal beam pattern in step 3.10, we get The second beamformer: in: Representation matrix The inverse matrix of Step 4.5: Assume that M > 2N + 1. Based on the second beamformer in step 4.4, use the minimum norm method and diagonal loading to obtain the third beamformer: in: Representation matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of ; η represents a small scalar; the value of η can be obtained by solving the following function: Among them: ∈ represents a small scalar that can be set freely; h DL (ω) represents the third beamformer; represents the third beamformer h DL The conjugate transpose of (ω); Step 4.6: Based on the second beamformer in step 4.4, combine the delay-sum beamformer to obtain beamformer 4: Where: μ is a freely settable constant, 0≤μ≤1; Representation matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of .
5. A speech noise reduction system based on a differential beamformer, characterized by: include: Speech signal acquisition and spectrum calculation module: collects noisy speech signals and converts them into the short-time Fourier transform domain to obtain the spectrum of the noisy speech signals; Steering vector determination module: Based on the spectrum of the noisy speech signal, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector for the microphone array; Ideal beam pattern determination module: Determines the first ideal beam pattern based on the actual noise reduction level and the theoretical beam pattern of the differential beamformer; determines the second ideal beam pattern based on the first ideal beam pattern using Chebyshev polynomials; and adjusts the angle of the second ideal beam pattern to the incident azimuth of the sound source. A third ideal beam pattern is obtained; Beamformer Design Module: Based on the microphone array's steering vector, the second ideal beam pattern, and the third ideal beam pattern, the first and second beamformers are designed using the zero-point constraint method. Based on the second beamformer, the third beamformer is designed using the minimum norm method and diagonal loading. Based on the second beamformer, the fourth beamformer is designed using a delay-sum beamformer. A frequency domain clean speech signal estimation module is configured to determine a frequency domain estimate of the clean speech signal based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in combination with an observation signal from the microphone array. Time-domain clean speech signal calculation module: Based on the frequency-domain estimate of the clean speech signal, the overlap-add or overlap-save method is used to transform the frequency-domain estimate of the clean speech signal into the time domain to obtain the denoised speech signal.
6. A speech noise reduction device based on a differential beamformer, characterized by: include: Memory: used for storing a computer program to implement the speech noise reduction method based on a differential beamformer according to any one of claims 1 to 4; Processor: configured to implement the speech noise reduction method based on a differential beamformer as described in any one of claims 1 to 4 when executing the computer program.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech noise reduction method based on a differential beamformer according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Self-adaptive sum and difference beam forming method
CN109799486A
Optimal modal beamformer for sensor arrays
US20120093344A1