Voice noise reduction method, system and device based on differential beam former and medium
By using technical means such as Chebishev polynomial and zero point constraint method in differential beamforming technology, an ideal beam pattern that can accurately control the width and side lobe amplitude of the beam main lobe in the existing technology is solved, and a more flexible and robust speech noise reduction effect is achieved.
Patent Information
- Application Number
- CN202510109472.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
When the existing differential beamforming technology is low in order, the beam main lobe width is wider and the side lobe amplitude cannot be accurately controlled, making it difficult to meet the needs of some voice noise reduction applications for narrower main lobes and consistent side lobe amplitudes.
Through the theoretical beam pattern based on the differential beamformer, the relationship between the main lobe width, side lobe amplitude and the order of the Chebishev polynomial is derived using the Chebishev polynomial to construct an ideal beam pattern that can accurately control the main lobe width and side lobe amplitude. Different beamformers are designed using zero point constraint method, diagonal loading method and combined delay summing beamformers.
It realizes precise control of the beam main lobe width and side lobe amplitude, enhances the flexibility and adaptability of the system, and can adjust the direction at any angle without beam pattern distortion. It is suitable for microphone arrays of any type, which is relatively robust to the microphone's self-noise and effectively performs voice noise reduction.
Smart Images

Figure CN119943077A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound signal processing, and in particular to a speech noise reduction method, system, device and medium based on a differential beamformer. Background Art
[0002] The beamformer is essentially a spatial filter that uses the spatial information contained in the signal collected by multiple sensors to extract the signal source in a specific direction and suppress the noise and interference sources in other directions. The earliest beamforming technology originated from the fields of radar, sonar, antenna, etc., and mainly processed narrowband signals. However, given that speech signals are typical broadband signals, traditional beamforming technology cannot be directly applied to the microphone array field to process speech signals. To this end, researchers have proposed a variety of beamforming technologies suitable for speech signal processing, including beamforming based on nested microphone arrays, model-based beamforming technology, super-directional beamforming technology, and differential beamforming technology. Among the above beamforming technologies, differential beamforming technology has received widespread attention and application due to its following advantages: 1. It has a frequency-invariant beam pattern; 2. It has high directivity; 3. It is small and compact and can be easily integrated into portable devices.
[0003] Although differential beamforming technology has the above advantages, it still has some problems. For example, when the order is low, the main lobe width of its beam pattern is wider. However, in some applications, in order to better collect the sound source signal, a narrower beam main lobe is required. For another example, in the current differential array design method, the amplitude of the beamformer side lobe cannot be accurately controlled. However, in some speech noise reduction applications, it may not be desirable to introduce too much distortion to the background noise field. In this case, it is necessary to control the amplitude of each side lobe to keep it consistent.
[0004] Chen Xi et al. proposed a differential beamformer in the paper “On the Robustness of the Superdirective Beamformer”, which can accurately control the robustness of the beamformer; Luo Xueqin proposed a differential beamformer that can steer in the entire space in the paper “Design of fully steerable broadband beamformers with concentric circular superarrays”; however, both of the above two beamformers cannot accurately control the main lobe width and sidelobe amplitude of the beamformer. Summary of the invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a speech noise reduction method, system, device and medium based on a differential beamformer. The method is based on the theoretical beam pattern of the differential beamformer and uses Chebyshev polynomials to derive the relationship between the main lobe width, the side lobe amplitude and the order of the Chebyshev polynomials, thereby constructing an ideal beam pattern that can accurately control the main lobe width and the side lobe amplitude; based on the constructed ideal beam pattern, different beamformers are constructed using a zero point constraint method, a diagonal loading method and a combined delay summation beamformer; these different beamformers can accurately control the main lobe width and the side lobe amplitude of the beam, can be steered at any angle without beam pattern distortion, can be applied to microphone arrays of any formation, are relatively robust to the self-noise of the microphone, and can effectively perform speech noise reduction; and have good applicability and robustness.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A speech noise reduction method based on a differential beamformer comprises the following steps:
[0008] Step 1: Collect a noisy speech signal and convert the noisy speech signal into a short-time Fourier transform domain to obtain a spectrum of the noisy speech signal;
[0009] Step 2: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector of the microphone array;
[0010] Step 3: Determine the first ideal beam pattern based on the actual required noise reduction degree and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern through the Chebyshev polynomial; adjust the angle of the second ideal beam pattern to the incident azimuth of the sound source in step 2 A third ideal beam pattern is obtained;
[0011] Step 4: Based on the steering vector of the microphone array in step 2, the second ideal beam pattern and the third ideal beam pattern in step 3, the first beamformer and the second beamformer are designed respectively by using the zero-point constraint method; based on the second beamformer, the third beamformer is designed by using the minimum norm method and the diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer;
[0012] Step 5: Based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in step 4, combined with the observation signal of the microphone array; determine the frequency domain estimation value of the clean speech signal based on the observation signal of the microphone array;
[0013] Step 6: Based on the frequency domain estimation value of the clean speech signal in step 5, the frequency domain estimation value of the clean speech signal is transformed into the time domain by using the overlap-add method or the overlap-save method to obtain the speech signal after noise reduction processing.
[0014] Furthermore, in step 2, based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source The steering vector of the microphone array is determined as follows:
[0015] Based on the incident pitch angle θ of the sound source s and the incident azimuth The position of the mth microphone is represented using Cartesian coordinates:
[0016]
[0017] Among them, θ m represents the elevation angle of the mth microphone; represents the azimuth angle of the mth microphone; r m It represents the distance between the mth microphone and the origin of the coordinate system; (·) T Represents the transpose operation of a vector or matrix; p m is the position of the mth microphone;
[0018] Based on the incident pitch angle θ of the sound source s and the incident azimuth Get the wave number of the sound source:
[0019]
[0020] Where c = represents the speed of sound; ω represents the angular frequency; f represents the frequency; θ s Indicates the pitch angle of the incident direction of the sound source; Indicates the azimuth of the incident direction of the sound source; k s is the wave number of the sound source;
[0021] Based on the position of the mth microphone and the wave number of the sound source, obtain the steering vector of the microphone array:
[0022]
[0023] Among them, k s is the wave number of the sound source; p mis the position of the mth microphone, and M represents the total number of microphones in the microphone array; (·) T Represents the transpose operation of a vector or matrix; = is the steering vector of the microphone array.
[0024] Furthermore, the step 3 is specifically as follows:
[0025] Step 3.1: Determine the order N of the beamformer based on the actual noise reduction requirements;
[0026] Step 3.2: Based on the order N in step 3.1, combined with the theoretical beam pattern of the differential beamformer, we can obtain The first ideal beam pattern at :
[0027]
[0028] Among them: a N,n (n=0,1,...,N) represents the real coefficient of the beam pattern; N represents the order of the beamformer; n represents the nth coefficient;
[0029] Step 3.3: Based on the order N in step 3.1, define the N-order Chebyshev polynomial:
[0030]
[0031] Where: cosh -1 x and coshx are inverse functions of each other; N represents the order of the beamformer;
[0032] Step 3.4: Based on the Nth order Chebyshev polynomial in step 3.3, let the main lobe amplitude correspond to The value of x0>1, the ratio of the main lobe amplitude to the side lobe amplitude is defined as:
[0033]
[0034] Where: L m Indicates the main lobe amplitude; L s represents the side lobe amplitude; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0035] Step 3.5: Based on the Nth order Chebyshev polynomial in step 3.3 and the ratio of the main lobe amplitude to the side lobe amplitude in step 3.4, we obtain:
[0036]
[0037] Where: x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0038] Step 3.6: Based on the entire interval [-1, x0] of the N-order Chebyshev polynomial in step 3.3, let
[0039]
[0040] Where: c1 and c2 are constants; Indicates azimuth;
[0041] Step 3.7: Based on the mapping relationship between π and x, by solving c1+c2=x0 and -c1+c2=-1, we get:
[0042]
[0043] Where: c1 and c2 are constants; x0 represents the position of the main lobe of the beam;
[0044] The mapping relationship between π and x is as follows:
[0045]
[0046] Where: π represents 180°; represents the azimuth; c1 and c2 are both constants;
[0047] Step 3.8: Substitute equation (11) into equation (10) to obtain:
[0048]
[0049] Where: x0 represents the position of the main lobe of the beam; - indicates the azimuth;
[0050] Step 3.9: Based on equations (8), (9), (12) and the first ideal beam pattern in step 3.2, we get The second ideal beam pattern is:
[0051]
[0052] Where: R represents the ratio of the main lobe amplitude to the side lobe amplitude; x0 represents the position of the main lobe of the beam; Indicates the azimuth; at this time, within the range of The root is: exist within the range of The root is:
[0053] Step 3.10: Adjust the angle of the second ideal beam pattern in step 3.9 to the angle in step 2.1. The third ideal beam pattern is obtained:
[0054]
[0055] in: Indicates the incident azimuth of the sound source; - represents the azimuth; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0056] Furthermore, the step 4 is specifically as follows:
[0057] Step 4.1: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraint h H (ω)d(ω,90°,0)=1, using zero point constraint, i.e. constraint The beam pattern in the direction is 0, Beam pattern The root of θ s = 90° of the first linear system:
[0058]
[0059] In formula (23), Defined as:
[0060]
[0061] Where: θ0 = [θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; i M Represents the identity matrix I of size M×M M The first column of
[0062] Step 4.2: Based on the first linear system in step 4.1 and the second ideal beam pattern in step 3.9, we get θ s = 90° when the first beamformer:
[0063]
[0064] in: Representation Matrix The inverse matrix of
[0065] Step 4.3: Let M = 2N + 1, based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraints Set the direction of zero point to Using the zero point constraint, we get The second linear system:
[0066]
[0067] In formula (17), Defined as:
[0068]
[0069] Where: i M Represents the identity matrix I of size M×M M The first column of
[0070] θ0=[θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array;
[0071] Step 4.4: Based on the second linear system in step 4.3 and the third ideal beam diagram in step 3.10, we get The second beamformer:
[0072]
[0073] in: Representation Matrix The inverse matrix of
[0074] Step 4.5: Assume that M>2N+1. Based on the second beamformer in step 4.4, the third beamformer is obtained by using the minimum norm method and diagonal loading method:
[0075]
[0076] in: Representation Matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of ; η represents a small scalar (positive number); the value of η can be obtained by solving the following function:
[0077]
[0078] Among them: ∈ represents a small scalar (positive number) set freely; h DL (ω) represents the third beamformer; represents the third beamformer h DL The conjugate transpose of (ω);
[0079] Step 4.6: Based on the second beamformer in step 4.4, the delay-sum beamformer is combined to obtain beamformer 4:
[0080]
[0081] Where: μ is a freely settable constant, 0≤μ≤1; Representation Matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of .
[0082] A speech noise reduction system based on a differential beamformer, comprising:
[0083] Speech signal acquisition and spectrum calculation module: collects noisy speech signals and converts them into short-time Fourier transform domain to obtain the spectrum of noisy speech signals;
[0084] Steering vector determination module: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector of the microphone array;
[0085] Ideal beam pattern determination module: Determine the first ideal beam pattern based on the actual required noise reduction degree and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern through Chebyshev polynomials; adjust the angle of the second ideal beam pattern to the incident azimuth angle of the sound source A third ideal beam pattern is obtained;
[0086] Beamformer design module: Based on the steering vector of the microphone array, the second ideal beam diagram and the third ideal beam diagram, the first beamformer and the second beamformer are designed respectively by using the zero-point constraint method; based on the second beamformer, the third beamformer is designed by using the minimum norm method and diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer;
[0087] A frequency domain clean speech signal estimation module: based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer, combined with the observation signal of the microphone array, determines a frequency domain estimation value of the clean speech signal;
[0088] Time domain clean speech signal calculation module: Based on the frequency domain estimation value of the clean speech signal, the frequency domain estimation value of the clean speech signal is transformed into the time domain by using overlap-add or overlap-preserve method to obtain the speech signal after noise reduction processing.
[0089] A speech noise reduction device based on a differential beamformer, comprising:
[0090] Memory: used for storing a computer program to implement the speech noise reduction method based on a differential beamformer as described above;
[0091] Processor: used to implement the above-mentioned speech noise reduction method based on differential beamformer when executing the computer program.
[0092] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the speech noise reduction methods based on a differential beamformer are implemented.
[0093] Compared with the prior art, the beneficial effects of the present invention are:
[0094] 1. Based on the theoretical beam pattern of the differential beamformer, the present invention uses Chebyshev polynomials to construct a second ideal beam pattern and a third ideal beam pattern that can accurately control the main lobe width and the side lobe amplitude; the main lobe beam width and the side lobe amplitude can be accurately controlled by the second ideal beam pattern and the third ideal beam pattern, thereby enhancing the flexibility and adaptability of the system.
[0095] 2. Based on the constructed ideal beam pattern, the present invention designs different beamformers by adopting the zero-point constraint method, the diagonal loading method and the combined delay-sum beamformer; the first beamformer, the second beamformer, the third beamformer and the fourth beamformer can all accurately control the main lobe width and the side lobe amplitude of the beam, can be steered at any angle without distortion, can be applied to microphone arrays of any formation, are relatively robust to the self-noise of the microphone, and can effectively perform speech noise reduction.
[0096] In summary, based on the theoretical beam pattern of the differential beamformer, the present invention uses Chebyshev polynomials to derive the relationship between the main lobe width, the side lobe amplitude and the order of the Chebyshev polynomials, and then constructs an ideal beam pattern that can accurately control the main lobe width and the side lobe amplitude; based on the constructed ideal beam pattern, different beamformers are constructed using a zero-point constraint method, a diagonal loading method and a combined delay-sum beamformer; these different beamformers can accurately control the main lobe width and the side lobe amplitude of the beam, can be steered at any angle without beam pattern distortion, can be applied to microphone arrays of any formation, are relatively robust to the self-noise of the microphone, can effectively perform speech noise reduction, and have good applicability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] Figure 1 The present invention is a flow chart of a speech noise reduction method based on a differential beamformer.
[0098] Figure 2 Schematic diagram of a stereo microphone array structure of the present invention.
[0099] Figure 3 For the present invention, N=3, θ s =90°, the desired beam pattern corresponding to different ratios of main lobe width to side lobe amplitude R; where, Figure 3 (a) in which N=3, θ s =90°, R = 5 dB; Figure 3 (b) in the equation is N=3, θ s =90°, R = 10 dB; Figure 3 (c) in is N=3, θ s =90°, R = 20 dB; Figure 3 (d) in is N=3, θ s =90°, and the desired beam pattern corresponding to R=30 dB.
[0100] Figure 4 For the present invention B NN =π / 3, θ s =90°, the desired beam patterns corresponding to different orders N; where, Figure 4 (a) in the equation is B NN =π / 3, θ s =90°, the desired beam pattern corresponding to N=4; Figure 4 (b) in the equation is B NN =π / 3, θ s =90°, the desired beam pattern corresponding to N=5; Figure 4 (c) in the equation is B NN =π / 3, θ s =90°, the desired beam pattern corresponding to N=6; Figure 4 (d) in the equation is B NN =π / 3, θ s =90°, and the desired beam pattern corresponding to N=7.
[0101] Figure 5 For the present invention, N=3, M=7, R=30 dB, θ s =90°, beam pattern of the third beamformer when the frequencies are different; where, Figure 5 (a) in the equation is N=3, M=7, R=30 dB, θ s =90°, beam pattern of the third beamformer when f=500Hz; Figure 5 (b) in the figure is N=3, M=7, R=30 dB, θ s =90°, beam pattern of the third beamformer when f=1000 Hz; Figure 5 (c) in which N = 3, M = 7, R = 30 dB, r = 0.02 m, θ s =90°, beam pattern of the third beamformer when f = 2000 Hz; Figure 5 (d) in is N=3, M=7, R=30 dB, θ s =90°, beam pattern of the third beamformer when f=4000 Hz.
[0102] Figure 6 For the present invention, N=3, M=7, R=30 dB, f=1000 Hz, θ s =90°, incident azimuth The beam patterns corresponding to the second beamformer at different times; wherein, Figure 6 (a) in which N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the second beamformer when Figure 6 (b) in the figure is N=3, M=7, R=30 dB, f=1000 Hz, θ s =90°, The beam pattern corresponding to the second beamformer when Figure 6 (c) in the figure is N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the second beamformer when Figure 6 (d) in which N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the second beamformer when .
[0103] Figure 7 For the present invention, N=3, M=7, R=30 dB, f=1000 Hz, θ s =90°, incident azimuth The beam patterns corresponding to the third beamformer at different times; wherein, Figure 7 (a) in which N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the third beamformer when ; Figure 7 (b) in the figure is N=3, M=7, R=30 dB, f=1000 Hz, θ s =90°, The beam pattern corresponding to the third beamformer when ; Figure 7 (c) in the figure is N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the third beamformer when ; Figure 7 (d) in which N = 3, M = 7, R = 30 dB, f = 1000 Hz, θ s =90°, The beam pattern corresponding to the third beamformer when . DETAILED DESCRIPTION
[0104] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0105] See also Figure 1 , a speech noise reduction method based on a differential beamformer, comprising the following steps:
[0106] Step 1: Collect a noisy speech signal and convert the noisy speech signal into a short-time Fourier transform domain to obtain a spectrum of the noisy speech signal;
[0107] Step 2: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector of the microphone array;
[0108] Figure 2 The schematic diagram of the stereo microphone array structure of this embodiment is as follows. The microphone array is composed of M omnidirectional microphones. The incident pitch angle of the sound source is θ s , the incident azimuth of the sound source is The position of each microphone in the microphone array and the wave number of the sound source are combined to obtain the steering vector of the microphone array.
[0109] In step 2, the incident pitch angle θ of the sound source is s and the incident azimuth of the sound source The steering vector of the microphone array is determined as follows:
[0110] Based on the incident pitch angle θ of the sound source s and the incident azimuth The position of the mth microphone is represented using Cartesian coordinates:
[0111]
[0112] Among them, θ m represents the elevation angle of the mth microphone; represents the azimuth angle of the mth microphone; r m It represents the distance between the mth microphone and the origin of the coordinate system; (·) T Represents the transpose operation of a vector or matrix; p m is the position of the mth microphone;
[0113] Based on the incident pitch angle θ of the sound source s and the incident azimuth Get the wave number of the sound source:
[0114]
[0115] Where c = represents the speed of sound; ω represents the angular frequency; f represents the frequency; θ s Indicates the pitch angle of the incident direction of the sound source; Indicates the azimuth of the incident direction of the sound source; k s is the wave number of the sound source;
[0116] Based on the position of the mth microphone and the wave number of the sound source, obtain the steering vector of the microphone array:
[0117]
[0118] Among them, k s is the wave number of the sound source; p m is the position of the mth microphone, and M represents the total number of microphones in the microphone array; (·) T Represents the transpose operation of a vector or matrix; = is the steering vector of the microphone array.
[0119] Step 3: Determine the first ideal beam pattern based on the actual required noise reduction degree and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern through the Chebyshev polynomial; adjust the angle of the second ideal beam pattern to the incident azimuth of the sound source in step 2 A third ideal beam pattern is obtained;
[0120] The step 3 is specifically as follows:
[0121] Step 3.1: Determine the order N of the beamformer based on the actual noise reduction requirement; the higher the required noise reduction amount, the higher the corresponding order N of the beamformer should be. In this embodiment, the order N of the beamformer can be flexibly adjusted as needed.
[0122] Step 3.2: Based on the order N in step 3.1, combined with the theoretical beam pattern of the differential beamformer, we can obtain The first ideal beam pattern at :
[0123]
[0124] Among them: a N,n (n=0,1,…,N) represents the real coefficient of the beam pattern, which determines the shape of the beam pattern; N represents the order of the beamformer; n represents the nth coefficient;
[0125] From equation (4), we can see that the beam pattern About The main goal of the design process of the differential microphone array in this embodiment is to convert the actual beam pattern Approximating the theoretical beam pattern The actual beam pattern of the beamformer is:
[0126]
[0127] Where: h(ω) represents the beamformer, represents the actual beam pattern, Indicates that the signal comes from the pitch angle θ and the azimuth angle The steering vector of the array.
[0128] In order to accurately control the beam main lobe width and side lobe amplitude, and obtain the narrowest beam main lobe width under a given side lobe amplitude, this embodiment uses Chebyshev polynomials to assist in designing an ideal beam pattern. The actual beam pattern of the beamformer h(ω) Approach
[0129] Step 3.3: Based on the order N in step 3.1, define the N-order Chebyshev polynomial:
[0130]
[0131] Where: cosh -1 x and coshx are inverse functions of each other; N represents the order of the beamformer;
[0132] Since Chebyshev polynomials are real coefficient polynomials, when x∈[-1,1], has N real roots. By taking It can be seen that the root is Place, that is:
[0133]
[0134] From equation (6), we can see that the Chebyshev polynomial has an equiripple characteristic in the interval [-1,1] (the amplitude of the polynomial is limited to ±1, that is, And all polynomials pass through the points (1,1) and (-1 , (-1) N ). When |x|>1,
[0135] Step 3.4: Based on the Nth order Chebyshev polynomial in step 3.3, let the main lobe amplitude correspond to The value of x0>1, the ratio of the main lobe amplitude to the side lobe amplitude is defined as:
[0136]
[0137] Where: L m Indicates the main lobe amplitude; L s represents the side lobe amplitude; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0138] Step 3.5: Based on the Nth order Chebyshev polynomial in step 3.3 and the ratio of the main lobe amplitude to the side lobe amplitude in step 3.4, we obtain:
[0139]
[0140] Where: x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude;
[0141] Therefore, the curve The point (x0, R) on the graph indicates the position where the main lobe amplitude is the largest.
[0142] Step 3.6: Based on the entire interval [-1, x0] of the N-order Chebyshev polynomial in step 3.3, let
[0143]
[0144] Where: c1 and c2 are constants; Indicates azimuth;
[0145] Step 3.7: Based on the mapping relationship between π and x, by solving c1+c2=x0 and -c1+c2=-1, we get:
[0146]
[0147] Where: c1 and c2 are constants; x0 represents the location of the main lobe of the beam; x∈[-1,x0] is related to Correspondingly; therefore, in order to control the entire range of the beam pattern, that is, The mapping relationship between π and x is shown in Table 1:
[0148] Table 1 Mapping relationship between θ and x
[0149]
[0150] Where: π represents 180°; represents the azimuth; c1 and c2 are both constants;
[0151] Step 3.8: Substitute equation (11) into equation (10) to obtain:
[0152]
[0153] Where: x0 represents the position of the main lobe of the beam; - indicates the azimuth;
[0154] Step 3.9: Based on equations (8), (9), (12) and the first ideal beam pattern in step 3.2, we get The second ideal beam pattern is:
[0155]
[0156] Where: R represents the ratio of the main lobe amplitude to the side lobe amplitude; x0 represents the position of the main lobe of the beam; - indicates the azimuth; at this time, within the range of The root is: exist within the range of The root is
[0157] Step 3.10: Adjust the angle of the second ideal beam pattern in step 3.9 to The third ideal beam pattern is obtained:
[0158]
[0159] in: Indicates the incident azimuth of the sound source; represents the azimuth; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude.
[0160] Through the following verification, when the order N is determined, the third ideal beam pattern can flexibly control the balance between the beam main lobe width and the side lobe amplitude.
[0161] From formula (7), we can see that The root is:
[0162]
[0163] Where: N represents the order of the beamformer; k represents the kth root; π represents 180°; express The angle where the kth root of is located;
[0164] Therefore, the root in x-space (x∈[-1,x0]) is:
[0165]
[0166] Where: x0 represents the position of the main lobe of the beam;
[0167] The corresponding The root is:
[0168]
[0169] in: express The location of the root; = means The angle where the kth root of is located;
[0170] From formula (17), we can know that the main lobe width of the beam is:
[0171]
[0172] in: express The first root of x0 represents the location of the main lobe of the beam;
[0173] In formula (18), B NN represents the main lobe width of the beamformer (the range between the first zero points on both sides of the main lobe), and equation (18) can be written as:
[0174]
[0175] When x0>1 and N≥1, we get:
[0176]
[0177] Where: x0 represents the location of the main lobe of the beam; N represents the order of the ideal beam pattern;
[0178] It can be observed that the beam main lobe width B NNIt increases with the increase of R (increase of x0) and decreases with the increase of N. Therefore, the balance between the main lobe width and the side lobe amplitude of the beam can be flexibly controlled under the condition that the order N is determined. When the order N is determined, the main lobe width of the beam can be changed by adjusting R. When R approaches 1 (i.e. x0→1), the beam width is the smallest, which means that the main lobe width is the same as the side lobe amplitude, that is:
[0179]
[0180] Among them: B NN represents the main lobe width of the beamformer; N represents the order of the ideal beam pattern;
[0181] Therefore, among different first-order differential microphone array beam patterns, the first-order dipole beam pattern with the same main lobe width and side lobe amplitude and a zero point at π / 2 has the smallest main lobe width, that is, B NN =π.
[0182] From formula (19), we can get
[0183]
[0184] When N and B NN When determined, R can be obtained by formula (22). However, it should be noted that B NN It will definitely not be narrower than π / N.
[0185] In order to approximate the desired beam pattern, the beamformer is designed using the zero-point constraint method in this embodiment. From equation (4), it can be seen that the desired beam pattern is about However, when designing the beam pattern of the array, it is necessary to consider The entire value range of Therefore, 2N zero points are used in the present invention, namely The first N zeros are The last N zeros are It should be noted that due to Passing point (-1, (-1) N ), so under the condition that the parameter R is constant, there will never be a zero at θ = π. Therefore, there are no higher-order (greater than 1) zeros.
[0186] Step 4: Based on the steering vector of the microphone array in step 2, the second ideal beam pattern and the third ideal beam pattern in step 3, the first beamformer and the second beamformer are designed respectively by using the zero-point constraint method; based on the second beamformer, the third beamformer is designed by using the minimum norm method and the diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer;
[0187] The step 4 is specifically as follows:
[0188] Step 4.1: Let M = 2N +1 , based on the steering vector of the microphone array in step 2.2.3, the combined distortion-free constraint h H (ω)d(ω,90°,0)=1, using zero point constraint, i.e. constraint The beam pattern in the direction is 0, Beam pattern The root of θ s = 90° of the first linear system:
[0189]
[0190] In formula (23), Defined as:
[0191]
[0192] Where: θ0 = [θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; i M Represents the identity matrix I of size M×M M The first column of
[0193] Step 4.2: Based on the first linear system in step 4.1 and the second ideal beam pattern in step 3.9, we get θ s = 90° when the first beamformer:
[0194]
[0195] in: Representation Matrix The inverse matrix of
[0196] If you want to steer the main lobe of the beam to In directions other than The undistortion constraint at this time is
[0197] Step 4.3: Let M = 2N + 1, based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraints Set the direction of zero point to Using the zero point constraint, we get The second linear system:
[0198]
[0199] In formula (17), Defined as:
[0200]
[0201] Where: i M Represents the identity matrix I of size M×M M The first column of 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array;
[0202] Step 4.4: Based on the second linear system in step 4.3 and the third ideal beam diagram in step 3.10, we get The second beamformer:
[0203]
[0204] in: Representation Matrix The inverse matrix of
[0205] When there are many microphones, the matrix The condition number of will be relatively high, and it is easy to be ill-conditioned. The inversion will cause a large error, causing the designed beam pattern to deviate too much from the ideal beam pattern. To solve this problem, this embodiment adopts a diagonal loading method.
[0206] Step 4.5: Assume that M>2N+1. Based on the second beamformer in step 4.4, the third beamformer is obtained by using the minimum norm method and diagonal loading method:
[0207]
[0208] in: Representation Matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of ; η represents a small scalar (positive number); the value of η can be obtained by solving the following function:
[0209]
[0210] Among them: ∈ represents a small scalar (positive number) set freely; h DL (ω) represents the third beamformer; represents the third beamformer h DL The conjugate transpose of (ω);
[0211] Step 4.6: Based on the second beamformer in step 4.4, the delay-sum beamformer is combined to obtain beamformer 4:
[0212]
[0213] Where: μ is a freely settable constant, 0≤μ≤1; Representation Matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of .
[0214] Step 5: Based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in step 4, a frequency domain estimation value of the clean speech signal is determined in combination with the observation signal of the microphone array;
[0215] Taking the third beamformer as an example, when the sound source signal is incident at an angle of θ s and the incident azimuth is When the direction of propagation reaches the microphone array, the observation signal of the microphone array is:
[0216]
[0217] Where: Y m (ω) represents the noisy speech signal collected by the mth microphone, that is, the observation signal; y(ω) represents the observation signal (that is, the noisy signal) vector of the microphone array; X(ω) represents the sound source signal; represents the clean speech signal vector received by the microphone array; v(ω) represents the additive noise signal vector received by the microphone array;
[0218] Based on the observation signal of the microphone array, that is, equation (31), the frequency domain estimate of the clean speech signal when beamformer 3 is applied is obtained:
[0219]
[0220] Where: Z(ω) is the frequency domain estimate of the sound source signal (clean speech signal) X(ω); superscript (·) H Represents the conjugate transpose operation of a vector or matrix;
[0221] In addition, the first beamformer, the second beamformer, and the fourth beamformer can respectively obtain frequency domain estimation values of corresponding clean speech signals according to the operations of the above steps.
[0222] Step 6: Based on the frequency domain estimation value of the clean speech signal in step 5, the frequency domain estimation value of the clean speech signal is transformed into the time domain by using the overlap-add method or the overlap-save method to obtain the speech signal after noise reduction processing.
[0223] A speech noise reduction system based on a differential beamformer, comprising:
[0224] Speech signal acquisition and spectrum calculation module: collects noisy speech signals and converts them into short-time Fourier transform domain to obtain the spectrum of noisy speech signals;
[0225] Steering vector determination module: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector of the microphone array;
[0226] Ideal beam pattern determination module: Determine the first ideal beam pattern based on the actual required noise reduction degree and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern through Chebyshev polynomials; adjust the angle of the second ideal beam pattern to the incident azimuth angle of the sound source A third ideal beam pattern is obtained;
[0227] Beamformer design module: Based on the steering vector of the microphone array, the second ideal beam diagram and the third ideal beam diagram, the first beamformer and the second beamformer are designed respectively by using the zero-point constraint method; based on the second beamformer, the third beamformer is designed by using the minimum norm method and diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer;
[0228] A frequency domain clean speech signal estimation module: based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer, combined with the observation signal of the microphone array, determines a frequency domain estimation value of the clean speech signal;
[0229] Time domain clean speech signal calculation module: Based on the frequency domain estimation value of the clean speech signal, the frequency domain estimation value of the clean speech signal is transformed into the time domain by using overlap-add or overlap-preserve method to obtain the speech signal after noise reduction processing.
[0230] A speech noise reduction device based on a differential beamformer, comprising:
[0231] Memory: used for storing a computer program to implement the speech noise reduction method based on a differential beamformer as described above;
[0232] Processor: used to implement the above-mentioned speech noise reduction method based on differential beamformer when executing the computer program.
[0233] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the speech noise reduction methods based on a differential beamformer are implemented.
[0234] The application effect of the present invention is described in detail below in conjunction with simulation experiments.
[0235] In order to better illustrate the effect of the present invention, the following specific examples are given to verify the effectiveness of the method proposed in the present invention (in the simulation of this part, θ s Always set to 90°).
[0236] In the simulation, the order N of the Chebyshev polynomial is set to 3. Figure 3 (a)-(d) in the figure show the changes in the desired beam pattern corresponding to different main lobe width and side lobe amplitude ratios R (unit: dB). It can be seen that when R takes different values, the side lobe amplitude is accurately controlled, and the main lobe beamwidth increases with the increase of R. When R is small, the beamwidth of the main lobe is narrower than the beamwidth of the super-pointing beamformer.
[0237] Figure 4 (a)-(d) in the figure give the main lobe beamwidth B NN Beam patterns corresponding to different polynomial orders N at a certain time. It can be observed that the method in the present invention can accurately control the main lobe beam width.
[0238] In the simulation, the order N of the Chebyshev polynomial is set to 3, R = 30 dB, M = 7 microphones are taken, a uniform circular array is used and the radius is set to r = 0.02 m. Figure 5 (a)-(d) in FIG. 1 show the changes in the desired beam patterns corresponding to the beamformers with different frequency zero-point constraints. It can be seen that the beam pattern of the beamformer designed by the method of the present invention is very close to the desired beam pattern, that is, Figure 3 In addition, if Figure 5 As shown in (b), compared with the beam pattern of the maximum front-to-back ratio beamformer, the third beamformer h DL The sidelobe amplitude of (ω) is precisely controlled and the main lobe is narrower.
[0239] Using the same array and Chebyshev polynomial parameters as in the above simulation, we investigate the second beamformer h NC (ω) and the third beamformer h DL (ω)’s steering capability. Figure 6(a)-(d) in Figure 2 describe the second beamformer h NC (ω) and the third beamformer h LS (ω) Desired signals in different directions at frequency f = 1000 Hz ( 60°, 90° and 120°) corresponding beam patterns; Figure 7 (a)-(d) in Figure 2 describe the third beamformer h DL (ω) Desired signals in different directions at frequency f = 1000 Hz ( 60°, 90° and 120°) corresponding beam patterns; as Figure 6 and Figure 7 As shown, the beam patterns of the second beamformer and the third beamformer proposed in the present invention can be steered at any angle without distortion.
[0240] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A speech noise reduction method based on a differential beamformer, characterized in that: The following steps are involved: Step 1: Collect a noisy speech signal and convert the noisy speech signal into a short-time Fourier transform domain to obtain a spectrum of the noisy speech signal; Step 2: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector of the microphone array; Step 3: Determine the first ideal beam pattern based on the actual required noise reduction degree and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern through the Chebyshev polynomial; adjust the angle of the second ideal beam pattern to the incident azimuth of the sound source in step 2 A third ideal beam pattern is obtained; Step 4: Based on the steering vector of the microphone array in step 2, the second ideal beam pattern and the third ideal beam pattern in step 3, the first beamformer and the second beamformer are designed respectively by using the zero-point constraint method; based on the second beamformer, the third beamformer is designed by using the minimum norm method and the diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer; Step 5: Based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer in step 4, a frequency domain estimation value of the clean speech signal is determined in combination with the observation signal of the microphone array; Step 6: Based on the frequency domain estimation value of the clean speech signal in step 5, the frequency domain estimation value of the clean speech signal is transformed into the time domain by using the overlap-add method or the overlap-save method to obtain the speech signal after noise reduction processing.
2. The speech noise reduction method based on differential beamformer according to claim 1, characterized in that: In step 2, the incident pitch angle θ of the sound source is s and the incident azimuth of the sound source The steering vector of the microphone array is determined as follows: Based on the incident pitch angle θ of the sound source s and the incident azimuth The position of the mth microphone is represented using Cartesian coordinates: Among them, θ m represents the elevation angle of the mth microphone; represents the azimuth angle of the mth microphone; r m It represents the distance between the mth microphone and the origin of the coordinate system; (·) T Represents the transpose operation of a vector or matrix; p m is the position of the mth microphone; Based on the incident pitch angle θ of the sound source s and the incident azimuth Get the wave number of the sound source: Where c = represents the speed of sound; ω represents the angular frequency; f represents the frequency; θ s Indicates the pitch angle of the incident direction of the sound source; Indicates the azimuth of the incident direction of the sound source; k s is the wave number of the sound source; Based on the position of the mth microphone and the wave number of the sound source, obtain the steering vector of the microphone array: Among them, k s is the wave number of the sound source; p m is the position of the mth microphone, and M represents the total number of microphones in the microphone array; (·) T Represents the transpose operation of a vector or matrix; = is the steering vector of the microphone array.
3. The speech noise reduction method based on differential beamformer according to claim 2, characterized in that: The step 3 is specifically as follows: Step 3.1: Determine the order N of the beamformer based on the actual noise reduction requirements; Step 3.2: Based on the order N in step 3.1, combined with the theoretical beam pattern of the differential beamformer, we can obtain The first ideal beam pattern at : Among them: a N,n (n=0,1,...,N) represents the real coefficient of the beam pattern; N represents the order of the beamformer; n represents the nth coefficient; Step 3.3: Based on the order N in step 3.1, define the N-order Chebyshev polynomial: Where: cosh -1 x and coshx are inverse functions of each other; N represents the order of the beamformer; Step 3.4: Based on the Nth order Chebyshev polynomial in step 3.3, let the main lobe amplitude correspond to The value of x0>1, the ratio of the main lobe amplitude to the side lobe amplitude is defined as: Where: L m Indicates the main lobe amplitude; L s represents the side lobe amplitude; R represents the ratio of the main lobe amplitude to the side lobe amplitude; Step 3.5: Based on the Nth order Chebyshev polynomial in step 3.3 and the ratio of the main lobe amplitude to the side lobe amplitude in step 3.4, we obtain: Where: x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude; Step 3.6: Based on the entire interval [-1, x0] of the N-order Chebyshev polynomial in step 3.3, let Where: c1 and c2 are constants; Indicates azimuth; Step 3.7: Based on the mapping relationship between π and x, by solving c1+c2=x0 and -c1+c2=-1, we get: Where: c1 and c2 are constants; x0 represents the position of the main lobe of the beam; The mapping relationship between π and x is as follows: Where: π represents 180°; represents the azimuth; c1 and c2 are both constants; Step 3.8: Substitute equation (11) into equation (10) to obtain: Where: x0 represents the position of the main lobe of the beam; - indicates the azimuth; Step 3.9: Based on equations (8), (9), (12) and the first ideal beam pattern in step 3.2, we get The second ideal beam pattern at : Where: R represents the ratio of the main lobe amplitude to the side lobe amplitude; x0 represents the position of the main lobe of the beam; Indicates the azimuth; at this time, within the range of The root is: exist within the range of The root is Step 3.10: Adjust the angle of the second ideal beam pattern in step 3.9 to the angle in step 2.
1. The third ideal beam pattern is obtained: in: Indicates the incident azimuth of the sound source; - represents the azimuth; x0 represents the position of the main lobe of the beam; R represents the ratio of the main lobe amplitude to the side lobe amplitude.
4. The speech noise reduction method based on differential beamformer according to claim 3, characterized in that: The step 4 is specifically as follows: Step 4.1: Let M = 2N + 1. Based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraint h H (ω)d(ω,90°,0)=1, using zero point constraint, i.e. constraint The beam pattern in the direction is 0, Beam pattern The root of θ s = 90° of the first linear system: In formula (23), Defined as: Where: θ0 = [θ 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; i M Represents the identity matrix I of size M×M M The first column of Step 4.2: Based on the first linear system in step 4.1 and the second ideal beam pattern in step 3.9, we get θ s = 90° when the first beamformer: in: Representation Matrix The inverse matrix of Step 4.3: Let M = 2N + 1, based on the steering vector of the microphone array in step 2.2.3, combine the distortion-free constraints Set the direction of zero point to Using the zero point constraint, we get The second linear system: In formula (17), Defined as: Where: i M Represents the identity matrix I of size M×M M The first column of 10 ,θ 20 ,…,θ N0 ,…,θ 2N0 ]=90°; M represents the total number of microphones in the microphone array; Step 4.4: Based on the second linear system in step 4.3 and the third ideal beam diagram in step 3.10, we get The second beamformer: in: Representation Matrix The inverse matrix of Step 4.5: Assume that M>2N+1. Based on the second beamformer in step 4.4, the third beamformer is obtained by using the minimum norm method and diagonal loading method: in: Representation Matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of ; η represents a small scalar (positive number); the value of η can be obtained by solving the following function: Among them: ∈ represents a small scalar (positive number) set freely; h DL (ω) represents the third beamformer; represents the third beamformer h DL The conjugate transpose of (ω); Step 4.6: Based on the second beamformer in step 4.4, the delay-sum beamformer is combined to obtain beamformer 4: Where: μ is a freely settable constant, 0≤μ≤1; Representation Matrix The inverse matrix of M Represents the identity matrix I of size M×M M The first column of .
5. A speech noise reduction system based on a differential beamformer, characterized in that: include: Speech signal acquisition and spectrum calculation module: collects noisy speech signals and converts them into short-time Fourier transform domain to obtain the spectrum of noisy speech signals; Steering vector determination module: Based on the spectrum of the noisy speech signal in step 1, the generalized cross-correlation method is used to estimate the incident pitch angle θ of the sound source s and the incident azimuth of the sound source Based on the incident pitch angle θ of the sound source s and the incident azimuth of the sound source determining a steering vector of the microphone array; Ideal beam pattern determination module: Determine the first ideal beam pattern based on the actual required noise reduction degree and the theoretical beam pattern of the differential beamformer; determine the second ideal beam pattern based on the first ideal beam pattern through Chebyshev polynomials; adjust the angle of the second ideal beam pattern to the incident azimuth angle of the sound source A third ideal beam pattern is obtained; Beamformer design module: Based on the steering vector of the microphone array, the second ideal beam diagram and the third ideal beam diagram, the first beamformer and the second beamformer are designed respectively by using the zero-point constraint method; based on the second beamformer, the third beamformer is designed by using the minimum norm method and diagonal loading method; based on the second beamformer, the fourth beamformer is designed by combining the delay-sum beamformer; A frequency domain clean speech signal estimation module: based on any one of the first beamformer, the second beamformer, the third beamformer, and the fourth beamformer, combined with the observation signal of the microphone array, determines a frequency domain estimation value of the clean speech signal; Time domain clean speech signal calculation module: Based on the frequency domain estimation value of the clean speech signal, the frequency domain estimation value of the clean speech signal is transformed into the time domain by using overlap-add or overlap-preserve method to obtain the speech signal after noise reduction processing.
6. A speech noise reduction device based on a differential beamformer, characterized in that: include: Memory: used to store a computer program to implement a speech noise reduction method based on a differential beamformer as described in any one of claims 1 to 4; Processor: used to implement the speech noise reduction method based on differential beamformer as described in any one of claims 1 to 4 when executing the computer program.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the speech noise reduction methods based on a differential beamformer are implemented.
Citation Information
Patent Citations
Self-adaptive sum and difference beam forming method
CN109799486A
Voice signal beam forming method based on interference noise spatial spectrum matrix
CN116312602A
Optimal modal beamformer for sensor arrays
US20120093344A1
Linear differential microphone arrays with steerable beamformers
US20220295175A1
Voice signal processing method and apparatus
WO2019169616A1
Cited By
N-order omnidirectional adjustable differential beam former with controllable zero point and design method
CN121072152A