A pickup method of a microphone array, an electronic device, and a storage medium
By using fixed beamforming and adaptive filter updates, the problem of error sensitivity in traditional microphone array pickup methods is solved, achieving higher quality speech pickup and enhancing the pickup performance of the microphone array.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- COLSONIC SUZHOU ELECTRONICS CO LTD
- Filing Date
- 2021-05-21
- Publication Date
- 2026-06-26
AI Technical Summary
Traditional microphone array pickup methods are sensitive to errors. Factors such as directional mismatch and microphone channel inconsistency can cause the desired signal to cancel out, thus reducing voice quality.
The method employs fixed beamforming, jamming processing, adaptive filter updates, and signal-to-noise ratio (SNR) estimation. Signals from unexpected directions of arrival are filtered out by first and second filters, while signals from expected directions of arrival are retained. SNR estimation is used to determine the timing of filter coefficient updates. By combining probability factors and filter constraints, speech quality is improved.
It effectively reduces the sensitivity to direction of arrival estimation error, improves speech quality, enhances pickup distance, and improves noise suppression capabilities.
Smart Images

Figure CN117037830B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention filed on May 21, 2021, with application number 202110556564.3. Technical Field
[0002] This invention belongs to the field of microphone array sound pickup, and relates to a robust microphone array sound pickup method, electronic device, and computer-readable storage medium. Background Technology
[0003] Video conferencing systems are essential tools for collaborative work, and online collaborative work is becoming increasingly popular. Voice pickup, as a crucial entry point for video conferencing systems, has therefore received widespread attention. Currently, the mainstream voice pickup method in video conferencing systems is single-microphone pickup. While single-microphone pickup is simple to implement, it is limited by factors such as sensitivity and complex sound reflection environments, resulting in a relatively short pickup distance. Microphone array pickup, on the other hand, utilizes more spatial information and has advantages such as high gain, strong noise suppression, and reverberation resistance, which can further extend the pickup distance.
[0004] Generalized Sidelobe Cancelling (GSC) algorithms have been widely used in microphone array sound pickup engineering because they can transform constrained optimality problems into unconstrained problems. Traditional GSC algorithms are quite sensitive to errors; factors such as directional mismatch, microphone channel inconsistency, and reverberation can all cause the desired signal to cancel out, thereby reducing speech quality. Although subsequent developments have made a series of improvements, there are still shortcomings. Summary of the Invention
[0005] The purpose of this invention is to provide a microphone array pickup method, electronic device, and computer-readable storage medium to further improve voice quality.
[0006] According to a first aspect of the present invention, a method for picking up sound using a microphone array includes the following steps:
[0007] S1. Perform fixed beamforming on the voice signal received by the microphone array, and point the beamforming direction of the microphone array toward the estimated expected direction of arrival.
[0008] S2. The voice signal processed in step S1 is blocked to block signals from the expected direction of arrival, and only signals from the unexpected direction of arrival are retained.
[0009] S3. Using the signal processed in step S2 as a reference signal, the signal with the unexpected direction of arrival in the speech signal processed in step S1 is filtered out by the first filter, while the signal with the expected direction of arrival is retained.
[0010] The sound pickup method also includes the following steps:
[0011] S4. Calculate the update factor of the first filter in step S3 according to the following formula (I), and update the coefficients of the first filter.
[0012]
[0013] in, SNR is the update factor of the first filter for the m-th microphone channel. f,d (ω,l) represents Y f,d The signal-to-noise ratio of (ω,l), Y f,d (ω,l) represents the signal Y after processing in step S1. f (ω,l) represents the delayed signal after delay processing, and its SNR. m (ω,l) represents the signal U after processing in step S2. m The signal-to-noise ratio (ω, l) is given by m = 1…M, where M is the number of microphone channels, ω is the angular frequency, and l is the frame index.
[0014] According to a preferred aspect, step S4 specifically includes:
[0015] S4-1, Estimate Y f,d The noise in (ω,l) will affect Y f,d Dividing the energy of (ω,l) by the noise yields the signal-to-noise ratio (SNR). f,d (ω,l);
[0016] S4-2, Estimate U m The noise in (ω,l) will make U m Dividing the energy of (ω,l) by the noise yields the signal-to-noise ratio (SNR). m (ω,l);
[0017] S4-3. Calculate the update factor according to formula (I), according to formula... The coefficients of the first filter are adaptively updated, where, These are the coefficients of the first filter in the current frame. These are the coefficients of the first filter in the next frame, μ is the step size factor, and Y(ω,l) is the signal output after processing in step S3. * For the conjugate of Y(ω,l), For U m Smooth energy of (ω,l).
[0018] According to a preferred aspect, in step S2, a second filter is used to block the speech signal processed in step S1, and the sound pickup method further includes the following steps:
[0019] S5. Calculate the update factor of the second filter in step S2 according to the following formula (II), and update the coefficients of the second filter.
[0020]
[0021] in, Let be the update factor of the second filter for the m-th microphone channel. For Y f The smooth energy of (ω,l), For U m Smooth energy of (ω,l), THR BM This is the preset threshold parameter.
[0022] More preferably, step S5 specifically includes:
[0023] S5-1, Estimating Y f Smooth energy of (ω,l)
[0024] S5-2, Estimating U m Smooth energy of (ω,l)
[0025] S5-3. Calculate the update factor according to formula (II), according to formula...
[0026] The second filter is adaptively updated, where These are the frequency domain coefficients of the second filter in the current frame. It is the intermediate frequency domain coefficient of the second filter in the next frame, U m (ω,l) * For U m The conjugate of (ω,l), It is Y f (ω,l) Signal Y after probability compensation c The smoothing energy of (ω,l), where μ is the step size factor;
[0027] frequency domain coefficients Convert to time domain coefficients Where n l+1 It is a discrete-time subscript, and according to the following formula... Make constraints,
[0028]
[0029] After the constraints are completed Then perform an FFT transform to convert it into the frequency domain coefficients of the second filter in the next frame. Proceed to the next round of filtering and coefficient updates, where low_bound... m (nl+1 ) and high_bound m (n l+1 These are the upper and lower limits of the preset filter coefficients, respectively;
[0030] The upper and lower limits of the filter coefficients are defined as follows:
[0031]
[0032] Where max{} represents the maximum number, t max In the allowed direction-of-arrival space [θ-θ err ,θ+θ err The maximum delay between the two channels, θ is the expected direction of arrival, θ err It is the maximum permissible directional error.
[0033] According to a preferred aspect, step S2 specifically includes:
[0034] S2-1. Delay the speech signal processed in step S1 to form signal Z. m,d (ω,l);
[0035] S2-2, Based on the signal Z from each microphone channel m,d The phase difference estimation signal of (ω,l) exists within a certain range [θ-θ] in the direction of arrival. err ,θ+θ err The probability of ], where θ is the expected direction of arrival of the wave. err This is the maximum permissible directional error;
[0036] S2-3, According to formula Y c (ω,l)=Prob(ω,l)Y f Y is obtained by performing probability compensation on (ω,l). c (ω,l);
[0037] S2-4, According to formula Filtered output, where, These are the frequency domain coefficients of the second filter for the m-th microphone channel;
[0038] S2-5. Adaptively update the coefficients of the second filter.
[0039] More preferably, step S2-2 is as follows:
[0040] S2-2-1, According to formula The phase difference is obtained by subtracting the phases of adjacent microphone channels. Where angle{} represents the phase of the signal, and unwrap{} represents the phase difference obtained by continuously adding or subtracting 2π. Z lies within the interval [-π, π]. m+1,d (ω,l) and Z m,d (ω,l) represent the delayed signals of two adjacent microphone channels, respectively;
[0041] S2-2-2, According to formula Phase difference Converted into time difference
[0042] S2-2-3, Based on the maximum allowable range of error angle θ err Converted to the maximum allowable time difference If the actual time difference is obtained exist If the signal is within this range, it is assumed that the desired signal is likely within the allowed direction-of-arrival space; if it is not within this range, it is assumed that the desired signal is likely not within the allowed direction-of-arrival space. A probability function is pre-defined. exist The value should be 1 if possible within this interval, and 0 if possible outside this interval, where s and α are preset parameters; based on the preset probability function Pr(t) and time difference. Calculate the probability Prob m (ω,l), then let the total probability factor of the signal existing in the allowed direction of arrival space be (ω,l).
[0043] S2-2-4. Correct the total probability factor Prob(ω,l) as shown in the following equation.
[0044]
[0045] Wherein, ω0 is the preset boundary frequency.
[0046] Furthermore, s satisfies when At that time, Pr(t) = 0.707.
[0047] According to a preferred aspect, in step S1, the microphone receives the signal X based on the estimated direction of arrival. m Z is obtained by performing a delay operation on (ω,l). m (ω,l), where X m (ω,l),m=1…M represents the STFT transformation of the signal received by the microphone array, transforming signal Z… m (ω,l) is fed into step S2; the delayed and aligned signals are added together to obtain the final signal. Signal Y f After delay processing of (ω,l), the signal Y is obtained. f,d (ω,l) is then fed into step S3.
[0048] According to a preferred aspect, in step S3, according to the formula Perform filtering output, where These are the coefficients of the first filter.
[0049] Preferably, the first filter is an NAF filter.
[0050] Preferably, the second filter is a CCAF filter.
[0051] According to a second aspect of the present invention, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the sound pickup method as described above.
[0052] According to a preferred and specific aspect, the electronic device is a remote conferencing device.
[0053] According to a third aspect of the present invention, a computer-readable storage medium stores a computer program that, when executed by a processor, implements the sound pickup method as described above.
[0054] The present invention adopts the above solution, which has the following advantages compared with the prior art:
[0055] The sound pickup method of the present invention can effectively filter out signals with a preset direction of arrival (DOA) across the entire frequency band, while retaining signals with non-preset DOA. This can effectively reduce the sensitivity to DOA estimation errors. Furthermore, by using signal-to-noise ratio (SNR) estimation to determine when to update the MC filter coefficients, the speech quality can be further improved. Attached Figure Description
[0056] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a schematic diagram of a sound pickup method according to an embodiment of the present invention;
[0058] Figure 2 This is a schematic diagram illustrating the updating principle of the first filter according to an embodiment of the present invention;
[0059] Figure 3 This is a schematic diagram of a microphone array;
[0060] Figure 4 The simulation results are shown for estimating the direction of human voice at 0 degrees.
[0061] Figure 5The simulation results are shown when the direction of the human voice is estimated to be 10 degrees. Detailed Implementation
[0062] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more readily understood by those skilled in the art. It should be noted that the description of these embodiments is for the purpose of aiding understanding the present invention, but does not constitute a limitation thereof.
[0063] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the word “comprising” as used in the specification of this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0064] Reference Figure 1 As shown, the voice signal picked up by the microphone array is processed by four modules: FBF, ABM, MC, and Control. The operation process of each part is described in detail below.
[0065] FBF (Fixed Beamforming) module: Fixed beamforming points the fixed microphone beamforming direction toward the estimated direction of arrival, enhancing the speech from the direction of arrival.
[0066] 1. For example Figure 1 As shown, X m (ω,l),m=1…M represents the STFT transformation of the microphone received signal, where M is the number of microphone channels, ω is the frequency, and l is the frame index.
[0067] 2. Figure 1 Steering is based on the estimated direction of arrival (DOA) of the microphone receiving signal X. m Z is obtained by performing a delay operation on (ω,l). m (ω,l) aligns the signals from the direction of arrival in time.
[0068] 3. Add the delayed and aligned signals together.
[0069] The ABM (Adaptive Blocking Matrix) module is used to block signals from the direction of arrival (θ), retaining only signals from non-direction-of-arrival (DOA) directions. A common fixing method is to delay and align the signal Z. mWe perform pairwise subtraction of (ω,l) because, theoretically, after alignment, the signals from the direction of arrival θ are consistent, and subtraction can yield the signals from the non-direction of arrival. However, in reality, the estimated direction of arrival... The error between the actual direction of arrival (θ) and the actual direction of arrival (θ) causes the signal output by the BM module to contain the signal in the direction of arrival θ, i.e., the leakage of the desired signal. This leads to the self-cancellation of the desired signal in the subsequent MC (Multiple-input Canceller) module. To solve this problem, this embodiment uses an adaptive filter that combines the spatial signal existence probability factor and CCAF (Coefficient-constrained adaptive filters) constraints to reduce the leakage of the desired signal.
[0070] 1. The delay module adds a delay to ensure the causality of the adaptive filter. The signal after the delay is Z. m,d (ω,l).
[0071] 2. Prob{} estimates the presence of a signal within a certain range [θ-θ] in the direction of arrival based on the phase difference of each channel signal. err ,θ+θ err The probability of θ err It is the maximum permissible directional error.
[0072] 2.1 The phase difference is obtained by subtracting the phases of adjacent channels, where angle{} represents the signal phase. Since the phase has a period of 2π, unwrap{} obtains the phase difference by continuously adding or subtracting 2π. It lies within the interval [-π, π].
[0073] 2.2 Converting phase difference into time difference ω is the angular frequency.
[0074] 2.3 If the estimated direction of arrival of the wave... If there is no error between the actual direction of arrival θ and the time difference It is 0 otherwise, based on the maximum allowable error angle θ. err Converted to the maximum allowable time difference If the actual time difference is If the signal falls within this range, it is assumed that the desired signal is highly likely to exist within the allowed direction-of-arrival space; otherwise, it is assumed that the desired signal is highly likely not to exist within the allowed direction-of-arrival space. A probability function is pre-defined. exist Ideally, the value should be 1, and ideally, it should be 0 outside this range. Here, s and α are preset parameters, where α controls the steepness of the transition from within to outside the preset time range; a larger value results in a steeper transition. After determining α, s is adjusted to satisfy the condition... At that time, Pr(t) = 0.707. Based on the preset probability function Pr(t) and time difference... Calculate the probability Prob m (ω,l), then let the total probability factor of the signal existing in the allowed direction of arrival space be (ω,l).
[0075] 2.4 Considering that the phase difference at mid-to-high frequencies may not be accurate in actual environments due to scattering, the probability factor at mid-to-high frequencies is not considered and is set to 1. Therefore, the final corrected total probability factor is: where ω0 is the preset boundary frequency:
[0076]
[0077] 3.Y c (ω,l)=Prob(ω,l)Y f (ω,l).
[0078] 4. Filtered output: U m (ω,l) and These are the output of the m-th channel and the second filter, respectively.
[0079] 5. Update the coefficients of the second filter; the second filter is... Figure 1 The CCAF filter in the model uses the commonly used NLMS algorithm for adaptive filter updates in the frequency domain.
[0080] U m (ω,l) * For U m The conjugate of (ω,l), where μ is the step size factor. For Y c The smooth energy of (ω,l), The update factor can only be 1 or 0, and is generated by the Control module.
[0081] After updating the frequency domain filter coefficients, it is necessary to convert the frequency domain coefficients to... Convert to time domain coefficients Where n l+1 It is a discrete-time subscript, and for Make constraints
[0082]
[0083] After the constraints are completed Then perform FFT transformation to convert Proceed to the next round of filtering and coefficient updates, where low_bound... m (n l+1 ) and high_bound m (n l+1 The upper and lower limits of the filter coefficients are preset separately. By presetting the upper and lower limits of the filter coefficients, the ABM output signal can retain only the signal other than the direction of arrival. The upper and lower limits of the filter coefficients are generally limited as follows:
[0084]
[0085] Where max{} represents the maximum number, t max In the allowed direction-of-arrival space [θ-θ err ,θ+θ err ], the maximum delay between the two channels.
[0086] The core of the CCAF algorithm is to constrain the filter to only filter signals in a preset direction of arrival (DOA) by imposing upper and lower limits on the filter coefficients, while retaining signals outside the preset DOA. However, the constraint chosen according to the above formula still results in signals outside the preset DOA at low frequencies, which is detrimental to the subsequent MC module's elimination of these signals. Using phase difference to determine whether a signal is within the preset DOA is more accurate at low frequencies. Therefore, by using phase difference to determine if a signal is within the preset DOA, if it exists, the probability is close to 1, and the CCAF reference input signal remains essentially unchanged, thus facilitating CCAF's removal of signals within the preset DOA. If it does not exist, the probability is close to 0, and the CCAF reference input signal is essentially zero. Therefore, no matter how it is updated, it cannot remove signals outside the preset DOA, which is beneficial for the subsequent MC module to further eliminate noise.
[0087] The ABM module utilizes an adaptive filter constrained by the joint spatial signal existence probability factor and CCAF (Coefficient-constrained adaptive filters) to effectively filter out signals with a preset direction of arrival across the entire frequency band, while retaining signals with outputs from directions other than the preset direction of arrival.
[0088] MC (Multiple-input Canceller) module: Utilizes ABM's module output U m (ω,l) is used as a reference signal to filter out signals in the FBF output signal that are not in the preset direction of arrival, and to maximize the retention of signals in the preset direction of arrival.
[0089] 1. Filtered output: in It is the first filter, i.e. Figure 1 The coefficients of the adaptive filter NAF in the model.
[0090] 2. First filter coefficient update: NAF uses the commonly used NLMS algorithm to perform adaptive filter update in the frequency domain, while limiting the energy of the filter coefficients. If the total energy exceeds the preset value, it is normalized according to the preset value; otherwise, it remains unchanged.
[0091]
[0092] Where Y(ω,l) * Let μ be the conjugate of Y(ω,l), and μ be the step size factor. For U m The smooth energy of (ω,l), The update factor can only be 1 or 0, and is generated by the Control module.
[0093] Control Module: Despite various constraints, ABM will still contain a small amount of signal with a preset direction of arrival (DOA). If this signal is speech, updating the filter in the MC module at this time will impair the output speech. To reduce speech impairment, it is necessary to determine when to update the filter coefficients. In the Control module, C refers to the comparator, SNR refers to calculating the signal-to-noise ratio, and E refers to calculating the smoothing energy.
[0094] produce:
[0095] 1. Estimate Y f Smooth energy of (ω,l)
[0096] 2. Estimate U of the m-th channel m Smooth energy of (ω,l)
[0097] 3. Among them THR BM It is a preset threshold parameter.
[0098] produce:
[0099] 1. Estimate Y f,d The signal-to-noise ratio (SNR) of (ω,l) f,d (ω,l):
[0100] 1.1 Estimating Y using noise methods f,d For noise in (ω,l), the commonly used monophonic noise estimation method is the mcra method. Refer to the book "Loizou, Philipos C, Speech Enhancement: Theory and Practice";
[0101] 1.2, Y f,d Divide the energy of (ω,l) by the noise in 1.1 to obtain the current signal-to-noise ratio (SNR). f,d (ω,l);
[0102] 2. Similarly, estimate U m The signal-to-noise ratio (SNR) of (ω,l) is SNR m (ω,l);
[0103] 3.
[0104] The update principle is described as follows:
[0105] See Figure 2 ,
[0106] make
[0107] v1(ω)=a1s(ω)+b1n(ω) (1)
[0108] v2(ω)=a2s(ω)+b2n(ω) (2)
[0109] g(ω)=v1(ω)-hv2(ω) (3)
[0110] Where s(ω) is the speech signal, n(ω) is the noise signal, ω is the angular frequency, a1, a2, b1, and b2 are the corresponding weighting coefficients, v1(ω) is the desired signal, and v2(ω) is the reference input signal, then the optimal problem expression is (the following is simplified, omitting the symbol ω):
[0111]
[0112] Where E{} represents the expected value. The optimal solution obtained by optimizing equation (4) is:
[0113]
[0114]
[0115] Substituting formulas (1), (2), and (6) into (3) yields...
[0116]
[0117] Define input signal-to-noise ratio
[0118]
[0119] Define the output signal-to-noise ratio
[0120]
[0121] The desired signal-to-noise ratio (SNR) of the output signal g after passing through the adaptive filter is... o It must be greater than the signal-to-noise ratio (SNR) of the original signal v1.
[0122]
[0123] in Substituting it into formula (10) and resolving it yields...
[0124]
[0125] in
[0126]
[0127] Substituting formula (12) into formula (11) yields
[0128]
[0129] Therefore, if you want to improve the signal-to-noise ratio, SNR1 and SNR2 must be less than 1.
[0130] The ABM module in this algorithm utilizes an adaptive filter constrained by the joint spatial signal existence probability factor and CCAF (Coefficient-constrained adaptive filters). It can effectively filter out signals with preset directions of arrival (DOA) across the entire frequency band while retaining signals with non-preset DOA. This can effectively reduce the sensitivity to DOA estimation errors. Furthermore, it uses signal-to-noise ratio (SNR) estimation to determine when to update the MC filter coefficients, further improving speech quality.
[0131] Simulation Example
[0132] Reference Figure 3 As shown, the microphone array used is a three-element uniformly distributed circular array. Angles are calculated by counter-clockwise rotation. The angles of the three-element array are [90, 210, 330] degrees, and the circumference radius is 4 cm. The target human voice is located at 0 degrees, and the noise source is located at 110 degrees. The signal-to-noise ratio is 0 dB. The boundary frequency point in the ABM algorithm is 300 Hz, α in the probability function is set to 4, the maximum allowable error direction is ±10 degrees, the filter order is 160, the step size factor is 0.1, and the delay p is 80. In MC, the filter order is 160, the step size factor is 0.1, the delay q is 100, and the square root of the total constraint energy of the filter is set to 0.2. The THR in the control module... BM We set it to 0.5. The signal-to-noise ratio used in the simulation is approximately 6dB.
[0133] Simulations were performed using the traditional GSC method and the robust-GSC method of this embodiment, and the results are compared below.
[0134] The estimated direction of the human voice is 0 degrees, meaning there is no error. See Table 1 for the results. Figure 4 .
[0135] Table 1
[0136] gsc robust-gsc Noise reduction (dB) 23.1421 18.0317 PESQ 1.7943 2.3794
[0137] The estimated direction of the human voice is 10 degrees, meaning there is an error of 10 degrees. See Table 2 for the results. Figure 5 .
[0138] Table 2
[0139] gsc robust-gsc Noise reduction (dB) 22.5463 17.8811 PESQ 1.3403 2.3817
[0140] Simulation results show that, under error-free conditions, although robust-GSC's noise reduction is slightly worse than traditional GSC, its PESQ value is significantly improved, resulting in a marked improvement in speech quality. However, under error-containing conditions, traditional GSC's speech quality deteriorates further, with almost the entire speech signal being eliminated. Therefore, the proposed robust-GSC does not significantly reduce either the noise reduction amount or the PESQ value.
[0141] The above embodiments are merely illustrative of the technical concept and features of the present invention, and are preferred embodiments. Their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and they should not be construed as limiting the scope of protection of the present invention. All equivalent transformations or modifications made according to the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for picking up sound using a microphone array, comprising the following steps: S1. Perform fixed beamforming on the voice signal received by the microphone array, and point the beamforming direction of the microphone array toward the estimated expected direction of arrival. S2. The second filter is used to block the speech signal processed in step S1, so as to block the signal from the expected direction of arrival and retain only the signal from the non-expected direction of arrival. S3. Using the signal processed in step S2 as a reference signal, the signal with the unexpected direction of arrival in the speech signal processed in step S1 is filtered out by the first filter, while the signal with the expected direction of arrival is retained. Its features are, The sound pickup method also includes the following steps: S4. Calculate the update factor of the first filter in step S3 according to the following formula (I), and update the coefficients of the first filter according to the calculated update factor. (I) in, For the first m The update factor of the first filter for each microphone channel. for The signal-to-noise ratio, The signal processed in step S1 The delayed signal after delay processing, The signal after processing in step S2 The signal-to-noise ratio, , M It is the number of microphone channels. It is angular frequency. l It is a frame index; S5. Calculate the update factor of the second filter in step S2 according to the following formula (II), where the second filter is a CCAF filter. The CCAF filter adaptively updates the coefficients of the second filter in the frequency domain using the NLMS algorithm based on the update factor of the second filter. (II) in, For the first m The update factor of the second filter for each microphone channel. for Smooth energy, for Smooth energy, This is the preset threshold parameter.
2. The sound pickup method according to claim 1, characterized in that, In step S4, the first filter is an adaptive filter, and the adaptive filter is adaptively updated in the frequency domain using the NLMS algorithm.
3. The sound pickup method according to claim 1, characterized in that, Step S5 specifically includes: S5-1, Estimation smooth energy ; S5-2, Estimation smooth energy ; S5-3. Calculate the update factor according to formula (II), according to formula... The second filter is adaptively updated, where These are the frequency domain coefficients of the second filter in the current frame. These are the intermediate frequency domain coefficients of the second filter in the next frame. for conjugate, yes Signal after probability compensation Smooth energy, Step size factor; frequency domain coefficients Convert to time domain coefficients ,in It is a discrete-time subscript, and according to the following formula... Make constraints, ; After the constraints are completed Then perform an FFT transform to convert them into the frequency domain coefficients of the second filter in the next frame. Then proceed to the next round of filtering and coefficient updates, where and These are the upper and lower limits of the preset filter coefficients, respectively; The upper and lower limits of the filter coefficients are defined as follows: ; Where max{} represents the maximum number. In the allowed direction of arrival space Maximum delay between the two channels It is the expected direction of the wave. It is the maximum permissible directional error.
4. A method for picking up sound using a microphone array, comprising the following steps: S1. Perform fixed beamforming on the voice signal received by the microphone array, and point the beamforming direction of the microphone array toward the estimated expected direction of arrival. S2. The voice signal processed in step S1 is blocked to block signals from the expected direction of arrival, and only signals from the unexpected direction of arrival are retained. S3. Using the signal processed in step S2 as a reference signal, the signal with the unexpected direction of arrival in the speech signal processed in step S1 is filtered out by the first filter, while the signal with the expected direction of arrival is retained. Its features are, Step S2 specifically includes: S2-1. Delay the speech signal processed in step S1 to form a signal. ; S2-2, Based on the signals from each microphone channel The phase difference estimation signal exists within a certain range of the direction of arrival. The probability, It is the expected direction of the wave. This is the maximum permissible directional error; S2-3, According to formula get ; S2-4, According to formula Filtered output, where, It is the first m Frequency domain coefficients of the second filter for each microphone channel; S2-5. Adaptively update the coefficients of the second filter; Specifically, step S2-2 is as follows: S2-2-1, According to formula The phase difference is obtained by subtracting the phases of adjacent microphone channels. Where angle{} is used to take the signal phase, and unwrap{} is used to take the phase of the signal through continuous addition or subtraction. Let the phase difference Located in the range Inside, and These are the delayed signals from two adjacent microphone channels, respectively. S2-2-2, According to formula Phase difference Converted into time difference ; S2-2-3, Based on the maximum allowable range of error angles Converted to the maximum allowable time difference If the actual time difference is obtained exist If the signal is within this range, it is assumed that the desired signal is likely within the allowed direction-of-arrival space; if it is not within this range, it is assumed that the desired signal is likely not within the allowed direction-of-arrival space. A probability function is pre-defined. ,exist The value inside this range should be as close to 1 as possible, and the value outside this range should be as close to 0 as possible. and These are preset parameters; based on a preset probability function. and time difference Calculate the probability Then let the total probability factor of the signal existing in the allowed direction-of-arrival space be . ; S2-2-4, Regarding the total probability factor Make corrections to achieve the following equation. ; in, This is the preset boundary frequency.
5. The sound pickup method according to claim 4, characterized in that, Satisfy when hour, .
6. The sound pickup method according to claim 4, characterized in that, In step S1, the microphone receives the signal based on the estimated direction of arrival. Obtain by performing a delayed operation ,in, STFT transformation of the signal received by the microphone array, converting the signal... The signal is fed into step S2; the delayed and aligned signals are then added together to obtain the final signal. , will signal The signal is obtained after delay processing. And send it to step S3; In step S3, according to the formula Perform filtering output, where These are the coefficients of the first filter.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the sound pickup method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the sound pickup method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Pickup method of microphone array, electronic equipment and storage medium
CN113470681A