First-order differential microphone array with steerable beamformer

By decomposing the target beam pattern into cardioid and dipole beam patterns, a differential microphone array design solves the problem of beamformers being difficult to flexibly turn in existing technologies, achieving better signal acquisition and noise suppression effects.

CN116325795BActive Publication Date: 2025-10-28NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180068171.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-10
Publication Date
2025-10-28
Estimated Expiration
2041-02-10

AI Technical Summary

Technical Problem

Existing differential microphone arrays struggle to effectively steer to directions other than the linear end-fire direction when flexible beamforming is required, resulting in poor signal acquisition and noise reduction performance.

Method used

By decomposing the target beam pattern into two sub-beam patterns, a first sub-beamformer is designed for the cardioid pattern and a second sub-beamformer for the dipole pattern. By combining the Hadamard product and zero-constraint method, a steerable beamformer is generated, enabling flexible beam pattern control.

Benefits of technology

It realizes steerable beamforming of differential microphone arrays, improves signal acquisition and noise reduction capabilities, and enhances the quality of voice signals and reduces noise interference, especially in smart devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116325795B_ABST
    Figure CN116325795B_ABST
Patent Text Reader

Abstract

A first-order differential microphone array (FODMA) with a steerable beamformer is constructed by specifying a target beamform of the FODMA at a steering angle θ, and then decomposing the target beamform into a first sub-beamformation and a second sub-beamformation based on the steering angle θ. A first sub-beamformer and a second sub-beamformer are generated for each filter signal from the microphones of the FODMA, wherein the first sub-beamformer is associated with a first sub-beamformation, and the second sub-beamformer is associated with a second sub-beamformation. A steerable beamformer is then generated based on the first and second sub-beamformers. Decomposing the target beamformation into the first and second sub-beamformers involves dividing the target beamformation into the sum of a first-order cosine (cardioid) first sub-beamformation and a first-order sine (dipole) second sub-beamformation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to differential microphone arrays, and more particularly, to constructing a first-order differential microphone array (FODMA) with a steerable differential beamformer. Background Technology

[0002] A differential microphone array (DMA) uses signal processing techniques to obtain a directional response to a source sound signal based on the difference between pairs of source signals received by the array's microphones. A DMA can comprise an array of microphone sensors that respond to the spatial derivative of the sound pressure field generated by the sound source. The microphones of a DMA can be arranged on a common planar platform according to the geometry of the microphone array (e.g., linear, circular, or other array geometries).

[0003] A beamformer DMA can be communicatively coupled to a processing device (e.g., a digital signal processor (DSP) or a central processing unit (CPU)) that includes circuitry programmed to implement a beamformer to calculate estimates of sound sources. A beamformer is a spatial filter that uses multiple versions of the sound signal captured by microphones in a microphone array to identify sound sources according to certain optimization rules. The beam pattern reflects the beamformer's sensitivity to plane waves incident on the DMA from a specific angular direction. DMAs, combined with appropriate beamforming algorithms, have been widely used in voice communication and human-machine interface systems to extract the speech signal of interest from unwanted noise and interference. Attached Figure Description

[0004] This disclosure is illustrated by way of example rather than limitation in the accompanying drawings.

[0005] Figure 1 This is a flowchart illustrating a method for constructing a first-order differential microphone array (FODMA) with a steerable beamformer according to an embodiment of the present disclosure.

[0006] Figure 2 This is a flowchart illustrating a method for constructing a first-order differential microphone array (FODMA) with a steerable beamformer according to an embodiment of the present disclosure.

[0007] Figure 3 The array geometry of microphones in a FODMA arranged as a uniform linear differential microphone array (LDMA) according to an embodiment of the present disclosure is shown.

[0008] Figure 4A A graph showing the DF value of FODMA as a function of the coefficients of the target beam pattern according to an embodiment of the present disclosure.

[0009] Figure 4BA graph showing the DF value of FODMA as a function of the selected steering angle according to an embodiment of the present disclosure.

[0010] Figure 5A A diagram showing the beam pattern of FODMA at a selected steering angle according to an embodiment of the present disclosure.

[0011] Figure 5B A graph showing the DF value of FODMA as a function of frequency according to an embodiment of the present disclosure.

[0012] Figure 5C A diagram showing the beam pattern of FODMA as a function of frequency according to an embodiment of the present disclosure.

[0013] Figure 5D A graph showing the approximate error between the target beam pattern of FODMA as a function of frequency and the beam pattern of a steerable beamformer according to an embodiment of the present disclosure.

[0014] Figure 6A A spectrogram of clean speech from a steerable beamformer according to an embodiment of the present disclosure is shown, wherein the speech source is located at a selected steering angle.

[0015] Figure 6B A spectrogram of a noisy speech signal from a steerable beamformer according to an embodiment of the present disclosure is shown, wherein the speech source is located at a selected steering angle.

[0016] Figure 6C A spectrogram of an enhanced speech signal from a steerable beamformer according to an embodiment of the present disclosure is shown, wherein the speech source is located at a selected steering angle.

[0017] Figure 7A A diagram showing the target beam pattern of FODMA and the beam pattern of a steerable beamformer according to an embodiment of the present disclosure.

[0018] Figure 7B A diagram showing the target beam pattern of FODMA and the beam pattern of a steerable beamformer according to an embodiment of the present disclosure.

[0019] Figure 8 This is a block diagram illustrating an example form of a computer system within which a set or sequence of instructions can be executed to cause the machine to perform any of the methods discussed herein. Detailed Implementation

[0020] DMA (Differentiation of Magnetic Fields) measures the derivatives (of different orders) of the sound signals captured by each microphone, where the collection of sound signals forms a sound field associated with the microphone array. For example, a first-order DMA beamformer formed using the difference between a pair of microphones (adjacent or non-adjacent) can measure the first derivative of the sound pressure field. A second-order DMA beamformer can be formed using the difference between a pair of two first-order differences of a first-order DMA. A second-order DMA can measure the second derivative of the sound pressure field by using at least three microphones. Typically, an Nth-order DMA beamformer can measure the Nth-order derivative of the sound pressure field by using at least N+1 microphones.

[0021] The beamformer pattern of a DMA can be quantified in one aspect by the directivity factor (DF), which is the ability of the beamformer to maximize its sensitivity in the viewing direction as a ratio of its average sensitivity over the entire space. The viewing direction is the angle of impact from which the desired sound source originates. The DF of a DMA beamformer can increase with the order of the DMA. However, higher-order DMAs may be very sensitive to noise generated by the hardware components of each microphone in the DMA itself, where sensitivity is measured in terms of white noise gain (WNG). The design of a DMA beamformer can focus on finding the optimal beamforming filter for a given array geometry (e.g., linear, circular, square, etc.) based on certain criteria (e.g., beamformer, DF, WNG, etc.).

[0022] First-order differential microphone arrays (FODMAs), combining a finely pitched uniform linear array with a first-order differential beamformer, have been widely used for sound and speech signal acquisition. In applications such as hearing aids and Bluetooth headsets, the direction of the sound source can be assumed, and beamformer directional beamforming is not actually required. However, in many other applications, such as smart TVs, smartphones, and tablets, directional beamformers may be necessary because the sound source location may not be directly impacted along the end-fire direction. For example, an LDMA can be mounted along the bottom of a smart TV with voice recognition capabilities to form a beam pattern along the wide edge of the TV. Therefore, beamformers capable of directional beamforming of such LDMAs would be useful for maximizing signal acquisition (e.g., user's voice) and noise reduction.

[0023] This disclosure provides a method for designing a linear differential microphone array (LDMA) with steerable beamformers. The method described herein involves dividing a target beam pattern into a sum of two sub-beam patterns, such as cardioid and dipole, where the sum is controlled by a steering angle. Two sub-beamformers are constructed: the first, similar to a conventional beamformer, is used to realize the cardioid sub-beam pattern, while the second is designed to filter the squared observation signal to approximate the dipole sub-beam pattern. The design of the second sub-beamformer focuses on estimating the spectral amplitude of the signal of interest, while not emphasizing the spectral phase, which is generally accepted in speech enhancement and noise reduction.

[0024] method

[0025] For ease of explanation, the method is depicted and described as a series of actions. However, the actions according to this disclosure can occur in various sequences and / or simultaneously, and together with other actions not presented and described herein. Furthermore, not all presented actions require the implementation of the method according to the disclosed subject matter. Additionally, these methods can also be represented by a series of interrelated states through state diagrams or events. Furthermore, it should be understood that the methods disclosed in this disclosure can be stored on an article of manufacture to facilitate the transfer and assignment of such methods to a computing device. As used herein, the term "article of manufacture" is intended to include a computer program accessible from any computer-readable device or storage medium. In one embodiment, these methods can be performed by... Figure 3 The LDMA300 is associated with hardware processing for execution.

[0026] Figure 1 This is a flowchart illustrating a method 100 for constructing a first-order differential microphone array (FODMA) with a steerable beamformer according to an embodiment of the present disclosure. As described herein, a steerable beamformer refers to a beamformer that can be steered away from the end-fire direction of the FODMA.

[0027] refer to Figure 1 At 102, the processing device can begin to perform operations to construct a first-order differential microphone array (FODMA) with a steerable beamformer, such as determining a signal model.

[0028] In one implementation, a uniform linear array of M microphones can be used to capture the signal of interest, for example... Figure 3 The LDMA300. In the frequency domain, the signal received at the m-th microphone, m = 1, 2, ..., M, can be expressed as:

[0029]

[0030] Where X(ω) is the signal of interest (also called the desired signal) received at the first microphone, Xm (ω) and V m (ω) represents the speech and additive noise signals received at the m-th microphone, respectively, where j is the imaginary unit. 2 =-1, ω=2πf is the angular frequency, f>0 indicates the time frequency, τ0=δ / c, δ is the microphone spacing, c is the speed of sound in air, which is generally assumed to be 340m / s, and θ is the source incident angle. In DMA, it is assumed that the spacing δ is much smaller than the minimum acoustic wavelength of the band of interest, such that ωτ0≤2π. For example, in the simulations and experiments described below, the values ​​of δ=1cm and δ=1.1cm are used for the spacing of the FODMA microphones. Since cosθ is an even function, the beam pattern of the linear array is symmetrical with respect to the line connecting all the sensors. Therefore, in the description below, the range of θ can be restricted to [0,π].

[0031] Traditionally, beamforming is achieved by applying a linear spatial filter h(ω) to the microphone's observed signal, i.e.

[0032]

[0033] in,

[0034]

[0035] It is the observed signal vector, and v(ω) is a noise signal vector defined similarly to the observed signal vector y(ω).

[0036]

[0037] It is a phase vector, and the superscripts * and H represent the complex conjugation and transpose conjugation operators, respectively. T is the transpose operator, and Z(ω) is an estimate of X(ω). One objective of beamforming is to determine the optimal filter under specific criteria such that Z(ω) is a good estimate of X(ω).

[0038] At 104, the processing device can specify the target beam pattern of FODMA at the steering angle θ.

[0039] Using linear microphone arrays and conventional beamforming methods, as described above at (2), the beam pattern of FODMA may lack steering flexibility, meaning its main lobe may be difficult to steer to directions other than the linear end-fire direction. In one implementation, to steer the main lobe to any direction within the range θ∈[0,π], the target frequency-independent beam pattern of FODMA can be expressed as:

[0040] B1(θ)=a0+a1 cos θ+a2 sin θ (5)

[0041] Where a0, a1, and a2 are real coefficients that determine the shape of the target beam pattern of FODMA.

[0042] At position 106, the processing device can decompose the target beam pattern into a first sub-beam pattern and a second sub-beam pattern based on the steering angle θ.

[0043] The target beam pattern of FODMA can be decomposed into two sub-beam patterns B. 1,1 (θ)+B 1,2 (θ), where:

[0044] B 1,1 (θ)=a0+a1 cos θ, (6)

[0045]

[0046] These are the first-order cosine (cardioid) pattern and the first-order sine (dipole) pattern, respectively. If a2 = 0, the target beam pattern degenerates into a special case of equation (2) above. Based on the properties of Fourier series expansion, any first-order beam pattern continuous in [0, 2π] can be represented by the target beam pattern (5). In the main lobe (or desired steering) direction θ = θ d The target beam pattern should be distortion-free, i.e., B1(θ) d ) = 1. Therefore, the following two conditions are satisfied:

[0047]

[0048] Given the target beam pattern in equation (5) above, the problem of differential beamforming becomes finding the beamforming filter h(ω) in (2) so that the resulting beam pattern is similar to the target beam pattern.

[0049] At 108, the processing device can generate a first sub-beamformer and a second sub-beamformer for each filtered signal from the microphone of FODMA, wherein the first sub-beamformer is associated with a first sub-beam pattern and the second sub-beamformer is associated with a second sub-beam pattern.

[0050] The processing device can generate two sub-beamformers h1(ω) and h2(ω), the output of which can be expressed as:

[0051]

[0052]

[0053] Where {M1, M2}≤M, and the definitions of h1(ω) and h2(ω) are similar to those of h(ω),

[0054]

[0055]

[0056] The definition of v1(ω) is similar to that of y1(ω). The definition is similar to ⊙ represents the Hadamard product (element-by-element product).

[0057]

[0058]

[0059] These are two phase vectors, and the definition of d2(ω, cosθ) is similar to that of d1(ω, cosθ).

[0060] At 110, the processing device can generate a steerable beamformer based on the first sub-beamformer and the second sub-beamformer.

[0061] Given Z1(ω) and Z2(ω), the estimate X(ω) of the desired signal can be obtained as follows:

[0062]

[0063] Where φ1(ω) is the spectral phase of the output of the sub-beamformer h1(ω) (or the phase estimate from the original noise phase or the clean speech spectrum can be used). The spectral phase has almost no impact on the quality of the estimated signal. Based on equations (9) and (10) above, the beam patterns of the two sub-beamformers can be defined as:

[0064]

[0065]

[0066] Equation (17) for defining the beam pattern of the second sub-beamformer (e.g., h2(ω)) is based on the above equation (10), which applies to the squared signal from the observed signal vector (e.g., ) is filtered. In one implementation, the cross terms in (10) can be ignored, which should not affect the validity of the beam pattern, since the signal of interest and any noise signal are assumed to be uncorrelated.

[0067] Therefore, the overall beam pattern of the designed beamformer is as follows:

[0068] B d (θ)=B1[h1(ω),θ]+B2[h2(ω),θ], (18)

[0069] In view of the above formula, beamforming in embodiments of this disclosure includes constructing filters h1(ω) and h2(ω) (e.g., first and second sub-beamformers) in an optimal manner such that their combination (e.g., a steerable beamformer of FODMA) produces beam pattern B. d (θ), for example (18) above, is similar to the target beam pattern given in equation (5) above.

[0070] The two sub-beamformers h1(ω) and h2(ω) can be determined using the zero-constraint method widely used in the design of differential beamformers. Based on M1≥2, h1(ω) can be constructed using the following linear system:

[0071] D(ω)h1(ω)=β1, (19)

[0072] in

[0073]

[0074]

[0075] The minimum norm solution of equation (19) can be expressed as:

[0076] h 1,MN (ω)=D H (ω)[D(ω)D H (ω)] -1 β1 (22)

[0077] Then, based on M2≥3, h2(ω) can be constructed using the following linear system:

[0078] T(ω)h2(ω)=β2, (23)

[0079] in

[0080]

[0081]

[0082] The minimum norm solution of equation (23) can be expressed as:

[0083] h 2,MN (ω)=T H (ω)[T(ω)T H (ω)] -1 β2. (26)

[0084] In the special cases of M1=2 and M2=3, from (22) and (26):

[0085] h 1,DI (ω)=D -1(ω)β1, (27)

[0086] h 2,DI (ω)=T -1 (ω)β2, (28)

[0087] "DI" stands for "direct inverse".

[0088] At 112, the processing device can terminate the operation of constructing FODMA with a steerable beamformer.

[0089] Figure 2 This is a flowchart illustrating a method 200 for constructing a first-order differential microphone array (FODMA) with a steerable beamformer according to an embodiment of the present disclosure. As described above, a steerable beamformer refers to a beamformer that can be steered away from the end-fire direction of the FODMA.

[0090] refer to Figure 2 At 202, the processing device can begin to perform operations to construct a first-order differential microphone array (FODMA) with a steerable beamformer, such as determining a signal model.

[0091] As mentioned above, regarding Figure 1 The signal of interest can be captured using a uniform linear array of M microphones, for example... Figure 3 The LDMA 300. In the frequency domain, the received signal at the m-th microphone, m = 1, 2, ..., M, can be expressed according to the above equation (1).

[0092] Traditionally, beamforming is achieved by applying a linear spatial filter h(ω) to the microphone's observed signal, i.e., equations (2), (3), and (4) above. As mentioned above, the goal of beamforming is to determine the optimal filter h(ω) so that the filtered signal from the FODMA microphone matches the signal of interest from the sound source (e.g., human voice).

[0093] At position 204, multiple (M) microphones can be organized on a substantially planar platform, the multiple microphones comprising a first subset (M1) of microphones and a second subset (M2) of microphones.

[0094] As follows about Figure 3More fully, FODMA can include uniformly distributed microphones (1, 2, ..., m, ..., M) arranged according to a linear array geometry on a common global platform. As described above with respect to the outputs of the two sub-beamformers h1(ω) and h2(ω) (see equations (9) and (10) above), signals from a set of microphones are used for each beamformer, where h1(ω) uses microphones from 1 to M1, and h2(ω) uses microphones from 1 to M2, where {M1, M2} ≤ M, and {} is the union operator.

[0095] At 206, the processing device can construct a first sub-beamformer based on a first subset (M1) of the microphone and a target beam pattern at a steering angle θ, wherein the first sub-beamformer is characterized according to a first-order cosine (cardioid) first sub-beam pattern.

[0096] Using linear microphone arrays and conventional beamforming methods, as described above (2), the beam pattern of FODMA may lack steering flexibility, i.e., its main lobe may be difficult to steer to directions other than the linear end-fire direction. As mentioned above, in one embodiment, in order to steer the main lobe to any direction in the range θ∈[0,π], the target frequency-independent beam pattern of FODMA can be expressed according to (5), where a0, a1, and a2 are real coefficients that determine the shape of the target beam pattern of FODMA.

[0097] As described above, according to (6) and (7), the target beam pattern of FODMA can be decomposed into two sub-beam patterns B. 1,1 (θ)+B 1,2 (θ), which represent the first-order cosine (cardioid) pattern and the first-order sine (dipole) pattern, respectively.

[0098] The processing device can generate two sub-beamformers h1(ω) and h2(ω), and the output of the first sub-beamformer can be represented as shown in (9) above:

[0099]

[0100] Where M1 is a subset of M, and h1(ω) is defined similarly to h(ω).

[0101]

[0102] As stated in (11), v1(ω) is defined similarly to y1(ω), and

[0103]

[0104] As described in (13), it is the phase vector.

[0105] At 208, the processing device can construct a second sub-beamformer based on a second subset (M2) of the microphone and the target beam pattern at the steering angle θ, wherein the second sub-beamformer is characterized according to a first-order sinusoidal (dipole) second sub-beam pattern.

[0106] As described above, according to (6) and (7), the target beam pattern of FODMA can be decomposed into two sub-beam patterns B. 1,1 (θ)+B 1,2 (θ), which represent the first-order cosine (cardioid) pattern and the first-order sine (dipole) pattern, respectively.

[0107] The processing device can generate two sub-beamformers h1(ω) and h2(ω), and the output of the second sub-beamformer can be represented as shown in (10) above:

[0108]

[0109] Where M2 is a subset of M, and h2(ω) is defined similarly to h(ω).

[0110]

[0111] As stated in (12), The definition is similar to ⊙ represents the Hadamard product (element-by-element product).

[0112]

[0113] As described in (14), the phase vector is d2(ω, cosθ), and d2(ω, cosθ) can be defined similarly to d1(ω, cosθ).

[0114] At 210, the processing device can generate a steerable beamformer based on the first sub-beamformer and the second sub-beamformer.

[0115] Given Z1(ω) and Z2(ω), the estimate of the desired signal X(ω) can be obtained at (15) as described above. The beam patterns of the two sub-beamformers can be defined as shown in (16) and (17), therefore, the overall beam pattern of the designed beamformer is:

[0116] B d (θ)=B1[h1(ω),θ]+B2[h2(ω),θ],

[0117] As shown in (18). Given the above formula, beamforming in embodiments of this disclosure includes constructing filters h1(ω) and h2(ω) (e.g., first and second sub-beamformers) in an optimal manner such that their combination (e.g., a steerable beamformer) produces beam pattern B. d(θ), for example (18) above, is similar to the target beam pattern given in equation (5) above.

[0118] At 212, the processing device can terminate the operation of constructing FODMA with a steerable beamformer.

[0119] system

[0120] Figure 3 The array geometry of the microphones of the FODMA 300, arranged as a uniform linear differential microphone array (LDMA) according to an embodiment of the present disclosure, is shown.

[0121] FODMA 300 may include uniformly distributed microphones (1, 2, ..., m, ..., M) arranged according to a linear array geometry on a common platform. The positions of these microphones can be specified relative to a reference point (e.g., microphone 1). The coordinates of microphones (2, ..., m, ..., M) of FODMA 300 can be specified by a distance mδ, where m = 1, 2, ..., M-1, representing the distance between the m-th microphone of FODMA 300 and the specified reference point: the distance between microphone 1 of FODMA 300 and itself is 0. Therefore, the vector ρ = [0, δ, 2δ, ..., mδ, ..., (M-1)δ]. T An array geometry 302 can be used to represent the microphones (1, 2, ..., m, ..., M) of FODMA 300, where T is the transpose operator. The maximum distance between two adjacent microphones (e.g., δ) can be assumed. max The wavelength λ of the impact sound wave will be smaller than that of the sound wave.

[0122] As described above regarding the outputs of the two sub-beamformers h1(ω) and h2(ω) (see equations (9) and (10) above), signals from a set of microphones are used for each beamformer, where h1(ω) uses microphones from 1 to M1 and h2(ω) uses microphones from 1 to M2, where {M1, M2} ≤ M, and {} is the union operator. Therefore, the two sub-beamformers h1(ω) and h2(ω) can use all M microphone sensors of the FODMA300 or a subset of the M microphone sensors (e.g., subarray 304).

[0123] Simulation and Experiment

[0124] Figure 4A Figure 400A shows the DF value of FODMA as a function of the coefficients of the target beam pattern according to an embodiment of the present disclosure.

[0125] For a practical or effective target beam pattern, the coefficients in equation (5) above should satisfy the conditions in (8) above. To determine the coefficients a0, a1, and a2, consider the following:

[0126] a0>0, 0<a1≤a0, and a2≥0, (29)

[0127] So that the target beam pattern B1(θ) can be decomposed into:

[0128] B1(θ)=B 1,1 (θ)+B 1,2 (θ), (30)

[0129] B 1,1 (θ)≥0 and B 1,2 (θ)≥0. Based on the condition in (29) above, it can be determined that for any value of a1: B 1,1 (a1, θ) = B 1,1 (-a1, π-θ).

[0130] Furthermore, taking the derivative of equation (5) with respect to θ and setting the result to zero, we obtain:

[0131]

[0132] Combining conditions (8) and (31), we can determine that:

[0133] a0+a1(cos θ d +tan θ d sin θ d )=1. (32)

[0134] The directionality factor (DF) of B1(θ) can be calculated as follows:

[0135]

[0136] It increases as the value of a0 decreases. Substituting equations (31) and (32) into (33), it can be seen that DF depends not only on the coefficients a0 and a1, but also on the steering angle θ. d .

[0137] Figure 4A Figure 400A plots DF as a function of a0. Each θ is set by making a0 = a1. d The starting point a0 at this point gives the maximum DF. It clearly shows that DF increases with θ. d The value decreases as the value increases. Therefore, a0 = a1 is used as the standard for all the following simulations and experiments described in this disclosure.

[0138] Figure 4BFigure 400B shows the DF value of FODMA as a function of steering angle according to an embodiment of the present disclosure.

[0139] Figure 4B Figure 400B shows the θ d The maximum DF of the function. DF with θ d The change from 0 to π / 2 first decreases and then increases, with the maximum value at θ. d =0, and the minimum value is at θ d =π / 3.

[0140] Based on respectively in Figure 4A and Figure 4B As shown in Figures 400A and 400B, the values ​​of a0, a1, and a2 can be determined as follows:

[0141]

[0142] For example, if θ d =π / 4, then as well as In this case, B 1,1 (θ) is a proportional cardioid, B 1,2 (θ) is a proportional dipole along the direction π / 2.

[0143] It should be noted that the decomposition of FODMA beam patterns mentioned earlier can be generalized to higher-order general cases. Based on the multi-level structure in the construction of DMA, the response of an Nth-order DMA is generally equal to the product of the responses of N FODMAs:

[0144]

[0145] Figure 5A Figure 500A shows the beam pattern of FODMA at a selected steering angle according to an embodiment of the present disclosure.

[0146] To study the performance of the method described in this paper, a uniform linear array consisting of three microphones (e.g., Figure 3 Simulations and experiments were conducted using a FODMA 300 with M=3. The spacing (δ) between adjacent microphones was 1 cm. Both the target and designed beam patterns were plotted on... Figures 5A-5D The beam pattern is displayed at f = 1 kHz and θ. d =60°. As clearly shown in Figure 500A, the designed and target beam patterns are almost identical (e.g., the lines representing each beam pattern on Figure 700A are indistinguishable from each other).

[0147] Figure 5B Figure 500B shows the DF value of FODMA as a function of frequency according to an embodiment of the present disclosure.

[0148] As can be seen from Figure 500B, for a specific θ d The value of DF is almost constant within the frequency range studied. This characteristic can be very important for processing wideband signals such as speech.

[0149] Figure 5C Figure 500C shows a beam pattern of FODMA as a function of frequency according to an embodiment of the present disclosure.

[0150] The frequency independence of the designed beam pattern is further verified by Figure 500C, in which the designed beam pattern is frequency invariant.

[0151] Figure 5D A graph showing the approximate error between the target beam pattern of FODMA as a function of frequency and the beam pattern of a steerable beamformer according to an embodiment of the present disclosure.

[0152] The distance between the designed beam pattern and the target beam pattern can be calculated using the following formula:

[0153]

[0154] The results are plotted in Figure 500D under the conditions of M1=2, M2=3, and δ=1cm. It can be easily seen that the difference between the designed beam pattern and the target beam pattern is very small in Figure 500D.

[0155] Figure 6A A spectrogram 600A of clean speech from a steerable beamformer according to an embodiment of the present disclosure is shown, wherein the speech source is located at a selected steering angle.

[0156] In another simulation ( Figures 6A-6C The described method is evaluated by examining their speech enhancement performance. This is done using previous simulations (see [link to previous simulation]). Figures 5A-5D The same microphone array was used in the study. The speech source (spoken by a female speaker) was taken from the TIMIT database of phoneme and word-transcribed speech (see TIMIT Acoustic Phoneme Continuous Speech Corpus. LinguisticData Consort, 1993), located at θ. d =60°. Place the car noise at 180° (end-firing direction) to simulate the noise source. Figures 6A-6C Spectral graphs of clean speech, noisy speech, and enhanced speech from the designed beamformer were plotted. (See also the spectrum of noisy speech). Figure 6B Compared to the previous version, it can be seen that the noise in the enhanced speech spectrum is significantly reduced (see [reference]). Figure 6CWe used signal-to-noise ratio (SNR) as a performance metric. When the input SNR was 5dB, the output SNR after beamforming was 18.25dB, which is consistent with the theoretical results of FODMA speech enhancement.

[0157] Figure 6B A spectrogram 600B of a noisy speech signal from a steerable beamformer according to an embodiment of the present disclosure is shown, wherein the speech source is located at a selected steering angle.

[0158] Figure 6C A spectrogram 600C of an enhanced speech signal from a steerable beamformer according to an embodiment of the present disclosure is shown, wherein the speech source is located at a selected steering angle.

[0159] As mentioned above, Figures 6A-6C Spectral graphs of clean speech, noisy speech, and enhanced speech from the designed beamformer were plotted. (See also the spectrum of noisy speech). Figure 6B Compared to the previous version, it can be seen that the noise in the enhanced speech spectrum is significantly reduced (see [reference]). Figure 6C ).

[0160] Figure 7A Figure 700A shows a target beam pattern for FODMA and a beam pattern of a steerable beamformer according to an embodiment of the present disclosure.

[0161] To further validate the performance of the method described in this paper, a uniform linear array consisting of three microphones was used. The uniform microphone spacing δ was 1.1 cm. The described beamforming algorithm was encoded into the DSP processor of the designed FODMA system. The system was then tested on top of a rotating platform in an anechoic chamber. A loudspeaker was placed on the same horizontal plane as the FODMA array to simulate a sound source. The platform rotated clockwise at 5° intervals. The beam pattern was obtained by measuring the FODMA array gain at each angle based on a reference input signal (e.g., the loudspeaker) and the beamforming output. The results at two different steering angles and frequencies are plotted on... Figure 7A and Figure 7B middle. Figure 7A Conditions: f = 610 Hz and θ d =60°.

[0162] As can be clearly seen from Figures 700A and 700B, the measured beam pattern (solid line) is close to the target beam pattern (dashed line), although there are some differences, which may be caused by a variety of reasons, such as measurement error.

[0163] Figure 7B Figure 700B shows the target beam pattern of FODMA and the beam pattern of a steerable beamformer according to an embodiment of the present disclosure.

[0164] Figure 7B Conditions: f = 2100 Hz and θ d =90°. As can be clearly seen from Figures 700A and 700B, the measured beam pattern (solid line) is close to the target beam pattern (dashed line), although there are some differences, which may be caused by a variety of reasons, such as measurement error.

[0165] Figure 8 This is a block diagram illustrating a machine in an example form of a computer system 800, in which an instruction set or sequence of instructions can be executed to cause the machine to perform any of the methods discussed herein.

[0166] In alternative implementations, the machine operates as a standalone device or can be connected (e.g., networked) to other machines. In a network deployment, the machine can operate as a server or client machine in a server-client network environment, or it can act as a peer machine in a peer-to-peer (or distributed) network environment. The machine can be an in-vehicle system, wearable device, personal computer (PC), tablet computer, hybrid tablet computer, personal digital assistant (PDA), mobile phone, or any machine capable of executing instructions (sequentially or otherwise) specifying the actions the machine should take. Furthermore, although only a single machine is shown, the term "machine" should also be considered to include any collection of machines that individually or jointly execute a set (or more sets) of instructions to perform any one or more methods discussed herein. Similarly, the term "processor-based system" should be considered to include any collection of one or more machines controlled or operated by a processor (e.g., a computer) to individually or jointly execute instructions to perform any one or more methods discussed herein.

[0167] Example computer system 800 includes at least one processor 802 (e.g., a central processing unit (CPU), a graphics processing unit (GPU) or both, a processor core, a computing node, etc.), main memory 804, and static memory 806, which communicate with each other via link 808 (e.g., a bus). Computer system 800 may also include a video display unit 810, an alphanumeric input device 812 (e.g., a keyboard), and a user interface (UI) navigation device 814 (e.g., a mouse). In one embodiment, display device 810, input device 812, and UI navigation device 814 are integrated into a touchscreen display. Computer system 800 may additionally include a storage device 816 (e.g., a drive unit), a signal generation device 818 (e.g., a speaker), a network interface device 820, and one or more sensors 822, such as a global positioning system (GPS) sensor, a compass, an accelerometer, a gyroscope, a magnetometer, or other sensors.

[0168] Storage device 816 includes machine-readable medium 824 on which one or more sets of data structures and instructions 826 (e.g., software) are stored, which are embodied in or utilized by any one or more methods or functions described herein. Instructions 826 may also reside wholly or at least partially within main memory 804, static memory 806, and / or processor 802 during execution by computer system 800, wherein main memory 804, static memory 806, and processor 802 also constitute machine-readable medium.

[0169] Although machine-readable medium 824 is shown as a single medium in the example embodiment, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instructions 826. The term "machine-readable medium" should also be used to include any tangible medium capable of storing, encoding, or carrying instructions that are executed by a machine and cause the machine to perform any one or more methods of this disclosure, or capable of storing, encoding, or carrying data structures used by or associated with such instructions. Specific examples of machine-readable media include volatile or non-volatile memories, such as, but not limited to, semiconductor storage devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0170] Instruction 826 can also be transmitted or received over communication network 828 using a transmission medium via network interface device 820 using any of a variety of well-known transmission protocols (e.g., HTTP). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, mobile phone networks, conventional telephone (POTS) networks, and wireless data networks (e.g., Wi-Fi, 3G and 4G LTE / LTE-A, or WiMAX networks). Input / output controller 830 can receive input and output requests from central processing unit 802 and then send device-specific control signals to the devices it controls (e.g., display device 810). Input / output controller 830 can also manage data flows to and from computer system 800. This frees central processing unit 802 from involvement in the details of controlling each input / output device.

[0171] language

[0172] Numerous details have been set forth in the foregoing description. However, it will be apparent to those skilled in the art who benefit from this disclosure that it can be practiced without these specific details. In some instances, well-known structures and apparatuses are shown in block diagrams rather than in detail to avoid obscuring this disclosure.

[0173] Some parts of the detailed description have been presented based on the algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate their work to others skilled in the art. Algorithms are, and are generally considered, a self-consistent sequence of steps that leads to a desired result. These steps are those that require physical operations on physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. Sometimes, primarily for common reasons, it has proven convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.

[0174] However, it should be remembered that all these and similar terms will be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless explicitly stated otherwise in the following discussion, it should be understood that throughout the description, discussions using terms such as “segmentation,” “analysis,” “determine,” “enable,” “identify,” “modify,” etc., refer to the actions and processes of a computer system or similar electronic computing device that manipulate and convert data represented as physical (e.g., electronic) quantities in the computer system’s registers and memories into other data represented as physical quantities within the computer system’s memory or other such information storage, transmission, or display devices.

[0175] The terms “example” or “exemplary” as used herein are intended to be used as examples, instances, or illustrations. Any aspect or design described herein as “example” or “exemplary” is not necessarily to be construed as preferred or superior to other aspects or designs. Rather, the use of words such as “example” or “exemplary” is intended to present concepts in a specific manner. As used in this application, the term “or” is intended to mean inclusive “or” rather than exclusive “or.” That is, unless otherwise stated or the context clearly indicates, “X includes A or B” is intended to mean any naturally inclusive permutation. That is, if X includes A; X includes B; or X includes A and B, then “X includes A or B” is satisfied in any of the foregoing cases. Furthermore, the articles “a” and “an” as used in this application and the appended claims should generally be construed as meaning “one or more” unless otherwise stated or clearly indicated from the context to be in the singular form. In addition, the terms “one embodiment” or “an implementation” or “an implementation” as used throughout do not imply the same embodiment or implementation unless so described.

[0176] Throughout this specification, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing throughout this specification do not necessarily refer to the same embodiment. Furthermore, the term "or" is intended to indicate inclusive "or" rather than exclusive "or".

[0177] It should be understood that the above description is intended to be illustrative and not restrictive. Many other embodiments will be apparent to those skilled in the art upon reading and understanding the above description. Therefore, the scope of this disclosure should be determined by reference to the appended claims and the full scope of their equivalents.

Claims

1. A method for constructing a first-order differential microphone array (FODMA) with a steerable beamformer, the method comprising: Organize the microphones on a basic planar platform; The target beam pattern of the FODMA is specified by the processing device at the steering angle θ; The processing device decomposes the target beam pattern into a first sub-beam pattern and a second sub-beam pattern based on the steering angle θ; The processing device generates a first sub-beamformer and a second sub-beamformer for each filter signal from the microphone of the FODMA, wherein the first sub-beamformer is associated with the first sub-beam pattern and the second sub-beamformer is associated with the second sub-beam pattern; as well as The steerable beamformer is generated by the processing device based on the first sub-beamformer and the second sub-beamformer.

2. The method according to claim 1, wherein, The steering angle θ∈[0,π].

3. The method according to claim 1, wherein, The processing device further includes decomposing the target beam pattern into a first sub-beam pattern and a second sub-beam pattern by: dividing the target beam pattern into a first-order cosine (cardioid) first sub-beam pattern and a first-order sine (dipole) second sub-beam pattern.

4. The method according to claim 1, wherein, The processing device for generating a first sub-beamformer and a second sub-beamformer for each filtered signal from the microphone of the FODMA further includes: the second sub-beamformer filtering the squared signal from the microphone of the FODMA to substantially match the second sub-beam pattern.

5. The method according to claim 4, further comprising: The second sub-beamformer ignores any signal correlation when filtering the squared signal from the microphone of the FODMA to substantially match the second sub-beam pattern.

6. The method according to claim 1, wherein, The process of generating the steerable beamformer by the processing device based on the first sub-beamformer and the second sub-beamformer further includes generating the steerable beamformer based on the spectral phase of the filtered signal from the first sub-beamformer.

7. The method according to claim 1, further comprising: The microphones of the FODMA are organized as a uniform linear differential microphone array (LDMA), wherein the microphones are equidistant along a straight line.

8. A method for constructing a first-order differential microphone array (FODMA) with a steerable beamformer, the method comprising: Multiple (M) microphones are organized on a basic plane platform, the multiple microphones comprising a first subset (M1) of microphones and a second subset (M2) of microphones. A first sub-beamformer is constructed by a processing device based on a first subset (M1) of the microphone and a target beam pattern at a turning angle θ, wherein the first sub-beamformer is characterized according to a first-order cosine (cardioid) first sub-beam pattern. The processing device constructs a second sub-beamformer based on a second subset (M2) of the microphone and a target beam pattern at the steering angle θ, wherein the second sub-beamformer is characterized according to a first-order sinusoidal (dipole) second sub-beam pattern; and The steerable beamformer is generated by the processing device based on the first sub-beamformer and the second sub-beamformer.

9. The method according to claim 8, wherein, The steering angle θ∈[0,π].

10. The method according to claim 8, wherein, The processing device for generating a first sub-beamformer and a second sub-beamformer for each filtered signal from the microphone of the FODMA further includes: the second sub-beamformer filtering the squared signal from the microphone of the FODMA to substantially match the second sub-beam pattern.

11. The method of claim 10, further comprising: The second sub-beamformer ignores any signal correlation when filtering the squared signal from the microphone of the FODMA to substantially match the second sub-beam pattern.

12. The method according to claim 8, wherein, The process of generating the steerable beamformer by the processing device based on the first sub-beamformer and the second sub-beamformer further includes generating the steerable beamformer based on the spectral phase of the filtered signal from the first sub-beamformer.

13. The method of claim 8, further comprising: The microphones of the FODMA are organized as a uniform linear differential microphone array (LDMA), wherein the microphones are equidistant along a straight line.

14. A first-order differential microphone array FODMA system with a steerable beamformer, the system comprising: A microphone located on a platform situated on a basic plane; as well as The processing device, communicatively coupled to the microphone, is configured to: Specify the target beam pattern of the FODMA at the steering angle θ; Based on the steering angle θ, the target beam pattern is decomposed into a first sub-beam pattern and a second sub-beam pattern; Generate a first sub-beamformer and a second sub-beamformer for each filtered signal from the microphone, wherein the first sub-beamformer is associated with a first sub-beam pattern, and the second sub-beamformer is associated with a second sub-beam pattern; and The steerable beamformer is generated based on the first sub-beamformer and the second sub-beamformer.

15. The FODMA system according to claim 14, wherein, The steering angle θ∈[0,π].

16. The FODMA system according to claim 14, wherein, The processing device is further configured to divide the target beam pattern into a sum of a first-order cosine (cardioid) sub-beam pattern and a first-order sine (dipole) sub-beam pattern.

17. The FODMA system according to claim 14, wherein, The processing device is further configured to filter the squared signal from the microphone using the second sub-beamformer to substantially match the second sub-beam pattern.

18. The FODMA system according to claim 17, wherein, The processing device is further configured to ignore any signal correlation when filtering the squared signal from the microphone using the second sub-beamformer to substantially match the second sub-beam pattern.

19. The FODMA system according to claim 14, wherein, The processing apparatus is further configured to generate the steerable beamformer based on the spectral phase of the filtered signal from the first sub-beamformer.

20. The FODMA system according to claim 14, wherein, The microphones of the FODMA are configured as a uniform linear differential microphone array (LDMA), with the microphones equidistantly spaced along a straight line.

21. A first-order differential microphone array FODMA system with a steerable beamformer, the system comprising: Multiple (M) microphones are located on a substantially planar platform, the multiple microphones comprising a first subset M1 of microphones and a second subset M2 of microphones; as well as The processing device, communicatively coupled to the plurality of microphones, is configured to: A first sub-beamformer is constructed based on a first subset M1 of the microphone and a target beam pattern at a steering angle θ, wherein the first sub-beamformer is characterized according to a first-order cosine (cardioid) first sub-beam pattern; A second sub-beamformer is constructed based on a second subset M2 of the microphone and a target beam pattern at the steering angle θ, wherein the second sub-beamformer is characterized according to a first-order sinusoidal (dipole) second sub-beam pattern; and The steerable beamformer is generated based on the first sub-beamformer and the second sub-beamformer.

22. The FODMA system according to claim 21, wherein: The steering angle θ∈[0,π]; and M1≥2 and M2≥3.

23. The FODMA system according to claim 21, wherein, The processing device is further configured to filter the squared signal from the microphone M2 using the second sub-beamformer to substantially match the second sub-beam pattern.

24. The FODMA system according to claim 23, wherein, The processing device is also configured to ignore any signal correlation when filtering the squared signal from the microphone from M2 using the second sub-beamformer to substantially match the second sub-beam pattern.

25. The FODMA system according to claim 21, wherein, The processing device is further configured to: The signal from the microphone in M1 is filtered using the first sub-beamformer; and The steerable beamformer is generated based on the spectral phase of the filtered signal from the first sub-beamformer.

26. The FODMA system according to claim 21, wherein, The microphones of the FODMA are configured as a uniform linear differential microphone array (LDMA), with M microphones equidistantly spaced along a straight line.

Citation Information

Patent Citations

  • Device and method for direction dependent spatial noise reduction

    CN102771144A

  • Conference system with a microphone array system and a method of speech acquisition in a conference system

    CN108370470A