Method and system for directional processing of speech information
The use of non-collinear microphone pairs with dual-delay subtraction algorithms enhances microphone array performance, achieving efficient and scalable directional audio processing with improved low-frequency performance and reduced computational complexity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- イギリス国
- Filing Date
- 2022-04-20
- Publication Date
- 2026-05-21
AI Technical Summary
Conventional directional microphone arrays face challenges in maintaining directivity across a wide bandwidth of audible frequencies, requiring larger sizes and compromising portability, while existing beamforming algorithms struggle with low-frequency performance and computational efficiency.
A method and apparatus using a digital signal processor for a microphone array with non-collinear pairs of microphones, employing dual-delay subtraction algorithms and frequency equalization to control directivity, allowing for compact, efficient, and scalable audio signal processing.
The system provides improved directivity and broadband suppression, maintaining good low-frequency performance with reduced computational cost, suitable for portable and mobile applications.
Smart Images

Figure 0007863576000003 
Figure 0007863576000004 
Figure 0007863576000005
Abstract
Description
[Technical Field]
[0001] This application relates to a method and apparatus for directional processing of speech information in a digital signal processor for a microphone array or speaker array. Specifically, this application relates to a method and apparatus for processing speech signals related to speech beamforming. [Background technology]
[0002] Audio beamforming can be applied to a pair of microphones, so that the outputs of that pair of microphones are combined to amplify sound from a given direction that is attenuated relative to sound from all other directions. When applied to speakers instead, audio beamforming drives a pair of speakers to provide a highly directional sound output. Each pair is sometimes called an array.
[0003] These directional microphone arrays are sometimes used as a solution to the "cocktail party effect," which can make it difficult for a person to focus on a single speaker or a specific sound source, such as a conversation, in a noisy environment, and can be even more difficult to distinguish if the sound is first recorded by the microphone and then played back, for example, through a hearing aid, or if the recording is played back later. This is partly because the neurological processing of the human brain in the live environment and associated sound filtering is lost, and therefore the intelligibility of speech or other sound sources is significantly reduced.
[0004] This problem arises in many areas of modern life where it is necessary to distinguish a single speaker or sound source from ambient noise, for example, to assist human hearing, to detect the location of an object known to be associated with a specific vocal signature such as an animal, vehicle, drone, or other aircraft, or to detect the direction of signs of life in an emergency rescue situation.
[0005] While known directional microphone arrays can exhibit good directivity in specific frequency bands, maintaining this directivity at lower frequencies using known methods requires microphone arrays of increasingly larger dimensions. For example, several arrays with a diameter of approximately 2 meters have been proposed. Thus, conventional techniques have an undesirable trade-off between the size of the microphone array and the uniformity of response and directivity across a wide bandwidth of audible frequencies.
[0006] One known example of a beamforming algorithm is the delayed-sum algorithm, in which a delay is added to one or more audio signals (received from each microphone or applied to each speaker) so that all signals in the intended beamforming direction are synchronized and added to each other in a constructive manner. Conversely, signals in other directions are not synchronized and tend to cancel each other out by adding them destructively. This can produce a desired narrow beam at high frequencies, but the beam at low frequencies is typically very wide in comparison. This means that the array is typically equally sensitive in all directions (omnidirectionally) at these low frequencies.
[0007] To overcome these drawbacks in the low-frequency range, conventional methods of delayed-sum beamforming algorithms use increasingly larger microphone / speaker arrays. Increasing the physical size of the array is also desirable because it reduces the number or influence of grating lobes corresponding to angles outside the intended beam, which makes the array highly sensitive in a particular frequency range. However, this physically larger size of the array is, naturally, undesirable for manufacturing, portability, and the practicality of such technology.
[0008] Another example of a beamforming algorithm is the delayed-subtraction algorithm, in which a delay is added to one of a pair of microphones, and then the audio signal is subtracted from it. If the sound source is perpendicular to the axis through which the pair of microphones passes, the sound will reach both microphones simultaneously. If no delay is added to one of the audio signals before subtraction, the audio signals from the sound source will cancel each other out at each microphone. This direction, where the audio signal is not detected (because it is canceled out), is called a null, and the response or sensitivity of the microphone pair in this direction will be zero. There will also be a second null on the opposite side of the axis.
[0009] By adding a delay to one of the audio signals before subtraction, these null directions are shifted from a direction perpendicular to the axis passing through the pair of microphones. Referring to Figure 1, the solid arrows correspond to the directions of arrival of sound waves entering a pair of microphones A and B separated by a distance d from each other, and the angle θ corresponds to the angle of entry of the sound waves relative to the axis passing through the pair of microphones. As shown in Figure 1, the following calculations are based on the assumption that the sound sources are far enough apart that the wavefront of the sound wave is close to a flat wavefront. In this case, the audio signal reaches microphone A at time t, and then reaches microphone B at time (t+T), and the null difference equation is as follows.
[0010] A(t)-B(t+T)=0 (1) Here, T is given by the following:
[0011]
number
[0012] B(t)-A(tT)=0 (3)
[0013] In two dimensions, there are two null directions, both at an angle θ with respect to the axis. In three dimensions, this forms a null cone with a conical aperture / apex angle of 2θ, as shown in Figure 2. Audio signals arriving from directions away from these nulls are not canceled out, but as a result of the nature of this arrangement, sensitivity to low frequencies where the wavelength of the signal is much greater than the distance between the microphones in the pair is reduced. To address this, it is common to filter the output of such a delayed subtraction pair to increase the gain in the lower frequency bands in order to obtain a flat frequency response of the audio signal emanating from the beamforming direction. However, this also raises the noise floor in these frequency bands, and therefore needs to be balanced when appropriately setting this white noise gain. This filter is sometimes called a frequency equalization filter or normalization filter.
[0014] For a delayed subtraction pair with a time delay set to show null at approximately ±110°, an example of a normalized (filtered with a frequency equalization filter as described above) response may be as shown in Figure 3, where the x-axis represents the azimuth angle in degrees from the intended beamforming direction, and the y-axis is a logarithmic scale of frequency in kHz, with brighter areas corresponding to higher dB sensitivity at that frequency and angle. Audio signal B of this microphone pair diff (t) can be written as follows:
[0015] B diff (t) = B(t) - A(t-T1) (4)
[0016] B diff(t) becomes zero (i.e., null) when Equation 2 holds. This pair can be regarded as a multi-sensor microphone and can be used as one of the additional pairs of microphones arranged in a linear array and can be considered to be input into a delay subtraction beamforming algorithm. Specifically, a pair of microphones A and B can be subjected to delay subtraction beamforming to generate an audio signal associated with pair AB, and another pair of microphones C and D (on the same line as the AB pair) can be subjected to delay subtraction beamforming to generate an audio signal associated with pair CD. Thereafter, the audio signal of AB and the audio signal of CD can be subjected to delay subtraction beamforming. This may be referred to as a collinear double delay subtraction beamforming array.
[0017] When the distance between the microphone CD pair is set equal to the distance between the microphone pair AB, the audio signal D of the microphone CD pair diff (t) can be rewritten as follows.
[0018] D diff (t)=D(t)-C(t-T1) (5)
[0019] Mathematically, B diff and D diff are regarded as a new pair of microphones for double delay subtraction, and the following equation is given.
[0020] D double-dif (t)=D diff (t)-B diff (t-T2) (6) D double-diff (t)=D(t)-C(t-T1)-B(t-T2)+A(t-T1-T2) (7) Here, T1 corresponds to the time difference in arrival of the audio signal between A and B and between C and D, and T2 corresponds to the time difference in arrival of the audio signal between the AB pair and the CD pair based on the distance between microphone B and microphone D.
[0021] This can be represented schematically as shown in FIG. 4, where there are two null cones. The θ1 null cone is for B diff and D diff corresponding to the audio signal and can be steered for a given spacing between microphones in each of the AB pair and the CD pair by adjusting the delay T1. The θ2 null cone is for D double-diff corresponding to the audio signal and can be steered for a given spacing between the AB pair and the CD pair by adjusting the delay T2.
[0022] The effect on the directivity frequency response of the linear microphone array of the double delay subtraction method to produce two null cones can be seen in FIG. 5, where the null cones are shown at ±90° and ±150°. Additional nulls result in wider audio signal rejection at angles away from the beamforming angle. However, especially at higher frequencies, there is still a significant amount of sensitivity between these two null cones, and it can be seen that these techniques perform worse than the delay-and-sum beamforming algorithm.
[0023] One option to enhance the controllability of the steering of the sensitivity of the microphone array is to use two pairs of zero-delay subtraction microphone pairs in a right-angled orientation rather than two pairs in a linear orientation using double delay subtraction. When the separation between microphones in each pair is significantly smaller than the wavelength of the sound of interest, the output of the zero-delay subtracted pair is close to the spatial derivative of the sound pressure δp / δx, where x is the distance along the axis of the pair and p is the sound pressure. This pair may be called a dipole, and a dipole with a narrow separation using subtraction may be called a differential pair.
[0024] Using two mutually perpendicular differential pairs to form the gradient vector
Number
[0025] Therefore, the inventors recognized the desire to provide a method and system for directional processing of audio information from a microphone array that is easy to design and implement, highly computationally efficient, and improves white noise gain. Specifically, a compact system with good low-frequency performance that is scalable to a system that can also provide good high-frequency performance, has extremely low computational cost for easy implementation. [Overview of the project]
[0026] The present invention is defined in the independent claims, which are referred to below. Advantageous features are described in the dependent claims.
[0027] A first aspect of this disclosure provides a method for directional processing of speech information in a digital signal processor for a microphone array. The method includes receiving a speech signal from each microphone of the microphone array, the microphone array comprising at least two pairs of parallel microphones that are not collinear with each other, the first pair of microphones comprising a first microphone and a second microphone, the second pair of microphones comprising a third microphone and a fourth microphone, and each microphone being arranged on a plane.
[0028] The method further includes adding a first time delay to the audio signal from the second microphone and subtracting the delayed audio signal from the second microphone from the audio signal from the first microphone to determine the audio signal associated with the first pair of microphones; adding a second time delay to the audio signal from the fourth microphone and subtracting the delayed audio signal from the fourth microphone from the audio signal from the third microphone to determine the audio signal associated with the second pair of microphones; and adding a third time delay to the audio signal associated with the second pair of microphones and subtracting the delayed audio signal associated with the second pair of microphones from the audio signal associated with the first pair of microphones to determine the audio signal associated with the microphone array, wherein the first, second, and third time delays are configured to control the directivity of the audio signal associated with the microphone array.
[0029] In this way, the method advantageously provides a directional frequency response with broadband suppression away from the direction of interest for sound source isolation / locating in the plane of the microphone array.
[0030] Optionally, the first time delay may be set based on the relative distance between the first microphone and the second microphone, the second time delay may be set based on the relative distance between the third microphone and the fourth microphone, and the third time delay may be set based on the relative distance between the first pair of microphones and the second pair of microphones.
[0031] Optionally, the distance between microphones in a pair of microphones may be uniform across all microphone pairs, resulting in the first and second time delays being equal to each other. By matching the microphone pairs in this way, the combined outputs of each pair have the same characteristics, thereby simplifying the associated processing of each output. Optionally, the microphone array may follow a specific embodiment of two pairs of microphones, with four microphones positioned at the vertices of a parallelogram.
[0032] Optionally, a frequency equalization filter may be applied to the audio signal associated with the microphone array. This frequency equalization filter can be tuned to compensate for attenuation in the frequency response of the output processed by the array's algorithm to provide a normalized, flat frequency response over the operating frequency range. These attenuations may result from defects in the frequency response of the physical microphones constituting the array, and / or from the delay subtraction process itself. In practice, lower frequencies are typically attenuated more significantly by the delay subtraction algorithm than higher frequencies, and therefore the frequency equalization filter will typically be tuned to increase the gain of these lower frequencies so that the output more accurately reflects the observed real-world sound.
[0033] Optionally, each microphone in the array may be selected to have an omnidirectional polarity pattern. This is advantageous because it simplifies the visualization / simulation of the array response during the design phase of fabricating a given system. Furthermore, by selecting microphones in the array to have the same response to each other, the processing of the audio signals detected by each physical microphone is simplified because it is not necessary to consider different polarity response patterns, for example, by introducing weighting to the audio signal associated with one of the microphones during subtraction calculations.
[0034] Optionally, the determined audio signals associated with the microphone array may correspond to a given direction based on first, second, and third time delay values, and the method may further include iteratively adjusting the first, second, and third time delays and determining a plurality of corresponding audio signals associated with the iteratively adjusted first, second, and third time delays. By comparing the plurality of corresponding audio signals, the direction closest to a desired audio signature or sound source can be determined. This is advantageous because it enables a sweep adjustment of the directivity of the microphone array for use in determining the direction of arrival of an audio signature being observed or to be observed.
[0035] Optionally, the method further includes determining one or more respective audio signals associated with one or more corresponding further microphone arrays arranged on a plane, and combining each audio signal associated with the microphone arrays to determine the audio signal associated with the array of microphone arrays using a beamforming algorithm. By combining these arrays of microphone arrays into a larger macroarray, improved performance is achieved compared to a single microphone array. Specifically, the directivity of the macroarray can be optimized by combining the null directions of each constituent microphone array to form a null that overlaps in a direction away from the beamforming direction or the direction of interest / maximum sensitivity direction. It is advantageous that the nulls of each microphone array can be aimed to overlap with weaknesses in the existing macroarray, or vice versa.
[0036] Optionally, the beamforming algorithm may be selected as a delayed-add beamforming algorithm. This advantageously combines delayed-add and delayed-subtract algorithms. For example, the delayed-subtract used in a microphone array can be optimized to provide broadband rejection in directions away from the beamforming direction even at lower frequencies, and the delayed-add for each microphone array can be optimized to narrow the width of the main beam in the beamforming direction even at higher frequencies.
[0037] Optionally, multiple microphone arrays may be tessellated such that at least one microphone is shared between two adjacent microphone arrays. This is advantageous because the output of a single physical microphone can be used as input to a delay subtraction algorithm of multiple adjacent microphone arrays, thereby reducing the total number of physical microphones required to form a macroarray.
[0038] A second aspect of the present disclosure provides an apparatus for directional processing of audio information for a microphone array. The apparatus includes one or more input units configured to receive audio signals from each microphone in the microphone array, the microphone array including at least two pairs of parallel microphones that are not collinear with each other. A first pair of microphones includes a first microphone and a second microphone, and the second pair of microphones includes a third microphone and a fourth microphone, each microphone arranged on a plane. The apparatus further includes a digital signal processor configured to add a first time delay to the audio signal from the second microphone and subtract the delayed audio signal of the second microphone from the audio signal from the first microphone to determine the audio signal associated with the first pair of microphones, and to add a second time delay to the audio signal from the fourth microphone and subtract the delayed audio signal of the fourth microphone from the audio signal from the third microphone to determine the audio signal associated with the second pair of microphones. The digital signal processor is further configured to add a third time delay to the audio signals associated with the second pair of microphones and to subtract the delayed audio signals associated with the second pair of microphones from the audio signals associated with the first pair of microphones in order to determine the audio signals associated with the microphone array. Here, the first, second, and third time delays are configured to control the directivity of the audio signals associated with the microphone array.
[0039] Optionally, a first time delay is set based on the relative distance between the first microphone and the second microphone, a second time delay is set based on the relative distance between the third microphone and the fourth microphone, and a third time delay is set based on the relative distance between the first pair of microphones and the second pair of microphones.
[0040] Optionally, the distance between microphones in a pair of microphones is uniform across all microphone pairs, and therefore the first and second time delays are set to be equal to each other.
[0041] Optionally, the microphone array includes two pairs of microphones, with the first, second, third, and fourth microphones positioned at the vertices of a parallelogram. Optionally, a digital signal processor is further configured to apply a frequency equalization filter to the audio signals associated with the microphone array. Optionally, the received audio signals correspond to the microphones of the microphone array having an omnidirectional polarity pattern.
[0042] Optionally, the digital signal processor is further configured to determine, based on first, second, and third time delay values, that an audio signal associated with a microphone array corresponds to a given direction; the digital signal processor is configured to iteratively adjust the first, second, and third time delays and determine a plurality of corresponding audio signals associated with the iteratively adjusted first, second, and third time delays; and further, to compare a plurality of corresponding audio signals to determine the direction that most closely corresponds to a desired audio signature or sound source.
[0043] Optionally, the digital signal processor is further configured to determine one or more audio signals associated with one or more corresponding further microphone arrays arranged on a plane, and to combine each audio signal associated with the microphone arrays to determine the audio signal associated with the array of microphone arrays using a beamforming algorithm. Optionally, a delayed-sum beamforming algorithm is selected as the beamforming algorithm.
[0044] Optionally, multiple microphone arrays are tessellated from one another such that at least one microphone is shared between two adjacent microphone arrays.
[0045] A third aspect of this disclosure discloses a system for directional processing of speech information for a microphone array. The system includes the apparatus of a second aspect of this disclosure and a microphone array. The microphone array includes at least two pairs of parallel microphones that are not collinear with each other, the first pair of microphones including a first microphone and a second microphone, the second pair of microphones including a third microphone and a fourth microphone, and each microphone is arranged on a plane.
[0046] A fourth aspect of the present disclosure provides a method for directional processing of audio information in a digital signal processor for a speaker array. The method includes receiving an audio signal transmitted by a speaker array and determining which audio signal to apply to each speaker of the speaker array, wherein the speaker array includes at least two pairs of parallel speakers that are not collinear with each other, the first pair of speakers including a first speaker and a second speaker, the second pair of speakers including a third speaker and a fourth speaker, and each speaker is arranged on a plane.
[0047] Each audio signal applied to each speaker is determined by: determining that the audio signal associated with the first pair of speakers is the audio signal transmitted by the speaker array; applying a delay and inversion algorithm to the audio signal transmitted by the speaker array using a first time delay to determine that the audio signal associated with the second pair of speakers is the audio signal associated with the first pair of speakers; applying a delay and inversion algorithm to the audio signal associated with the first pair of speakers using a second time delay to determine that the audio signal applied to the first speaker is the audio signal associated with the first pair of speakers; determining that the audio signal applied to the third speaker is the audio signal associated with the second pair of speakers; and applying a delay and inversion algorithm to the audio signal associated with the second pair of speakers using a third time delay to determine that the audio signal applied to the fourth speaker is the audio signal associated with the second pair of speakers. The first, second, and third time delays are configured to control the directivity of the audio signals transmitted by the speaker array.
[0048] In this way, this method is advantageous because it provides a speaker array that can transmit a highly directional sound wave beam with broadband suppression away from the direction of the sound wave beam, while maintaining a relatively uniform frequency response for transmission within the beamwidth of the speaker array.
[0049] A fifth aspect of the present disclosure provides an apparatus for directional processing of speech information in a digital signal processor for a speaker array. The apparatus includes an input unit configured to receive a speech signal transmitted by a speaker array, the speaker array including at least two pairs of parallel speakers that are not collinear with each other, the first pair of speakers including a first speaker and a second speaker, the second pair of speakers including a third speaker and a fourth speaker, and each speaker is arranged in a plane.
[0050] The device further includes a digital signal processor configured to determine which audio signal to apply to each speaker of the speaker array by determining that the audio signal associated with a first pair of speakers is an audio signal transmitted by the speaker array, applying a delay and inversion algorithm to the audio signal transmitted by the speaker array using a first time delay in order to determine which audio signal is associated with a second pair of speakers, determining that the audio signal applied to the first speaker is an audio signal associated with the first pair of speakers, applying a delay and inversion algorithm to the audio signal associated with the first pair of speakers using a second time delay in order to determine which audio signal is applied to the second speaker, determining that the audio signal applied to the third speaker is an audio signal associated with the second pair of speakers, and applying a delay and inversion algorithm to the audio signal associated with the second pair of speakers using a third time delay in order to determine which audio signal is applied to the fourth speaker. The first, second, and third time delays are configured to control the directivity of the audio signal transmitted by the speaker array.
[0051] Optionally, a first time delay is set based on the relative distance between a first speaker and a second speaker, a second time delay is set based on the relative distance between a third speaker and a fourth speaker, a third time delay is set based on the relative distance between a first pair of speakers and a second pair of speakers, the distance between speakers in a pair of speakers is uniform across all speaker pairs, the first time delay and the second time delay are equal, and the first, second, third and fourth speakers are located at the vertices of a parallelogram.
[0052] Optionally, the digital signal processor is further configured to determine one or more respective audio signals associated with one or more corresponding further speaker arrays arranged on a plane, using a beamforming algorithm.
[0053] Hereinafter, embodiments of the present invention will be described only as examples, with reference to the accompanying drawings. [Brief explanation of the drawing]
[0054] [Figure 1] This diagram shows sound waves approaching a pair of microphones at an angle to the axis of the microphone pair. [Figure 2] This figure shows a pair of microphones representing a null cone in three dimensions. [Figure 3] This figure shows an example of the normalized directional frequency response of the output of a delay subtraction algorithm applied to a pair of microphones. [Figure 4] This figure shows two pairs of collinear microphones in three dimensions, illustrating the two null cones resulting from the double-delay subtraction algorithm. [Figure 5] This figure shows an example of the normalized directional frequency response of the output of a double-delay subtraction algorithm applied to two pairs of colinear microphones. [Figure 6] This is a block diagram showing a system 10 according to a microphone embodiment of the present disclosure. [Figure 7] This diagram shows the arrangement of four microphones in the shape of a parallelogram. [Figure 8] This block diagram shows a dual-delay subtraction array algorithm for four microphones. [Figure 9] This figure shows two null cones as a result of applying a double-delay subtraction algorithm to the microphone arrangement shown in Figure 7. [Figure 10] This figure shows an example of the normalized directional frequency response of the output of a double-delay subtraction algorithm applied to two pairs of microphones that are not on the same line. [Figure 11] This figure shows an example of the normalized directional frequency response of the output of a delay-sum algorithm applied to multiple subarrays, each using a double-delay subtraction algorithm. [Figure 12] This is a block diagram of system 40 according to the speaker embodiment of the present disclosure. [Modes for carrying out the invention]
[0055] This disclosure relates to a method and apparatus for directional processing of speech information in a digital signal processor for a speech transducer array, and a corresponding system including a speech transducer array. In one embodiment, the speech transducer is a microphone, and the microphone array can be used to detect sound emitted from a given direction. However, those skilled in the art will see that the same operating principle can also be applied in reverse to an array of speakers instead of microphones. In this further embodiment, the speaker array can be used to generate a highly directional sound wave beam.
[0056] Figure 6 is a block diagram showing a system 10 according to one embodiment of the present disclosure. The system 10 includes a device 12 and a microphone array 14. The device 12 includes a plurality of input units 16 for receiving audio signals from each of the microphones in the microphone array, and a digital signal processor 18 that communicates with each of the input units 16. The digital signal processor also communicates with an output module 20 for outputting the results of digital signal processing.
[0057] It can be seen that preamplifiers and analog-to-digital converters for the audio signals captured by the microphones of the microphone array can be placed at any point along the communication path between the microphones 14 and the digital signal processor 18. For example, individual microphones may be connected to one or more inputs 16 of the device 12 for analog-to-digital conversion within the device 14, or analog-to-digital conversion may be performed before the device 12 so that the audio signals from all microphones in the microphone array are received in the digital domain by a single input 16 in the device.
[0058] In one embodiment, the microphone array 14 may include four microphones arranged at the vertices of a parallelogram, as shown in Figure 7. Specifically, the four microphones may be divided into a first pair of microphones A and B and a second pair of microphones C and D. The distance between microphones in each microphone pair is d1, and the distance between each pair of microphones is d2. The inventors have recognized that a delayed subtractive beamforming algorithm can be used with such a microphone array to provide an improved beamforming output.
[0059] The first pair of microphones (AB) and the second pair of microphones (CD) are parallel to each other but not arranged in a straight line, and are therefore referred to as a non-collinear pair of microphones in this specification. For sound received from a sound source in a given direction, the double delay subtraction algorithm follows the same analysis as described with respect to Equation 7 above. Specifically, the sound signal detected by the microphone array 14 can be output as follows.
[0060] D double-diff (t)=D(t)-C(t-T1)-B(t-T2)+A(t-T1-T2) (8)
[0061] The logic of equation (8) is also shown in Figure 8, with the addition of the normalization / frequency equalization filter described above to correct the frequency response of the microphone array 14 to approximate a flat frequency response.
[0062] This double delay subtraction algorithm also produces two null cones in this case, but since these two pairs are not collinear, these two null cones will be oriented in different directions. As can be seen in Figure 8, the null cones corresponding to the two delay subtraction pairs are oriented along the direction of the axes of those pairs and have an angle θ1 based on the distance d1 and the time delay T1. Next, proceeding to the double delay subtraction step of this algorithm, the outputs of the first pair of microphones A and B and the second pair of microphones C and D are mathematically treated as outputs received in their respective single hypothetical / virtual microphones. Therefore, the axis passing through these two virtual microphones will be along the direction between microphones B and D (or similarly A and C), and thus the additional null cone corresponding to the double subtraction combination will be oriented along this direction, having an angle θ2 based on the distance d2 and the time delay T2.
[0063] In the plane of the microphone array, these two null cones correspond to four null directions (two pairs of nulls). By appropriately configuring the geometry of the microphone array and adjusting the delay applied to the microphone signal, the angles of the null directions can be adjusted / steering to position the nulls in order to construct a series of nulls that connect to each other, forming holes in the sensitivity of the microphone array 14 to reduce the sensitivity of the microphone array in directions other than the direction of interest. Narrow-angle null cones are more effective than wide-angle null cones (e.g., within a 90° region), and therefore, by orienting the null cones differently, it is advantageous to aim the narrow null cones at adjacent angles to improve the rejection of the audio signal in that region.
[0064] When designing such an array, the distance d1 between microphones in each pair should be set to less than half of the shortest wavelength of interest (ideally less than a quarter wavelength). However, minimizing this distance increases the low-frequency noise present in the system output, and therefore, a compromise between these two may be necessary when designing the array. The same logic applies to the selection of the array distance d2. In certain use cases, selecting d1 to be equal to d2 can further simplify the system's operation.
[0065] While the fineness of adjustment of the half-angle θ1 using the value of T1 (see Equation 2) is theoretically infinite, this method and system are most computationally efficient when the value of the T1 duration is set to an integer number of samples, taking into account the sample rate of the recording of the audio signal received by the microphone. This is because, if the delay T1 is not equal to an integer number of samples, the delay value of the audio signal at the delayed microphone would need to be interpolated by a digital signal processor from the values between two adjacent samples. An accurate interpolated audio signal can be determined using a scaled sink function based on the Nyquist-Shannon sampling theorem using known methods. However, optimizing the choice of T1 to an integer number of samples avoids the need for this additional processing.
[0066] This method differs conceptually from conventional techniques that focus on generating an electronically steerable main beam without doing so, where the main beam is the range in which the array's response is at maximum gain, having a uniform response pattern with respect to the steered direction of the microphone array.
[0067] By positioning the array microphones in this non-collinear arrangement, nulls with intersecting orientations can provide wider interference suppression away from the main beam in the array's plane. For example, Figure 10 shows a plot of the directional frequency response for a non-collinear arrangement, which can be directly compared to that of Figure 5. This improvement provides greater flexibility in the null arrangement, allowing the use of in-plane beamforming directions to still be aimed by narrow null cones in an angular range away from the main beam in a non-collinear arrangement, as narrower null cones typically result in a wider suppression range around the target null angle.
[0068] In summary, the proposed arrangement of parallel pairs of microphones that are not collinear brings improvements to various known systems for the following reasons: The use of dual-delay subtractive beamforming for 2D arrays is easy to implement and computationally efficient, making it suitable for low-power and / or mobile systems. • The delay subtraction method is inherently broadband (effective at all frequencies below an upper limit dependent on microphone spacing), and therefore does not require splitting the speech signal into various different frequency bands for independent calculations using corresponding weightings, which would be computationally expensive (this is particularly advantageous for applications involving broadband signals such as speech). • The system's directivity is improved in separating higher-frequency sounds within the plane of the microphone array, while still maintaining good low-frequency performance. • Low-frequency performance can be achieved by increasing the spacing between microphones compared to conventional technology, and this further reduces the system noise at low frequencies (improves white noise gain) compared to conventional differential systems with closerly spaced microphones but the same low-frequency pattern. • Greater spacing between microphones improves system flexibility for designing configurations that use time delays that are integer multiples of the audio sample length to reduce the complexity of the corresponding digital signal processing.
[0069] The systems and methods described above can be applied to arrays of microphones, each having omnidirectional (i.e., the same sensitivity to audio signals coming from all directions) and a flat frequency response (i.e., the same sensitivity to all frequencies). However, this is not mandatory, and microphones with non-omnidirectional polar response patterns can also be used. If the microphones in a pair do not have the same polar response pattern, it may be necessary to take their differences into account by introducing weighting to the audio signals associated with one of the microphones during subtraction calculations. If the microphones have a common directional polar response pattern, this can improve the directivity of the resulting microphone array and simplify the processing of each audio signal, as this additional weighting is not required.
[0070] During the design phase, the array response can be viewed as a superposition of the omnidirectional array response to the actual response of each individual directional microphone in the array. This can be calculated by representing the directional microphone response and the omnidirectional response using complex numbers (representing both amplitude gain and phase shift) and multiplying them together. This simple calculation allows for the design of the final system by considering the performance of each predicted response separately or by multiplying them together. Microphones in such an array will preferably have the same polar response pattern and be oriented in the same direction.
[0071] The above use case involves two pairs of parallel microphones positioned at the vertices of a parallelogram (as shown in Figures 7 and 9), such as a rhombus, rectangle, or square. However, configurations with other shapes are also possible. For example, the distances between each pair of microphones do not have to be equal, thereby changing the parallelogram to a more general scalene quadrilateral. In that case, the distances between the microphones in one pair will be different from the distances between the other parallel pair, thus requiring different time delays for each pair of microphones. Similarly, to match the gain levels of the two pairs, the gain of the audio signal output of one pair may need to be weighted.
[0072] Furthermore, the same process can be extended to larger arrays using additional microphones. For example, the process may be extended to a triple-delay subtraction algorithm applied to four pairs of microphones, thereby applying delay subtraction to each of the four pairs, and the resulting four virtual microphones are used as inputs to the aforementioned double-delay subtraction algorithm. This results in three null cones and therefore six null directions within the plane of the microphone array. While this extension improves the directivity of beamforming, the use of such additional microphones also unnecessarily introduces additional self-inducing microphone noise.
[0073] Alternatively, the system can be extended by treating the aforementioned array as a subarray of multiple corresponding subarrays that combine to form a macroarray. Each subarray can be considered a new virtual microphone with a directional response that can be coupled to each other using superposition. Specifically, for the exact same subarray, the response of one subarray can be calculated as described above, and then the response associated with the array arrangement of the virtual microphones (each corresponding to one position in the subarray) can be calculated by considering each virtual microphone as an omnidirectional microphone. Finally, the response of the macroarray can be considered as the complex product of the two responses calculated in the previous step.
[0074] Each null in the subarray corresponds to the zero response / sensitivity of the incoming audio signal at a given angle, and by multiplying these responses together in this way, these nulls are preserved so that the nulls in the macroarray include the nulls in each subarray and any nulls associated with the array arrangement of the virtual microphones. This is a powerful array design technique because it is much easier to design the array to have a desired set of characteristics, as the two superimposed responses can be designed independently and aimed.
[0075] Optionally, it may be advantageous to use a delayed-sum algorithm to determine the response of the virtual microphone array arrangement, while using a double delayed-subtract algorithm to determine the response of the constituent subarrays. In this way, the consistent beamwidth at low frequencies of the delayed-subtract technique can be superimposed with the narrow beamwidth at high frequencies of the delayed-sum technique to provide a narrower main beam that can be steered in the plane of the macroarray. When combining subarrays in this way, normalization using a frequency equalization filter can be applied only once to the determined response of the macroarray, rather than filtering each subarray individually before combining.
[0076] The effect of this further improvement is shown in Figure 11, where the subarrays are coupled in this way so that the side lobes of each subarray (or the virtual microphone corresponding to the subarray) are hidden in the nulls (or other low-response regions) of the macroarray, and vice versa. This coupling also means that a single subarray does not need to prioritize responses with a particularly narrow main beam (and instead, the focus can be on achieving more consistent signal rejection away from the main beam by bundling the nulls more closely together to remove or eliminate side lobes, etc.), because the coupling of subarrays further narrows the main beam of the macroarray. Next, rejection away from the main beam can be achieved by directing a nearly continuous spread of nulls, which can be called a wall of nulls, to couple to achieve a high level of voice rejection over a wide range of approach angles.
[0077] Furthermore, a delayed sum assembly of subarrays with relatively wide main beam widths allows for adaptive adjustment of the steering of the narrower macroarray main beam within the wider main beam of the subarray, depending on the delay value used in the delayed sum algorithm.
[0078] Furthermore, increasing the number of sub-arrays within a macroarray, rather than increasing the number of microphones within a sub-array, reduces white noise gain and the overall noise level of the system.
[0079] By using subarrays with four microphones arranged in a parallelogram, the formation of macroarrays can utilize tessellation to join adjacent subarrays, which offers the further advantage of allowing the positions of microphones in one subarray to overlap with the positions of microphones in adjacent subarrays, so that the output of a single microphone can be used as input to the beamforming algorithm of both subarrays. In fact, with parallelogram tessellation, each microphone can be positioned to contribute to the inputs of up to four subarrays. This obviously results in a significant reduction in the total number of physical microphones required to implement the macroarray. As an example, tessellation of 25 subarrays using a parallelogram shape, which would otherwise require 100 physical microphones, can be implemented with only 36 physical microphones.
[0080] Such tessellation is made possible by using larger subarrays than those found in the prior art, based on considerations of the wavelength limit of the prior art. Subarray shapes other than parallelograms can also be tessellated in the same way, and this disclosure is not limited thereto.
[0081] As described above, this disclosure provides a system and method for distinguishing a single speaker or sound source from ambient noise for use in a wide range of applications, such as detecting the location of a personal hearing aid or an object known to be associated with a specific voice signature, such as an animal, vehicle, drone or other aircraft, or for detecting the direction of signs of life in an emergency rescue situation. To determine the direction associated with the sound source of a voice signature, the main beam of the macroarray can be swept to pinpoint the direction in which the strongest voice signal is detected by iteratively adjusting various time delays associated with the beamforming of the macroarray.
[0082] However, the present invention is not limited to this embodiment, and these same principles for processing the output of a microphone array can also be used to process the input to a speaker array to provide a highly directional audio output. Those skilled in the art will see that the operation of speakers is simply the inverse of that of microphones, and that the above teachings can be simply reversed to apply them to embodiments of speaker arrays.
[0083] Figure 12 shows an exemplary system 40 for this embodiment of the present disclosure that provides directional processing of audio information. The system 40 includes a device 42 and a speaker array 44. The device 42 includes an input 46, a digital signal processor 48, and one or more outputs 50. It will be seen that the power amplifiers necessary to drive the speakers can be placed at any point along the communication path between the speakers 44 and the digital signal processor 48. For example, the power amplifiers may be part of the device 42 or placed in close proximity to the speaker array 44.
[0084] In one embodiment, the speaker array may include at least two pairs of parallel speakers that are not collinear with each other, such that the first pair of speakers includes a first speaker and a second speaker, and the second pair of speakers includes a third speaker and a fourth speaker, with each of these speakers positioned on a plane. By configuring these speakers to have a common directivity / dispersion pattern, the associated processing required to determine the audio signal applied to each speaker can be simplified. Alternatively, differences in the directivity / dispersion patterns of a pair of speakers can be addressed by introducing weighting to the audio signal associated with one of the speakers at each stage involving inversion.
[0085] The input unit 46 of the device 40 can be configured to receive audio signals transmitted by the speaker array and transmit them to the digital signal processor 48 for processing. The digital signal processor 48 can be configured to determine which audio signals to apply to each speaker of the speaker array 44 by determining that the audio signals associated with a first pair of speakers are the audio signals transmitted by the speaker array, and by applying a delay and inversion algorithm to the audio signals transmitted by the speaker array using a first time delay in order to determine which audio signals are associated with a second pair of speakers.
[0086] The delay and inversion algorithm adds a delay to the audio signal and then inverts it by multiplying it by -1. This corresponds to the delay-subtraction algorithm in the microphone embodiment, but in the speaker embodiment, the audio signals are split rather than combined, so the concepts of addition and subtraction do not apply. Therefore, inversion in the speaker embodiment is equivalent to subtraction in the microphone embodiment.
[0087] Next, to determine that the audio signal to be applied to the first speaker is the audio signal associated with the first pair of speakers, and to determine the audio signal to be applied to the second speaker, a delay and inversion algorithm can be applied to the audio signal associated with the first pair of speakers using a second time delay. Similarly, to determine that the audio signal to be applied to the third speaker is the audio signal associated with the second pair of speakers, and to determine the audio signal to be applied to the fourth speaker, a delay and inversion algorithm can be applied to the audio signal associated with the second pair of speakers using a third delay. By controlling the first, second, and third time delays, the device 40 can control the directivity of the audio signals transmitted by the speaker array 44.
[0088] Similar to microphone arrangements, one embodiment of this aspect of the present disclosure can provide a speaker array comprising four speakers arranged at the vertices of a parallelogram, such that the distance between a first pair of speakers is equal to the distance between a second pair of speakers, and thus the first time delay and the second time delay are set to be equal. In this case as well, the speaker array can be considered a subarray of a larger macroarray of speakers formed by tessellating the corresponding subarrays of speakers using the teachings described above with respect to microphone arrays.
Claims
1. A method for directional processing of speech information in a digital signal processor for a microphone array, Receiving an audio signal from each microphone in a microphone array, wherein the microphone array includes at least two pairs of parallel microphones that are not collinear with each other, the first pair of microphones includes a first microphone and a second microphone, the second pair of microphones includes a third microphone and a fourth microphone, and each microphone is arranged on a plane, and receiving an audio signal from each microphone. A first time delay is added to the audio signal from the second microphone, and the delayed audio signal from the second microphone is subtracted from the audio signal from the first microphone in order to determine the audio signal associated with the first pair of microphones. A second time delay is added to the audio signal from the fourth microphone, and the delayed audio signal from the fourth microphone is subtracted from the audio signal from the third microphone in order to determine the audio signal associated with the second pair of microphones. This includes adding a third time delay to the audio signal associated with the second pair of microphones, and subtracting the delayed audio signal associated with the second pair of microphones from the audio signal associated with the first pair of microphones in order to determine the audio signal associated with the microphone array, A method wherein the first, second, and third time delays are configured to control the directivity of the audio signal associated with the microphone array.
2. The method according to claim 1, wherein the first time delay is set based on the relative distance between the first microphone and the second microphone, the second time delay is set based on the relative distance between the third microphone and the fourth microphone, and the third time delay is set based on the relative distance between the first pair of microphones and the second pair of microphones.
3. The method according to claim 1, wherein the distance between microphones in a pair of microphones is uniform across the entirety of at least two pairs of microphones, and the first time delay is equal to the second time delay.
4. The method according to claim 1, wherein the microphone array includes two pairs of microphones, and the first, second, third, and fourth microphones are arranged at the vertices of a parallelogram.
5. The determined audio signal associated with the microphone array corresponds to a given direction based on the first, second, and third time delay values, and the method The first, second, and third time delays are iteratively adjusted, and a plurality of corresponding audio signals associated with the iteratively adjusted first, second, and third time delays are determined. The method according to claim 1, further comprising comparing the plurality of corresponding audio signals to determine the direction that most closely corresponds to a desired audio signature or sound source.
6. The method according to any one of claims 1 to 5, further comprising determining one or more audio signals associated with one or more corresponding further microphone arrays arranged on the plane, and combining each audio signal associated with the microphone arrays to determine the audio signal associated with the array of microphone arrays using a beamforming algorithm.
7. The method according to claim 6, wherein the beamforming algorithm is a delayed-sum beamforming algorithm.
8. A device for directional processing of audio information for a microphone array, A plurality of input units configured to receive audio signals from each microphone in the microphone array, wherein the microphone array includes at least two pairs of parallel microphones that are not collinear with each other, the first pair of microphones includes a first microphone and a second microphone, the second pair of microphones includes a third microphone and a fourth microphone, and each microphone is arranged on a plane. A digital signal processor is configured to add a first time delay to the audio signal from the second microphone, subtract the delayed audio signal from the second microphone from the audio signal from the first microphone to determine the audio signal associated with the first pair of microphones, add a second time delay to the audio signal from the fourth microphone, subtract the delayed audio signal from the fourth microphone from the audio signal from the third microphone to determine the audio signal associated with the second pair of microphones, add a third time delay to the audio signal associated with the second pair of microphones, subtract the delayed audio signal associated with the second pair of microphones from the audio signal associated with the first pair of microphones to determine the audio signal associated with the microphone array, A device in which the first, second, and third time delays are configured to control the directivity of the audio signal associated with the microphone array.
9. The apparatus according to claim 8, wherein the first time delay is set based on the relative distance between the first microphone and the second microphone, the second time delay is set based on the relative distance between the third microphone and the fourth microphone, and the third time delay is set based on the relative distance between the first pair of microphones and the second pair of microphones.
10. The apparatus according to claim 8, wherein the distance between microphones in a pair of microphones is uniform across the entirety of at least two pairs of microphones, and the first time delay is equal to the second time delay.
11. The apparatus according to claim 8, wherein the microphone array includes two pairs of microphones, and the first, second, third, and fourth microphones are arranged at the vertices of a parallelogram.
12. The apparatus according to claim 8, wherein the digital signal processor is further configured to determine, based on the values of the first, second, and third time delays, that the audio signals associated with the microphone array correspond to a given direction, the digital signal processor is configured to iteratively adjust the first, second, and third time delays and to determine a plurality of corresponding audio signals associated with the iteratively adjusted first, second, and third time delays, and further to compare the plurality of corresponding audio signals to determine the direction that most closely corresponds to a desired audio signature or sound source.
13. The apparatus according to claim 8, wherein the digital signal processor is further configured to determine one or more audio signals associated with one or more corresponding further microphone arrays arranged on the plane, and to combine each audio signal associated with the microphone arrays to determine the audio signal associated with the array of microphone arrays using a beamforming algorithm.
14. The apparatus according to claim 13, wherein the beamforming algorithm is a delayed-sum beamforming algorithm.
15. The apparatus according to claim 13, wherein the plurality of microphone arrays are tessellated from one another such that at least one microphone is shared between two adjacent microphone arrays.
16. A system for directional processing of audio information for a microphone array, A microphone array comprising at least two pairs of parallel microphones that are not collinear with each other, wherein the first pair of microphones comprises a first microphone and a second microphone, the second pair of microphones comprises a third microphone and a fourth microphone, and each microphone is arranged on a plane; A system comprising the apparatus described in any one of claims 8 to 15.
17. A method for directional processing of audio information in a digital signal processor for speaker arrays, Receiving audio signals transmitted by the speaker array, The process includes determining which audio signal to apply to each speaker in a speaker array, wherein the speaker array includes at least two pairs of parallel speakers that are not on the same line, the first pair of speakers includes a first speaker and a second speaker, the second pair of speakers includes a third speaker and a fourth speaker, each speaker is arranged on a plane, and the process includes determining which audio signal to apply to each speaker. Determining that the audio signal associated with the first pair of speakers is the audio signal transmitted by the speaker array, Applying a delay and inversion algorithm to the audio signal transmitted by the speaker array, using a first time delay, in order to determine the audio signal associated with the second pair of speakers, Determining that the audio signal applied to the first speaker is the audio signal associated with the first pair of speakers, In order to determine the audio signal applied to the second speaker, a delay and inversion algorithm is applied to the audio signal associated with the first pair of speakers using a second time delay, Determining that the audio signal applied to the third speaker is the audio signal associated with the second pair of speakers, This includes applying a delay and inversion algorithm to the audio signal associated with the second pair of speakers, using a third time delay, in order to determine the audio signal to be applied to the fourth speaker, A method wherein the first, second, and third time delays are configured to control the directivity of the audio signal transmitted by the speaker array.
18. A device for directional processing of audio information in a digital signal processor for speaker arrays, The speaker array includes at least two pairs of parallel speakers that are not on the same line as each other, the first pair of speakers includes a first speaker and a second speaker, the second pair of speakers includes a third speaker and a fourth speaker, and each speaker is arranged on a plane, and the input unit is configured to receive an audio signal transmitted by the speaker array, The digital signal processor includes, Determining that the audio signal associated with the first pair of speakers is the audio signal transmitted by the speaker array, Applying a delay and inversion algorithm to the audio signal transmitted by the speaker array, using a first time delay, in order to determine the audio signal associated with the second pair of speakers, Determining that the audio signal applied to the first speaker is the audio signal associated with the first pair of speakers, In order to determine the audio signal applied to the second speaker, a delay and inversion algorithm is applied to the audio signal associated with the first pair of speakers using a second time delay, Determining that the audio signal applied to the third speaker is the audio signal associated with the second pair of speakers, In order to determine the audio signal applied to the fourth speaker, a delay and inversion algorithm is applied to the audio signal associated with the second pair of speakers using a third time delay, It is configured to determine which audio signal to apply to each speaker in the speaker array, A device in which the first, second, and third time delays are configured to control the directivity of the audio signal transmitted by the speaker array.
19. The apparatus according to claim 18, wherein the second time delay is set based on the relative distance between the first speaker and the second speaker, the third time delay is set based on the relative distance between the third speaker and the fourth speaker, the first time delay is set based on the relative distance between the first pair of speakers and the second pair of speakers, the distance between speakers in a pair of speakers is uniform over the entirety of the at least two pairs of speakers, the second time delay and the third time delay are equal, and the first, second, third and fourth speakers are arranged at the vertices of a parallelogram.
20. The apparatus according to claim 18 or 19, wherein the digital signal processor is further configured to determine one or more audio signals associated with one or more corresponding further speaker arrays arranged on the plane, using a beamforming algorithm.