Hybrid Audio Beamforming System
The hybrid audio beamforming system addresses issues of wide beamwidth and suboptimal directionality by using a time-domain beamformer for upper frequencies and a frequency-domain beamformer for lower frequencies, resulting in improved beam directionality and reduced resource usage.
Patent Information
- Application Number
- JP2023545980
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-28
- Filing Date
- 2022-01-27
- Publication Date
- 2026-02-16
- Estimated Expiration
- 2042-01-27
AI Technical Summary
Conventional audio beamforming systems suffer from wider beamwidths and suboptimal directionality, especially at lower frequencies, leading to unwanted audio detection and increased computational and memory resource usage.
A hybrid audio beamforming system that combines a time-domain beamformer for upper frequency bands and a frequency-domain beamformer for lower frequency bands, using distinct beamforming techniques to generate narrower beams with improved directionality and reduce resource usage.
The hybrid system achieves narrower beams and enhanced directionality across different frequency ranges while minimizing computational and memory resources, improving overall performance and reducing unwanted audio detection.
Smart Images

Figure 0007814400000001 
Figure 0007814400000002 
Figure 0007814400000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 142,711, filed January 28, 2021, which is hereby incorporated by reference in its entirety.
[0002] This application relates generally to audio beamforming systems, and more particularly to a hybrid audio beamforming system having narrower beams and improved directivity through the use of a time-domain beamformer for processing upper frequency band signals of an audio signal and a frequency-domain beamformer for processing lower frequency band signals of the audio signal. [Background technology]
[0003] Conference environments, such as conference rooms, boardrooms, videoconferencing applications, and the like, may involve the use of microphones to capture sound from various audio sources in such environments. Such audio sources may include, for example, the person speaking. The captured sound may be disseminated to a nearby audience within the environment through amplified speakers (for sound reinforcement) and / or to others remote from the environment (such as through television broadcasts and / or webcasts). The type of microphones and their placement within a particular environment may depend on the location of the audio source, physical space requirements, aesthetics, room layout, and / or other considerations. For example, in some environments, microphones may be placed on a table or lectern near the audio source. In other environments, microphones may be mounted overhead, for example, to capture sound from throughout the room. Thus, microphones are available in a variety of sizes, form factors, mounting options, and wiring options to suit the needs of a particular environment.
[0004] Conventional microphones typically have a fixed polar pattern and several manually selectable settings. To capture sound in a conference environment, many conventional microphones can be used to immediately capture audio sources in the environment. However, conventional microphones also tend to capture undesirable audio, such as room noise, echoes, reverberation, and other harmful audio elements. Capturing these undesirable noises is exacerbated with the use of many microphones.
[0005] Array microphones with multiple microphone elements can offer benefits such as steerable coverage or pickup patterns with beams or lobes, which allow the microphone to focus on desired audio sources and reject undesirable sounds such as room noise. The ability to manipulate the audio pickup pattern offers the benefit of allowing for less precision in microphone placement, thus making the array microphone more forgiving. Additionally, array microphones offer the ability to pick up multiple audio sources using a single array microphone or unit, again due to the ability to manipulate the pickup pattern.
[0006] Beamforming is used to combine signals from the microphone elements of an array microphone to achieve a specific sound collection pattern with one or more beams or lobes. However, due to the longer wavelengths of sounds at lower frequencies, for wideband audio signals, conventional beamforming algorithms (e.g., time-domain Delay sum operation (delay and sum operating)The width of the beams generated using frequency-domain beamforming may be wider than configured or desired. Furthermore, the beam directionality may not be optimal when using conventional beamforming algorithms for wideband audio signals. Wider beamwidths and suboptimal beam directionality may result in unwanted audio detection, reduced performance of the array microphone, and dissatisfaction among users of the array microphone. Additionally, using frequency-domain beamforming across the entire frequency range may be computationally and memory resource intensive.
[0007] Thus, there is an opportunity for an audio beamforming system that addresses these concerns, and more specifically, for a hybrid audio beamforming system with narrower beams and improved directivity through the use of a time-domain beamformer to process the upper frequency band signals of the audio signal and a frequency-domain beamformer to process the lower frequency band signals of the audio signal. Summary of the Invention
[0008] The present invention is intended to solve the above problems by providing an audio beamformer system and method designed to, among other things: (1) provide a time-domain beamformer to generate a first beamformed signal based on upper frequency band signals obtained from an audio signal and using a time-domain beamforming technique; (2) provide a frequency-domain beamformer to generate a second beamformed signal based on lower frequency band signals obtained from the audio signal and using a first frequency-domain beamforming technique for a first group of the lower frequency band signals and a second frequency-domain beamforming technique for a second group of the lower frequency band signals; (3) output a beamformed output signal based on the first beamformed signal generated by the time-domain beamformer and the second beamformed signal generated by the frequency-domain beamformer; (4) have improved beam width and directionality, especially at lower frequencies; and (5) reduce the use of computational and memory resources by avoiding the use of frequency-domain beamforming across the entire frequency range.
[0009] In one embodiment, the beamforming system includes a first beamformer configured to generate a first beamformed signal based on a first frequency band signal obtained from a plurality of audio signals, a second beamformer configured to generate a second beamformed signal based on a second frequency band signal obtained from the plurality of audio signals, and an output generation unit in communication with the first and second beamformers, where the first beamformer is configured to process the first frequency band signal using a first beamforming technique, the second beamformer is configured to process the second frequency band signal using a second beamforming technique, and the output generation unit is configured to generate a beamformed output signal based on the first beamformed signal and the second beamformed signal.
[0010] In another embodiment, a beamforming system includes a first beamformer configured to generate a first beamformed signal based on upper frequency band signals obtained from a plurality of audio signals, a second beamformer configured to generate a second beamformed signal based on lower frequency band signals obtained from the plurality of audio signals, and an output generation unit in communication with the first and second beamformers. The first beamformer is configured to process the upper frequency band signals using a time-domain beamforming technique, and the second beamformer is configured to process a first group of lower frequency band signals using a first frequency-domain beamforming technique and a second group of lower frequency band signals using a second frequency-domain beamforming technique. The output generation unit is configured to generate a beamformed output signal based on the first beamformed signal and the second beamformed signal.
[0011] In another embodiment, a method includes receiving a plurality of audio signals; and generating a first beamformed signal based on an upper frequency band signal obtained from the plurality of audio signals using a time domain beamforming technique; frequency Using area beamforming techniques, a signal is obtained from multiple audio signals. Lower Based on the frequency band signal, 2 and generating a beamformed output signal based on the first beamformed signal and the second beamformed signal.
[0012] In another embodiment, a beamforming system includes a first beamformer configured to generate a first beamformed signal based on first frequency band signals obtained from a plurality of audio signals, a second beamformer configured to generate a second beamformed signal based on second frequency band signals obtained from the plurality of audio signals, and an output generation unit in communication with the first and second beamformers. The first beamformer is configured to process the first frequency band signals using a time-domain beamforming technique, and the second beamformer is configured to process a first group of second frequency band signals using a first frequency-domain beamforming technique and a second group of second frequency band signals using a second frequency-domain beamforming technique. The output generation unit is configured to generate a beamformed output signal based on the first beamformed signal and the second beamformed signal.
[0013] These and other embodiments, as well as various substitutions and aspects, will become apparent and will be more fully understood from the following detailed description and accompanying drawings that set forth illustrative embodiments that illustrate various ways in which the principles of the invention may be employed. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram of a hybrid audio beamforming system for use with an array microphone, according to some embodiments. [Figure 2] 2 is a flowchart illustrating operations for beamforming multiple microphone audio signals using the hybrid audio beamforming system of FIG. 1 in accordance with some embodiments. [Figure 3] 1 is a flowchart illustrating operations for beamforming upper frequency band signals derived from audio signals of multiple microphones and using a time domain beamformer, according to some embodiments. [Figure 4]1 is a flowchart illustrating operations for beamforming sub-band signals derived from audio signals of multiple microphones and using a frequency domain beamformer, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0015] The following description describes, shows, and illustrates one or more specific embodiments of the present invention in accordance with its principles. This description is provided not to limit the present invention to the embodiments described herein, but to explain and teach the principles of the present invention so that those skilled in the art can understand these principles and, with that understanding, apply them not only to the embodiments described herein but also to other embodiments that may be conceived in accordance with these principles. The scope of the present invention encompasses all such embodiments that may fall within the scope of the appended claims, either literally or under the doctrine of equivalents.
[0016] It should be noted that in the description and drawings, similar or nearly similar elements may be labeled with the same reference number. However, sometimes these elements may be labeled with different numbers, for example, where such labeling facilitates a clearer description. Furthermore, the drawings described herein are not necessarily drawn to scale, and in some instances, proportions may be exaggerated to more clearly show certain features. Such labeling and drawing practices do not necessarily suggest an underlying substantial objective. As stated above, this specification is intended to be taken as a whole and interpreted in accordance with the principles of the present invention as taught herein and as understood by those skilled in the art.
[0017] The hybrid audio beamforming system and method described herein can enable an array microphone to have narrower beams, improved beam directionality, and better overall performance across different frequency ranges. The hybrid audio beamforming system may include a time-domain beamformer configured to process an upper frequency band signal using a time-domain beamforming technique and a frequency-domain beamformer configured to process a group of lower frequency band signals using multiple frequency-domain beamforming techniques. The upper frequency band signal and the lower frequency band signal may be obtained from an audio signal, such as an audio signal from a microphone element of an array microphone. The hybrid audio beamforming system may generate a beamformed output signal based on a first beamformed signal from the time-domain beamformer and a second beamformed signal from the frequency-domain beamformer.
[0018] The frequency-domain beamformer may convert the time-domain audio signal to the frequency domain using a transform, such as a discrete Fourier transform (DFT) with a hop size smaller than the DFT block size. The frequency-domain beamformer may utilize a first frequency-domain beamforming technique to process a first group of sub-frequency band signals, such as the lower frequency components of the sub-frequency band signals. The frequency-domain beamformer may also utilize a second frequency-domain beamforming technique to process a second group of sub-frequency band signals, such as the higher frequency components of the sub-frequency band signals. By using multiple frequency-domain beamforming techniques in the frequency-domain beamformer, the frequency-domain beamformer may generate narrower beams with improved directionality for audio in the lower frequency range. The beamformed signal from the frequency-domain beamformer can be converted to the time domain, such as an inverse DFT, and the converted time-domain signal may be further smoothed using a weighted overlap-add (WOLA) method.
[0019] Therefore, combining a time-domain beamformer using a time-domain beamforming technique with a frequency-domain beamformer using a frequency-domain beamforming technique can result in more optimal beamwidths and directionality across different frequency ranges while using the same set of microphone elements in an array microphone. In addition, the increased computational and memory resources required when using frequency-domain beamforming across the entire frequency range can be avoided. The latency, computational resources, and storage capacity of weighting coefficients for the beamformer can therefore be minimized through the use of the hybrid audio beamforming system and method described herein.
[0020] 1 is a block diagram of a hybrid audio beamforming system 100. The hybrid audio beamforming system 100 may include microphone elements 102a, b, c,..., z included in an array microphone; a lower frequency band signal path 103 including a low-pass filter 104, a decimator 106, a frequency-domain beamformer 108, an interpolator 110, and a low-pass filter 112; an upper frequency band signal path 113 including a high-pass filter 114, a time-domain beamformer 116, and a delay element 118; a weight determination unit 120; and an output generation unit 122. The various components included in the hybrid audio beamforming system 100 may be implemented using software executable by a computing device having a processor and memory and / or by hardware (e.g., discrete logic circuits, application-specific integrated circuits (ASICs), programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.).
[0021] An array microphone including microphone elements 102a,b,c,...,z can detect sounds from audio sources at various frequencies. Array microphones can be utilized, for example, in conference rooms or boardrooms, where the audio sources can be one or more speakers and / or other desired sounds. There can be other sounds in the environment that can be harmful, such as noise from ventilation systems, other people, audiovisual equipment, electronic devices, etc. In a typical situation, the audio sources might be seated in chairs at a table, although other configurations and arrangements of audio sources are contemplated and possible.
[0022] The array microphone may be placed on a table, lectern, tabletop, etc., so that sounds from an audio source, such as voices spoken by a speaker, can be detected and captured. The array microphone may include any number of microphone elements 102 a, b, c,..., z, and the hybrid beamforming audio system 100 may be used to form multiple collection patterns so that sounds from the audio source are more consistently detected and captured. The microphone elements 102 a, b, c,..., z may be arranged in any suitable layout, including concentric rings, and / or harmonically nested. In embodiments, the microphone elements 102 a, b, c,..., z may be arranged generally symmetrically or asymmetrically. In other embodiments, the microphone elements 102 a, b, c,..., z may be, for example, disposed on a substrate, placed within a frame, or individually suspended. One embodiment of an array microphone is described in commonly assigned US Pat. No. 9,565,493, which is incorporated herein by reference in its entirety.
[0023] In some embodiments, the microphone elements 102a,b,c,...,z can each be a MEMS (microelectromechanical system) microphone. In other embodiments, the microphone elements 102a,b,c,...,z can be electret condenser microphones, dynamic microphones, ribbon microphones, piezoelectric microphones, and / or other types of microphones. In embodiments, the microphone elements 102a,b,c,...,z can be unidirectional microphones that are sensitive primarily in one direction. In other embodiments, the microphone elements 102a,b,c,...,z can have other directional or polar patterns, such as cardioid, subcardioid, or omnidirectional.
[0024] Each of the microphone elements 102 a, b, c,..., z in the array microphone may detect sound and convert the sound into an audio signal. Components within the array microphone, such as an analog-to-digital converter, a processor, and / or other components, may process the audio signal and ultimately generate one or more digital audio output signals. The digital audio output signals may, in some embodiments, conform to the Dante standard for audio transmission over Ethernet or may conform to other standards. In other embodiments, the microphone elements 102 a, b, c,..., z in the array microphone may output analog audio signals so that other components and devices (e.g., processors, mixers, recorders, amplifiers, etc.) external to the array microphone 100 may process the analog audio signals.
[0025] If the microphone elements 102a, b, c,..., z are simply used with a conventional beamformer (e.g., operating in the time domain), Delay sumIn some cases, especially at lower frequencies, the beamwidth may be wider than desired and the beam directionality may not be optimal. This may be due to the long wavelengths at these lower frequencies. Furthermore, beamforming at lower frequencies in the time domain may result in excessive sidelobes, relatively high latency, and / or higher computational load during processing.
[0026] However, as described in more detail herein, the lower frequency band signal path 103 (including the frequency-domain beamformer 108) and the upper frequency band signal path 113 (including the time-domain beamformer 116) may be in communication with the microphone elements 102 a, b, c,..., z. Specifically, the frequency-domain beamformer 108 may be used to process lower frequency band signals derived from the audio signals of the microphone elements 102 a, b, c,..., z. The lower frequency band signals may be, for example, 0 to 12 kHz. The time-domain beamformer 116 may be used to process upper frequency band signals derived from the audio signals of the microphone elements 102 a, b, c,..., z. The upper frequency band signals may be, for example, 12 to 24 kHz. Thus, using the hybrid audio beamforming system 100 may result in beamwidths that are narrower and have improved directionality across different frequencies, including lower frequencies.
[0027] One embodiment of a process 200 for hybrid beamforming of audio signals with an array microphone is shown in FIG. 2. Process 200 can be utilized to output a beamformed output signal from an array microphone using the hybrid audio beamforming system 100 shown in FIG. 1, where the beamformed output signal has a narrower beam and improved directionality. One or more processors and / or other processing components (e.g., analog-to-digital converters, encryption chips, etc.) within or external to system 100 may perform any, some, or all of the steps of process 200. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) may also be utilized in conjunction with the processor and / or other processing components to perform any, some, or all of the steps of process 200.
[0028] In step 202, the weight determination unit 120 may determine weighting coefficients for the frequency-domain beamformer 108 (which processes lower frequency band signals) and the time-domain beamformer 116 (which processes upper frequency band signals) based on the desired position and width of the beam. In some embodiments, the desired position and width of the beam may be determined programmatically or algorithmically using an automated decision-making scheme, such as automatic focusing, positioning, and / or unfolding of the beam. Embodiments of such schemes are described in commonly assigned U.S. patent applications Ser. Nos. 16 / 826,115 and 16 / 887,790, which are incorporated herein by reference in their entireties. In other embodiments, the desired position and width of the beam may be configured by a user, for example, through a user interface on an electronic device in communication with the weight determination unit 120.
[0029] The desired position of the beam may be determined or configured as a particular three-dimensional coordinate relative to the position of the array microphone, for example, in Cartesian coordinates (i.e., x, y, z), or spherical coordinates (i.e., radial distance r, polar angle θ (theta), azimuthal angle φ (phi)), etc. The desired width of the beam may be determined or configured, for example, in steps (e.g., narrow, medium, wide, etc.) or as an angle of field (e.g., angle, change in angle, percentage change, etc.).
[0030] In some embodiments, some or all of the weighting coefficients for the various positions and widths of the beam may be predetermined and stored in memory within weight determination unit 120 or in communication with weight determination unit 120. In other embodiments, some or all of the weighting coefficients for the various positions and widths of the beam may be calculated on the fly to reduce the amount of memory required for storing the weighting coefficients, e.g., to operate in the frequency domain in a relatively efficient and low latency manner. Delay sum For beamforming techniques, it may be possible to calculate such weighting factors on the fly. The calculation may take advantage of constant gains and uniform incremental phase shifts for all microphone elements 102a,b,c,...,z.
[0031] In an embodiment, for a particular beamforming technique (e.g., minimum variance distortionless response operating in the frequency domain), the weighting factors for various positions and widths of the beam may be generated using static noise covariance to obtain narrower beamwidths, or using dynamic noise covariance for improved signal-to-noise ratio.
[0032] Audio signals from microphone elements 102a, b, c, z may be received in step 204 at the lower frequency band signal path 103 (in an embodiment, at the low-pass filter 104) and also at the upper frequency band signal path 113 (in an embodiment, at the high-pass filter 114). In step 206, a first beamformed signal may be generated using the time-domain beamformer 116 and through the use of a time-domain beamforming technique based on the upper frequency band signal obtained from the audio signals from microphone elements 102a, b, c, z received in step 204. The upper frequency band signal may include mid- and higher frequencies, for example, 12-24 kHz. The time-domain beamforming technique used in the time-domain beamformer 116 may utilize the weighting coefficients determined in step 202. One embodiment of step 206 is described below with respect to FIG. 3.
[0033] In step 208, a second beamformed signal may be generated using the frequency-domain beamformer 108 based on lower frequency band signals obtained from the audio signals from the microphone elements 102a,b,c,...,z received in step 204, and through the use of frequency-domain beamforming techniques for different groups of the lower frequency band signals. The audio signals may be converted from the time domain to the frequency domain to produce lower frequency domain signals that are utilized in the frequency-domain beamformer 108. The lower frequency band signals may include signals having lower frequencies, e.g., 0-12 kHz, compared to the upper frequency band signals. The frequency-domain beamforming technique used in the frequency-domain beamformer 108 may utilize the weighting coefficients determined in step 202. One embodiment of step 208 is described below with respect to FIG. 4. In an embodiment, steps 206 and 208 may occur substantially simultaneously or at different times.
[0034] The beamformed output signal may be generated by the output generation unit 122 in step 210. The beamformed output signal may be generated by combining the first beamformed signal and the second beamformed signal generated by the time-domain beamformer 116 and the frequency-domain beamformer 108, respectively. In an embodiment, the first beamformed signal and the second beamformed signal may be combined by the output generation unit 122 by being added together to generate the beamformed output signal. The beamformed output signal may be a digital signal, such as, for example, a signal compliant with the Dante standard for audio transmission over Ethernet. In an embodiment, the beamformed output signal may be output to a component or device (e.g., a processor, mixer, recorder, amplifier, etc.) external to the hybrid audio beamforming system 100 and / or the array microphone.
[0035] FIG. 3 illustrates one embodiment of a process 206 for time-domain beamforming of an upper frequency band signal using an upper frequency band signal path 113 that includes a time-domain beamformer 108. Process 206 illustrated in FIG. 3 may correspond to step 206 of process 200 illustrated in FIG. 2. In process 206 of FIG. 3, an audio signal received in step 204 of process 200 may be filtered in step 302 by a high-pass filter 114. The high-pass filter 114 may be configured to pass audio signals having frequencies in an upper frequency range, e.g., 12-24 kHz. In an embodiment, the spectral response of the high-pass filter 114 may be matched to the spectral response of the low-pass filter 104 (in the lower frequency band signal path 103) to flatten the spectral response of a wideband signal, i.e., the beamformed output signal.
[0036] In step 304, the upper frequency band signals from the high pass filter 114 may be processed by the time domain beamformer 116 using a time domain beamforming technique. In an embodiment, the time domain beamformer 116 comprises: Delay sumAs previously mentioned, the weighting coefficients used by the time-domain beamformer 116 may be received from the weight determination unit 120 in step 202 based on the desired position and width of the beam.
[0037] In step 306, the signal generated by the time-domain beamformer 116 may be delayed by the delay element 118 to generate a first beamformed signal, which is provided to the output generation unit 122. As previously mentioned, the output generation unit 122 may combine the first and second beamformed signals in step 210 of process 200. The delay element 118 may add an appropriate amount of delay to the signal from the time-domain beamformer 116 to align the signal with the second beamformed signal generated by the lower frequency band signal path 103. This is because the lower frequency band signal path 103 has greater latency due to its additional components (i.e., the low-pass filters 104, 112, the decimator 106, and the interpolator 110) and due to the frequency-domain beamformer 108. Therefore, the amount of delay added by the delay element 118 may be based on the difference in latency between the lower frequency band signal path 103 and the upper frequency band signal path 113.
[0038] 4 shows one embodiment of a process 208 for frequency-domain beamforming of a lower frequency band signal using a lower frequency band signal path 103 that includes a frequency-domain beamformer 108. Process 208 shown in FIG. 4 may correspond to step 208 of process 200 shown in FIG. 2. In process 208 of FIG. 4, an audio signal received in step 204 of process 200 may be filtered by a low-pass filter 104 in step 402. Low-pass filter 104 may be configured to pass audio signals having frequencies in the lower frequency range, for example, 0-12 kHz.
[0039] The filtered signal from the low-pass filter 104 may be processed by the decimator 106 at step 404 to generate a lower frequency band signal for processing by the frequency-domain beamformer 108. Specifically, the decimator 106 may downsample the filtered signal by a particular factor to a lower sampling rate compared to the sampling rate of the audio signal received at step 204. The filtered signal may be downsampled to simplify the computational complexity and processing by the frequency-domain beamformer 108. In an embodiment, the decimator 106 may downsample the filtered signal from the 48 kHz sampling rate of the audio signal to half the 24 kHz sampling rate. In other embodiments, the decimator 106 may downsample the filtered signal by a different factor to another appropriate sampling rate.
[0040] In step 405, the decimated filtered signal may be converted from the time domain to the frequency domain using an appropriate frequency transform, such as a fast Fourier transform, a short-time Fourier transform, a discrete Fourier transform, a discrete cosine transform, or a wavelet transform. The lower frequency band signals may be processed using frequency domain beamforming techniques to avoid problems with excessive sidelobes and the need to use high-order filter banks that may arise when using time domain beamforming techniques on lower frequency band signals.
[0041] In steps 406 and 408, the frequency-domain beamformer 108 may process the two groups of sub-band signals using different frequency-domain beamforming techniques. Although Figure 4 shows the sub-band signals being processed in two groups, in embodiments it is contemplated and possible for the frequency-domain beamformer 108 to process more than two groups of sub-band signals using two or more frequency-domain beamforming techniques.
[0042] In an embodiment, the sub-band signal in the frequency domain may be transformed using a weighted overlap-add (WOLA) methodology. The WOLA methodology may decompose the sub-band signal into overlapping frames having a specific size to reduce artifacts at the boundaries between frames. The frames may be transformed into frequency bins using a frequency transform. The frequency bins may be divided into a first group (e.g., lower frequency components of the sub-band signal) and a second group (e.g., upper frequency components of the sub-band signal).
[0043] In an embodiment, the frame size of the WOLA methodology may be configurable to enable a trade-off between (1) latency in the lower frequency band signal path 103 and (2) computational resource and memory usage. Specifically, if the frame size is equal to or less than the block size of the frequency transform, the latency of the lower frequency band signal path 103 may be reduced while utilizing relatively high computational resources and memory. The block size of the FFT transform and the frame size may be expressed in terms of the number of samples. For example, the latency of the lower frequency band signal path 103 when the block size of the FFT transform is 256 and the frame size is 256 may be greater than the latency of the lower frequency band signal path 103 when the frame size is 128 or 192 (and the block size of the FFT transform remains 256) using a zero-padding method to fill the entire block of data for the FFT.
[0044] In step 406, the first group of sub-frequency band signals may be processed by the frequency-domain beamformer 108 using a first frequency-domain beamforming technique. In an embodiment, the first group may be lower-frequency components of the sub-frequency band signals, and the first frequency-domain beamforming technique may be a superdirective beamforming technique, such as a minimum variance distortion-free response (MVDR) beamforming technique. In other embodiments, the first frequency-domain beamforming technique may be another suitable superdirective beamforming technique. The frequency range of the lower-frequency components of the sub-frequency band signals may depend on the physical aperture size of the microphone array with which the beamformer is used, such as a frequency corresponding to less than the aperture size. For example, in an embodiment, the lower-frequency components of the sub-frequency band signals may be in the range of approximately 0 to 1 kHz or approximately 0 to 2 kHz. As previously mentioned, the weighting coefficients used by the first frequency-domain beamforming technique in the frequency-domain beamformer 116 may be received from the weight determination unit 120 in step 202 based on the desired position and width of the beam.
[0045] In step 408, the second group of lower frequency band signals may be processed by the frequency domain beamformer 108 using a second frequency domain beamforming technique. In an embodiment, the second group may be higher frequency components of the lower frequency band signals, and the second frequency domain beamforming technique may be Delay sumThe second frequency-domain beamforming technique may be a beamforming technique. In other embodiments, the second frequency-domain beamforming technique may be any other suitable beamforming technique. The frequency range of the upper frequency components of the lower frequency band signal may also depend on the physical aperture size of the microphone array with which the beamformer is being used, such as a frequency corresponding to one to two octaves higher than the aperture size. For example, in embodiments, the lower frequency components of the lower frequency band signal may be in the range of approximately 1 kHz or 2 kHz or higher. As previously described, the weighting coefficients used by the second frequency-domain beamforming technique in the frequency-domain beamformer 116 may be received from the weight determination unit 120 in step 202 based on the desired position and width of the beam. In embodiments, steps 406 and 408 may occur substantially simultaneously or at different times.
[0046] In step 409, the signals generated by the frequency domain beamformer 108 (based on the first and second frequency beamforming techniques) may be transformed from the frequency domain to the time domain using an appropriate inverse frequency transform, such as an inverse fast Fourier transform, an inverse short-time Fourier transform, an inverse discrete Fourier transform, an inverse discrete cosine transform, or an inverse wavelet transform. In an embodiment, the transformation of the signals from the frequency domain to the time domain may use the WOLA methodology as previously described.
[0047] At step 410, the converted signal (based on the signal generated by the frequency-domain beamformer 108) may be processed by the interpolator 110. Specifically, the interpolator 110 may upsample the signal generated by the frequency-domain beamformer 108 to a higher sampling rate by a particular factor. In an embodiment, the interpolator 110 may upsample the signal to twice the 48 kHz sampling rate. In other embodiments, the interpolator 110 may upsample the signal by a different factor to another suitable sampling rate.
[0048] Low-pass filter 122 may filter the upsampled signal from interpolator 110 in step 412 and generate a second beamformed signal that is provided to output generation unit 122. Output generation unit 122 may combine the first and second beamformed signals in step 210 of process 200, as previously described. Low-pass filter 122 may be configured to pass components of the upsampled signal having frequencies in the lower frequency range, for example, between 0 and 12 kHz.
[0049] 2-4 illustrate that the audio signal may be divided into groups for processing: an upper frequency band signal, a lower frequency component of the lower frequency band signal, and an upper frequency component of the lower frequency band signal. It should be noted that it is contemplated that the audio signal may be divided into groups for processing based on any suitable frequency range. Furthermore, any of the groups may be divided into groups for processing based on any suitable frequency range. Furthermore, any of the groups may be divided into groups for processing based on any suitable frequency range, including superdirective beamforming techniques in the frequency domain, ... Delay sum Beamforming techniques and / or time domain Delay sum It can be processed by beamforming techniques.
[0050] Any process description or block in a diagram should be understood to represent a module, segment, or portion of code that contains one or more executable instructions for implementing a particular logical function or step in a process; alternative implementations are within the scope of the embodiments of the invention, in which functions may be performed in a different order than that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved, as will be understood by those skilled in the art.
[0051] This disclosure describes how to make and use various embodiments in accordance with the technology, rather than limiting their true, intended, and fair scope and spirit. The above description is not intended to be exhaustive or to be limited to the precise forms disclosed. Modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described to best illustrate the principles of the described technology and to illustrate its practical application and to enable the technology to be utilized in various embodiments and with various modifications as suited to the particular uses contemplated. All such modifications and variations are within the scope of the embodiments, as determined by the appended claims, as they may be amended during the pendency of this patent application, when interpreted in accordance with the breadth to which they are fairly, legally, and equitably entitled, and all equivalents thereof.
Claims
1. 1. A beamforming system comprising: a first beamformer configured to generate a first beamformed signal based on a first frequency band signal obtained from a plurality of audio signals, the first beamformer configured to process the first frequency band signal using a first beamforming technique including a delay-and-sum beamforming technique performed in the time domain; a second beamformer configured to generate a second beamformed signal based on a second frequency band signal obtained from the plurality of audio signals, the second beamformer configured to process the second frequency band signal using a second beamforming technique, the second frequency band signal comprising a first group and a second group, the second beamformer further configured to process the first group using a superdirective beamforming technique performed in the frequency domain and to process the second group using a delay-and-sum beamforming technique in the frequency domain; an output generation unit in communication with the first and second beamformers, the output generation unit configured to generate a beamformed output signal based on the first beamformed signal and the second beamformed signal; A beamforming system comprising:
2. The beamforming system of claim 1 , wherein the first beamforming technique comprises a time-domain beamforming technique and the second beamforming technique comprises a frequency-domain beamforming technique.
3. The second beamforming technique includes a first frequency-domain beamforming technique and a second frequency-domain beamforming technique; the second beamformer is further configured to process the first group using the first frequency-domain beamforming technique and to process the second group using the second frequency-domain beamforming technique. The beamforming system of claim 1 .
4. 4. The beamforming system of claim 3, wherein the first and second frequency domain beamforming techniques are based on a weighted overlap-add (WOLA) methodology with a frame size less than or equal to a block size of a frequency domain transform.
5. The beamforming system of claim 4 , wherein the frame size is configurable.
6. The beamforming system of claim 1 , wherein the superdirective beamforming technique comprises a minimum variance distortionless response (MVDR) beamforming technique performed in the frequency domain.
7. the first frequency band signal comprises an upper frequency band signal; the second frequency band signal comprises a lower frequency band signal; the first group of sub-band signals comprises lower frequency components of the sub-band signals; the second group of lower frequency band signals comprises upper frequency components of the lower frequency band signals; The beamforming system of claim 1 .
8. The beamforming system of claim 1 , wherein the first frequency band signal comprises an upper frequency band signal and the second frequency band signal comprises a lower frequency band signal.
9. 1. A method comprising: receiving a plurality of audio signals; generating a first beamformed signal based on first frequency band signals obtained from the plurality of audio signals using a first beamforming technique including a delay-and-sum beamforming technique performed in the time domain; generating second beamformed signals based on second frequency band signals obtained from the plurality of audio signals using a second beamforming technique, the second frequency band signals comprising a first group and a second group, and generating the second beamformed signals includes processing the first group using a superdirective beamforming technique performed in the frequency domain and processing the second group using a delay-and-sum beamforming technique in the frequency domain; generating a beamformed output signal based on the first beamformed signal and the second beamformed signal; A method comprising:
10. The method of claim 9 , wherein the first beamforming technique comprises a time-domain beamforming technique and the second beamforming technique comprises a frequency-domain beamforming technique.
11. The second beamforming technique includes a first frequency-domain beamforming technique and a second frequency-domain beamforming technique; generating the second beamformed signal includes processing the first group using the first frequency domain beamforming technique and processing the second group using the second frequency domain beamforming technique.
10. The method of claim 9.
12. The method of claim 11 , wherein the first and second frequency domain beamforming techniques are based on a weighted overlap-add (WOLA) methodology with a frame size less than or equal to a block size of a frequency domain transform.
13. The method of claim 12 , wherein the frame size is configurable.
14. The method of claim 9 , wherein the superdirective beamforming technique comprises a minimum variance distortionless response (MVDR) beamforming technique performed in the frequency domain.
15. the first frequency band signal comprises an upper frequency band signal; the second frequency band signal comprises a lower frequency band signal; the first group of sub-band signals comprises lower frequency components of the sub-band signals; the second group of lower frequency band signals comprises upper frequency components of the lower frequency band signals; 10. The method of claim 9.
16. 10. The method of claim 9, wherein the first frequency band signal comprises an upper frequency band signal and the second frequency band signal comprises a lower frequency band signal.
17. An array microphone, a plurality of microphone elements, each configured to generate one of a plurality of audio signals; a beamformer configured to generate a beamformed output signal based on the plurality of audio signals, the beamformer comprising a plurality of beamformers each configured to process a first frequency band signal and a second frequency band signal using a different beamforming technique, the first frequency band signal and the second frequency band signal being derived from a plurality of audio signals; a first beamformer of the plurality of beamformers configured to process the first frequency band signals using a delay-and-sum beamforming technique in the time domain; an array microphone, wherein a second beamformer of the plurality of beamformers is configured to process the first group of second frequency band signals using a superdirective beamforming technique performed in the frequency domain, and is configured to process the second group of second frequency band signals using a delay-and-sum beamforming technique in the frequency domain.
Citation Information
Patent Citations
Method and apparatus for canceling noise from sound input through microphone
US20090141907A1
Endfire linear array microphone
US20190387311A1
Target sound enhancement device and car navigation system
WO2012160602A1
Sound collection device
WO2020026727A1