Microphone array beam forming method, system and device

By synchronously calibrating, differentially processing, and optimizing the MVDR algorithm of the microphone array signal, combined with frequency response equalization and dynamic range compression, the problem of unsatisfactory effects caused by environmental complexity in microphone array beamforming is solved, achieving higher-precision speech enhancement and stability.

CN120602827APending Publication Date: 2025-09-05SHENZHEN SHIDU DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510755350.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing microphone array beamforming methods fail to effectively consider complex factors such as reverberation and noise in actual application environments, resulting in unsatisfactory beamforming effects and affecting speech enhancement effects.

Method used

By synchronously calibrating the original sound signal collected by the omnidirectional microphone, differential processing is performed to obtain the cardioid directional unit signal, and the MVDR algorithm is used to calculate the signal covariance matrix and steering vector. Combined with frequency response equalization and dynamic range compression, beamforming processing is realized, and real-time performance monitoring and parameter adjustment are carried out.

Benefits of technology

It improves the accuracy of beamforming, reduces environmental noise interference, ensures stable operation of the system in different acoustic environments, improves voice quality and system performance, and adapts to diverse application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602827A_ABST
    Figure CN120602827A_ABST
Patent Text Reader

Abstract

The invention relates to a microphone array beam forming method, system and device, and the method comprises the steps: obtaining an original sound signal collected by an omnidirectional microphone, and carrying out the synchronous calibration processing of the original sound signal, and obtaining a corresponding preprocessing signal; carrying out differential processing on the pre-processed signal to obtain a plurality of heart-shaped directivity unit signals, and carrying out phase correction and gain balance optimization to obtain optimized heart-shaped unit signals; performing signal covariance matrix and guide vector calculation on the optimized heart-shaped unit signal according to a preset MVDR algorithm to obtain a corresponding optimal weight; performing beam forming processing on the optimal weight and the optimized heart-shaped unit signal to obtain an initial beam signal; performing frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; and performing performance monitoring and parameter adjustment based on the enhanced output signal to obtain a final real-time beam signal. According to the invention, stable operation of the system in different acoustic environments can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microphones, and in particular to a microphone array beamforming method, system and device. Background Art

[0002] Microphone arrays, as a highly efficient sound acquisition technology, have been widely used in voice interaction, far-field sound pickup, and other fields. With the increasing popularity of smart devices and remote conferencing systems, accurately localizing and enhancing sound sources to ensure clear voice communication quality has become a key research focus. Existing beamforming methods often focus solely on single technical parameters, such as spatial filtering or directional optimization of the signal, while ignoring the complex influence of factors such as reverberation and noise in actual application environments. This one-sided approach can lead to suboptimal beamforming results, compromising the ultimate voice enhancement effect. Summary of the Invention

[0003] The main purpose of the present invention is to provide a microphone array beamforming method, system and device, which can ensure that the system can operate stably in different acoustic environments and reduce adverse effects such as reverberation.

[0004] To achieve the above object, the present invention provides a microphone array beamforming method, comprising: Acquire an original sound signal collected by an omnidirectional microphone, perform synchronous calibration processing on the original sound signal, and obtain a corresponding preprocessed signal; Performing differential processing on the preprocessed signal to obtain a plurality of cardioid directional unit signals, and performing phase correction and gain balance optimization to obtain an optimized cardioid unit signal; Calculating the signal covariance matrix and the steering vector of the optimized cardioid unit signal according to a preset MVDR algorithm to obtain the corresponding MVDR optimal weight; Performing beamforming processing on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal; performing frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; Performance monitoring and parameter adjustment are performed based on the enhanced output signal to obtain a final real-time beam signal.

[0005] Furthermore, the acquiring of the original sound signal collected by the omnidirectional microphone and performing synchronous calibration processing on the original sound signal to obtain a corresponding preprocessed signal include: Acquire the sound signal collected by the omnidirectional microphone to obtain the original sound signal; Performing analog-to-digital synchronous conversion on the original sound signal to obtain a synchronous signal group; performing a DC offset elimination process on the synchronization signal group to obtain a de-biased signal group; Performing delay compensation filtering on the de-biased signal group to obtain a delay compensation signal group; Resampling the delay compensation signal group according to a preset sampling frequency to obtain a resampled signal group; performing acoustic calibration processing on the resampled signal group to obtain a calibration signal group; Performing frame processing on the calibration signal group according to a preset frame length to obtain a frame signal; Windowing is performed on the framed signal to obtain the preprocessed signal.

[0006] Furthermore, performing differential processing on the preprocessed signal to obtain a plurality of cardioid directional unit signals, and performing phase correction and gain balance optimization to obtain an optimized cardioid unit signal, includes: Grouping the preprocessed signals into groups of two adjacent microphones to obtain multiple groups of microphone signal pairs; Performing amplitude difference calculation on each pair of microphone signals to obtain an initial differential signal; Performing Fourier transform on the initial differential signal to obtain a frequency domain differential signal; Calculating the time delay parameters between adjacent microphones based on the frequency domain differential signal to obtain a phase compensation coefficient; Performing a convolution operation on the frequency domain difference signal and the phase compensation coefficient to obtain a phase correction signal; Performing gain modulation on the phase correction signal according to a preset cardioid directivity gain function to obtain a gain modulated signal; Performing an inverse Fourier transform on the gain modulated signal to obtain a time-domain cardioid unit signal; Calculating the average energy distribution according to the time-domain cardioid unit signal to obtain an energy correction coefficient; The time-domain cardioid unit signal and the energy correction coefficient are weighted to obtain the optimized cardioid unit signal.

[0007] Furthermore, the signal covariance matrix and steering vector calculation of the optimized cardioid unit signal according to the preset MVDR algorithm to obtain the corresponding MVDR optimal weights include: Performing frame processing on the optimized cardioid unit signal to obtain multiple signal frames; Performing windowing processing on the optimized cardioid unit signal according to the signal frame and a preset time window function to obtain a windowed signal; Performing autocorrelation and cross-correlation operations on the windowed signal to obtain a correlation signal matrix; Regularizing the correlation signal matrix according to the correlation signal matrix and a preset regularization coefficient to obtain a signal covariance matrix; Performing delay calculation on the optimized cardioid unit signal according to the array geometric layout parameters and the target sound source distance parameters to obtain delay information; Constructing a steering vector for the optimized cardioid unit signal according to the delay information and preset reference point parameters to obtain an initial steering vector; Performing weight optimization calculation on the optimized cardioid unit signal according to the signal covariance matrix and the initial steering vector to obtain the MVDR optimal weight; The formula for weight optimization calculation includes: ; is the optimal weight of the MVDR, is the signal covariance matrix, is the initial steering vector, and H is the conjugate transpose, which is used to ensure that the calculation result is a real value.

[0008] Furthermore, the performing beamforming processing on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal includes: performing sub-band decomposition processing on the optimized cardioid unit signal to obtain multiple frequency band signals; performing phase alignment processing on the plurality of frequency band signals to obtain an alignment signal matrix; Performing frame processing on the aligned signal matrix to obtain multiple signal subframes; Performing windowing processing on the signal subframe to obtain a windowed signal sequence; Performing Fourier transform processing on the signal sequence to obtain a frequency domain signal vector; Performing spatial correlation analysis on the frequency domain signal vector to construct a signal space feature matrix; Performing a complex matrix multiplication operation on the signal space feature matrix and the MVDR optimal weight to obtain a beam output vector; Performing inverse Fourier transform processing on the beam output vector to obtain a windowed signal sequence; Performing sub-band synthesis processing on the de-windowed signal sequence to obtain a reconstructed signal; Performing time domain compensation processing on the reconstructed signal to obtain the initial beam signal; The calculation formula for performing the complex matrix multiplication operation includes: ; is the beam output vector, is the optimal weight of the MVDR, is the signal space feature matrix.

[0009] Furthermore, performing frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal includes: Performing time-frequency transformation on the initial beam signal to obtain a time-frequency analysis signal; Performing energy spectrum smoothing analysis on the time-frequency analysis signal to obtain a smoothed energy spectrum; Dividing the time-frequency analysis signal into sub-bands according to the smoothed energy spectrum to obtain a plurality of equalized sub-band components; performing amplitude processing on the multiple equalized sub-band components respectively to obtain amplitude sub-band components; Calculating sub-band coefficients according to the amplitude sub-band components to obtain corresponding frequency response compensation coefficients; Performing a product operation on the frequency response compensation coefficient and the amplitude sub-band component to obtain a compensation sub-band component; Performing dynamic range detection on the compensation sub-band component to obtain a dynamic range parameter; Performing gain adjustment on the compensation sub-band component according to the dynamic range parameter to obtain a compression sub-band component; Performing sub-band synthesis on the compressed sub-band components to obtain a synthesized time-frequency signal; Performing inverse time-frequency transform on the synthesized time-frequency signal to obtain the enhanced output signal.

[0010] Furthermore, the performing performance monitoring and parameter adjustment based on the enhanced output signal to obtain a final real-time beam signal includes: Performing short-term performance parameter calculation on the enhanced output signal to obtain performance monitoring parameters; Performing spectrum analysis on the enhanced output signal according to the performance monitoring parameters to obtain spectrum characteristic parameters; Performing threshold judgment and data statistics on the spectrum characteristic parameters to obtain performance evaluation indicators; Calculating beamforming weights according to the performance evaluation index to obtain optimized beam weights; The enhanced output signal is beam reconstructed according to the optimized beam weight to obtain the real-time beam signal.

[0011] The present invention further provides a microphone array beamforming method system, which is applied to any of the above-mentioned microphone array beamforming methods, comprising: An acquisition module is used to acquire an original sound signal collected by an omnidirectional microphone, perform synchronous calibration processing on the original sound signal, and obtain a corresponding preprocessed signal; An analysis module, configured to perform differential processing on the preprocessed signal to obtain a plurality of cardioid directional unit signals, and perform phase correction and gain balance optimization to obtain an optimized cardioid unit signal; An association module is configured to calculate a signal covariance matrix and a steering vector for the optimized cardioid unit signal according to a preset MVDR algorithm to obtain a corresponding MVDR optimal weight; a processing module, configured to perform beamforming processing on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal; a control module, configured to perform frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; An execution module is used to perform performance monitoring and parameter adjustment based on the enhanced output signal to obtain a final real-time beam signal.

[0012] The present invention also provides a microphone array beamforming method and device, comprising: Memory, used to store programs; The processor is configured to execute the program to implement each step of any one of the above-mentioned microphone array beamforming methods.

[0013] The present invention provides a microphone array beamforming method, system, and device, which have the following beneficial effects: By synchronously calibrating and preprocessing the original sound signal, combined with differential processing to generate a cardioid driver signal, the system more accurately captures the spatial information of the target sound source, thereby improving beamforming accuracy and providing a reliable foundation for subsequent signal enhancement. By performing phase correction and gain balance optimization on the cardioid driver signal, and calculating the covariance matrix based on the MVDR algorithm, the system achieves refined processing of signals from sound sources in different directions, helping to enhance the target sound source and reduce ambient noise interference. Beamforming based on the optimized cardioid driver signal ensures stable operation in various acoustic environments, reduces adverse effects such as reverberation, and improves overall system performance. Frequency response equalization and dynamic range compression of the initial beam signal enable the development of a more optimized signal enhancement strategy. Real-time monitoring and parameter adjustment enable adaptive optimization of the system, effectively improving voice quality. Furthermore, by considering the complexity of actual application environments, the system can flexibly adjust the beamforming strategy based on different scenario characteristics and changing requirements, making the system more adaptable to diverse application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flow chart of a microphone array beamforming method provided by the present invention; Figure 2 This is a structural diagram of a microphone array beamforming system provided by the present invention; Figure 3 This is a structural diagram of a microphone array beamforming device provided by the present invention.

[0015] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0018] Reference Figure 1 As shown, the present invention provides a microphone array beamforming method, comprising: Step S1: obtaining an original sound signal collected by an omnidirectional microphone, performing synchronous calibration processing on the original sound signal, and obtaining a corresponding preprocessed signal; Step S2: performing differential processing on the preprocessed signal to obtain multiple cardioid directional unit signals, and performing phase correction and gain balance optimization to obtain an optimized cardioid unit signal; Step S3: Calculate the signal covariance matrix and steering vector of the optimized cardioid unit signal according to the preset MVDR algorithm to obtain the corresponding MVDR optimal weight; Step S4: performing beamforming processing on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal; Step S5: performing frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; Step S6: Perform performance monitoring and parameter adjustment based on the enhanced output signal to obtain the final real-time beam signal.

[0019] Based on the above steps, the detailed process is as follows: Step S1: Collect sound signals in the environment using an omnidirectional microphone array. Omnidirectional microphones are characterized by their ability to pick up sound evenly from all directions, providing complete sound field information for subsequent beamforming processing. Because the multiple microphones in the array may differ in physical position and electrical characteristics, synchronous calibration is required to ensure temporal consistency of the signals. This calibration process includes delay compensation, amplitude calibration, and phase alignment. In specific implementation, the delay differences between the microphones must first be measured, and then the sampling points must be aligned through interpolation or resampling to ensure that the signals of all channels are completely synchronized in time. At the same time, the frequency response of each microphone must be calibrated using a test signal to compensate for response differences caused by manufacturing errors and obtain a calibrated preprocessed signal. The quality of this step directly affects the effectiveness of subsequent processing and is the foundation of the entire beamforming system.

[0020] Step S2: After obtaining the pre-processed signal, the omnidirectional signal is converted into a signal with cardioid directional characteristics through differential processing. This process is achieved by subtracting the adjacent microphone signals, and different directional patterns can be formed by adjusting the subtraction weights. The cardioid directional pattern has the characteristics of forward gain and backward suppression, which can preliminarily suppress noise and interference from non-desired directions. After obtaining the basic cardioid directional signal, phase correction and gain balance optimization are required. The purpose of phase correction is to compensate for the phase distortion introduced by differential processing and ensure that the phase relationship of each frequency band is correct. Gain balance optimization is to make each cardioid unit have a consistent response in the desired direction. This usually requires an adaptive algorithm to continuously adjust the gain coefficient of each unit until the optimal balance is achieved.

[0021] Step S3: Based on the optimized cardioid unit signal, the signal covariance matrix is ​​estimated using the preset MVDR algorithm. This matrix contains information about the signal's spatial correlation. A steering vector is calculated based on the location of the target sound source. This steering vector describes the phase relationship between the desired signal as it reaches the array from various directions. The covariance matrix and steering vector are then applied to the MVDR optimization criterion, and the optimal weight vector is obtained by solving the optimization problem. This weight vector ensures a gain of 1 (distortion-free response) when the beam is pointed in the target direction, while minimizing the impact of interference and noise from other directions. This step is the core of the entire beamforming system and directly determines the system's spatial selectivity and noise suppression capabilities.

[0022] Step S4: Perform complex multiplication of each cardioid unit signal with the corresponding optimal weight, and then superimpose all weighted signals to obtain the initial beam signal. This process is actually performed in the frequency domain. It is necessary to first perform a short-time Fourier transform (STFT) on the signal, complete weighting and superposition in the frequency domain, and then return to the time domain through inverse transformation. In actual processing, in order to avoid spectrum leakage and ensure the accuracy of time-frequency analysis, a suitable window function is usually used, and the overlap-addition method is adopted to ensure the continuity of the signal. The key to this step is to accurately implement the application of weights to ensure that the optimal weights calculated theoretically can be accurately executed in actual signal processing, thereby achieving the expected directional gain and noise suppression effects.

[0023] Step S5: Frequency response equalization and dynamic compression After obtaining the initial beam signal, frequency response equalization and dynamic range compression processing are required to improve the sound quality and usability of the signal. The purpose of frequency response equalization is to compensate for the frequency response distortion introduced by the beamforming process and ensure that the output signal has a flat response characteristic across the entire frequency band. This is usually achieved by designing a set of equalization filters, which can be achieved by parametric equalization or graphic equalization. Dynamic range compression is to process changes in signal intensity to avoid distortion caused by excessively strong signals or information loss caused by excessively weak signals. The compression process requires setting appropriate parameters such as thresholds, compression ratios, attack times, and release times to ensure that the dynamic characteristics of the signal are maintained while being limited to an appropriate range. The processing results of this step directly affect the listening quality of the final output signal, and the processing parameters need to be adjusted according to the specific application scenario.

[0024] Step S6: System performance is evaluated through real-time monitoring of the enhanced output signal, and necessary parameter adjustments are made. Performance monitoring encompasses multiple aspects: real-time calculation and tracking of metrics such as signal-to-noise ratio, directivity, frequency response flatness, and dynamic range. Based on these monitoring results, the system adaptively adjusts parameters across various processing steps, including synchronization calibration parameters, the cardioid driver gain coefficient, constraints in MVDR weight calculation, the equalizer frequency response, and compressor dynamic parameters. This step aims to enable the system to adapt to diverse acoustic environments and usage scenarios, maintaining optimal operating conditions. It also provides the system with real-time optimization capabilities, enabling rapid response to environmental changes and ensuring the stability of the beamforming effect.

[0025] The present invention provides a microphone array beamforming method that can more accurately capture the spatial information of the target sound source by synchronously calibrating and preprocessing the original sound signal and combining it with the cardioid directional unit signal obtained by differential processing, thereby improving the accuracy of beamforming and providing a reliable foundation for subsequent signal enhancement. By performing phase correction and gain balance optimization on the cardioid unit signal and calculating the covariance matrix based on the MVDR algorithm, refined processing of sound source signals in different directions is achieved, which helps to improve the enhancement effect of the target sound source and reduce environmental noise interference. Beamforming processing based on the optimized cardioid unit signal can ensure that the system can operate stably in different acoustic environments, reduce adverse effects such as reverberation, and improve the overall performance of the system. By performing frequency response equalization and dynamic range compression on the initial beam signal, a more reasonable signal enhancement strategy is formulated, and adaptive optimization of the system is achieved through real-time monitoring and parameter adjustment, thereby effectively improving voice quality. At the same time, by considering the complexity of the actual application environment, the beamforming strategy can be flexibly adjusted according to the characteristics of different scenarios and changes in demand, making the system more adaptable to diverse application scenarios.

[0026] In one embodiment, obtaining an original sound signal collected by an omnidirectional microphone and performing synchronous calibration processing on the original sound signal to obtain a corresponding preprocessed signal includes: The sound signals in the environment are collected through an omnidirectional microphone array. Multiple omnidirectional microphones are arranged in a preset geometric pattern, and each microphone collects the sound signals in the sound field in real time to form the original sound signal set.

[0027] The original sound signal is synchronously sampled and quantized by the analog-to-digital converter. The converter uses 24-bit quantization accuracy and a sampling rate of 48kHz. During the analog-to-digital conversion process, the same clock source is used to drive all channels, ensuring synchronization of the signals across all channels and outputting a synchronized digital signal group.

[0028] The signal's DC offset is removed using a high-pass filter. The filter's cutoff frequency is set to 20 Hz, removing the DC component and very low-frequency interference from the signal. The filter uses a Butterworth structure, offering a flat passband response and outputting a debiased signal.

[0029] Delay compensation filtering uses an FIR filter bank. Based on the geometric position of the microphone array, the relative delay between channels is calculated. The FIR filter has an order of 256, and the filter coefficients are designed using a Hamming window. This filtering process compensates for the delay caused by differences in sound wave propagation paths and outputs a delay-compensated signal bank.

[0030] The resampling process converts the signal sampling rate to a uniform 16kHz. A polyphase filter structure with a filter length of 512 points is used in the resampling process. This ensures the real-time requirements of subsequent processing and outputs a resampled signal group.

[0031] Acoustic calibration equalizes and compensates for the amplitude response of each channel's signal. An equalization filter is designed based on the microphone's sensitivity curve and frequency response characteristics. The equalization filter uses a 64th-order FIR structure to compensate for the uneven frequency response of the microphone and output a calibration signal set.

[0032] Framing divides the continuous signal into short time frames. Each frame is set to 512 samples, with a 50% overlap between adjacent frames. This frame length setting maintains frequency domain resolution while meeting the quasi-steady-state assumption for speech signals, and outputs the framed signal.

[0033] Windowing uses a Hanning window function. The window function's length matches the frame length, and weighted processing is performed on the framed signal. This process reduces spectral leakage and improves the accuracy of frequency domain analysis, ultimately outputting the preprocessed signal.

[0034] After the above preprocessing steps, the sound signal quality is significantly improved. The preprocessed signal's signal-to-noise ratio is increased, and its time-frequency characteristics are more stable, laying a good foundation for subsequent beamforming processing. This preprocessing solution has demonstrated good robustness and processing effectiveness in practical engineering applications.

[0035] This embodiment ensures the synchronization and quantization accuracy of multi-channel signal acquisition and effectively avoids signal distortion by adopting 24-bit high-precision analog-to-digital conversion and unified clock source drive. High-pass filters and FIR filter groups are used to eliminate DC offset and compensate for delay, eliminating low-frequency interference in the signal and achieving accurate compensation for sound wave propagation delay. Polyphase filters and equalization filters are used in the resampling and acoustic calibration links to maintain signal quality while reducing the sampling rate and compensate for the uneven frequency response of the microphone. The optimal parameter configuration is selected for framing and windowing processing, which meets the real-time processing requirements while ensuring the accuracy of frequency domain analysis, and significantly improves the time-frequency characteristics of the preprocessed signal. The entire preprocessing scheme has a reasonable structure and the processing links are closely connected, so that the output signal has a higher signal-to-noise ratio and more stable characteristics, providing a high-quality input data foundation for beamforming processing.

[0036] In one embodiment, differential processing is performed on the preprocessed signal to obtain multiple cardioid directional unit signals, and phase correction and gain balance optimization are performed to obtain optimized cardioid unit signals, including: The processing is based on the acoustic signals collected by multiple microphones. After obtaining the pre-processed signal, the signals of two adjacent microphones are grouped as a group. Assuming that the microphone array contains N microphones, N-1 groups of microphone signal pairs are obtained after grouping. For each group of microphone signal pairs, an amplitude difference calculation is performed to obtain the initial differential signal. The amplitude difference calculation adopts the direct subtraction method of adjacent microphone signals, including: subtracting the signal of the i-th microphone from the signal of the i+1-th microphone, where the value of i ranges from 1 to N-1. It also includes differentiating the two microphone signals S1 and S2 to obtain the directional signal D=S1-S2 to achieve a cardioid directivity response.

[0037] This differential calculation highlights subtle differences in the signals collected by adjacent microphones, effectively extracting spatial phase information during sound wave propagation. Each set of differential signals reflects the time-amplitude characteristics of sound waves propagating between adjacent microphones, providing basic data for subsequent phase correction and beamforming.

[0038] The differential signal is converted to the frequency domain via a Fourier transform, resulting in a frequency-domain differential signal. This frequency-domain differential signal contains both amplitude and phase spectrum information. The time delay parameter between adjacent microphones is calculated based on the phase spectrum characteristics of the frequency-domain differential signal. The time delay parameter is obtained by analyzing the slope of the phase spectrum and reflects the time difference between sound waves reaching adjacent microphones. The phase compensation coefficient is derived from the time delay parameter to compensate for phase deviation caused by the distance between the microphones.

[0039] Phase correction is achieved by convolving the frequency-domain differential signal with the phase compensation coefficient. Phase correction aims to align the phase responses of individual frequency components, synchronizing the signals in the time domain. Gain modulation is achieved by multiplying the phase-corrected signal with the preset cardioid gain function. The cardioid gain function defines the desired gain response in different directions.

[0040] The gain-modulated signal is converted back to the time domain via an inverse Fourier transform to obtain a time-domain cardioid driver signal. Based on the time-domain cardioid driver signal, its energy distribution characteristics, including statistical features such as short-time energy and RMS value, are calculated to derive an energy correction coefficient. This energy correction coefficient is used to balance the output energy of each cardioid driver.

[0041] The time-domain cardioid driver signal and the energy correction coefficient are weighted to produce an optimized cardioid driver signal. This weighting employs a linear weighting scheme to ensure that each driver contributes an appropriate weight to the beamforming process. This optimized cardioid driver signal exhibits excellent directional performance and acoustic properties, laying the foundation for subsequent beamforming.

[0042] Through the above signal processing flow, this method realizes the conversion from the original microphone signal to a beam signal with a cardioid directivity characteristic, providing reliable technical support for practical applications.

[0043] This embodiment effectively solves the problems of insufficient directionality and phase misalignment in traditional microphone array beamforming by performing differential processing and phase correction on the pre-processed signal. Based on the differential calculation of adjacent microphone signal pairs, combined with the convolution operation of the phase compensation coefficient, the system can accurately capture the direction information of the sound source, significantly improving the directional performance of the beamforming. The introduction of the cardioid directivity gain function realizes the precise modulation of the signal gain, ensuring that the beam has the best acoustic response characteristics in the desired direction. Through the weighted processing mechanism of the energy correction coefficient, this method effectively balances the output energy distribution of each cardioid unit, overcoming the beam distortion problem caused by energy imbalance in traditional methods. The overall solution adopts a processing strategy that combines frequency domain and time domain, which not only ensures the real-time performance of the algorithm, but also improves the anti-noise ability of the system, providing reliable technical support for acoustic signal acquisition in actual application environments.

[0044] In one embodiment, the signal covariance matrix and steering vector are calculated for the optimized cardioid unit signal according to a preset MVDR algorithm to obtain the corresponding MVDR optimal weight, including: The optimized cardioid unit signal is framed. Assuming a sampling frequency of 16kHz, a frame length of 256 samples, and a frame shift of 128 samples, multiple signal frames are generated, with a 50% overlap between frames to ensure signal continuity. Framing is used to divide a continuous signal into short, stable segments for easier processing.

[0045] The optimized cardioid unit signal is windowed based on the signal frame and a preset time window function. In this embodiment, a Hanning window is selected as the time window function, with a window length equal to the frame length, i.e., 256 points. The purpose of windowing is to reduce signal endpoint effects, lower spectral leakage, and improve the accuracy of signal spectrum analysis.

[0046] Perform autocorrelation and cross-correlation on the windowed signal to generate a correlation signal matrix. Specifically, for each signal frame, the cross-correlation and autocorrelation between the signals from each microphone unit are calculated to form a 4x4 correlation signal matrix. The autocorrelation operation is used to estimate the power spectrum of the signal, and the cross-correlation operation is used to estimate the correlation between the signals.

[0047] The correlation signal matrix is ​​regularized based on the correlation signal matrix and a preset regularization coefficient to obtain a signal covariance matrix. In this embodiment, the regularization coefficient is set to 0.01 to improve the condition number of the correlation signal matrix and prevent numerical instability in subsequent inverse matrix operations. Regularization is typically achieved by adding a small quantity to the diagonal of the correlation signal matrix. This small quantity is typically the regularization coefficient multiplied by the largest eigenvalue of the matrix.

[0048] The optimized cardioid unit signal is delayed based on the array geometry and the target sound source distance parameters to obtain delay information. In this embodiment, the array geometry is a linear array, the microphone spacing is 0.05 meters, and the target sound source is located 30 degrees in front and 2 meters away. Based on the speed of sound and the geometric relationship, the delay of each microphone unit receiving the target sound source signal can be calculated. For example, the delay of the first microphone is 0, and the delay of the second microphone is (d*sin(θ) / c), where d is the microphone spacing, θ is the target direction angle, and c is the speed of sound.

[0049] A steering vector is constructed for the optimized cardioid unit signal based on the delay information and preset reference point parameters to obtain an initial steering vector. In this embodiment, the reference point parameter is set to the center point of the array. Based on the delay information of each microphone unit, a complex steering vector is constructed to represent the array's directivity toward the target direction. Each element of the steering vector corresponds to a complex weight for a microphone unit, with the magnitude and phase of the weight determined by the delay information.

[0050] Based on the signal covariance matrix and the initial steering vector, the weights of the optimized cardioid unit signal are optimized to obtain the optimal MVDR (Minimum Variance Distortionless Response) weights. The MVDR algorithm solves an optimization problem to minimize the output noise power while maintaining no distortion of the target signal. Specifically, the MVDR weight vector can be calculated using the following formula: Among them, the calculation formula for obtaining the signal covariance matrix includes: ; The calculation formula for the initial steering vector includes: ; The formula for calculating the cardioid driver output is: ; The formulas for weight optimization calculation include: ; is the optimal weight of MVDR, is the signal covariance matrix, is the initial steering vector, H is the conjugate transpose, which is used to ensure that the calculation result is a real value. is the target direction angle, 、 All are optimized cardioid unit signals. is the signal space feature matrix, is a constant.

[0051] In one embodiment, beamforming is performed on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal, including: The optimized cardioid signal is first decomposed into multiple frequency bands by a filter bank. This decomposition ensures that each frequency band covers a specific frequency range and has concentrated signal energy, making it easier to process.

[0052] Phase alignment is performed on these frequency band signals. By calculating and adjusting the phase difference of the signals in each frequency band, the phase delay caused by the different microphone positions in the microphone array is eliminated, ensuring the temporal consistency of the signals.

[0053] The phase-aligned signal matrix is ​​framed and divided into multiple time subframes. The frame length should be selected to balance the time resolution and frequency resolution to capture the instantaneous characteristics of the signal.

[0054] Each subframe signal is then multiplied by a window function to perform windowing to reduce spectral leakage. Common window functions include Hanning window and Hamming window.

[0055] The windowed signal sequence is converted to the frequency domain through Fourier transform to obtain the frequency domain signal vector, which reveals the frequency components of the signal.

[0056] Perform spatial correlation analysis on the frequency domain signal vector and construct a signal space feature matrix to reflect the energy distribution of the signal in different directions.

[0057] The beam output vector is generated by performing a complex matrix multiplication on this matrix and the MVDR optimal weights, which are calculated using the minimum variance distortionless response criterion and aim to enhance the target signal and suppress interference.

[0058] The beam output vector is converted back to the time domain through inverse Fourier transform to obtain the windowed signal sequence.

[0059] The de-windowed signal sequence is subjected to sub-band synthesis to restore the complete frequency structure of the signal, and then time domain compensation processing is performed to eliminate the distortion introduced during the processing, and finally the initial beam signal is obtained.

[0060] The calculation formula for complex matrix multiplication includes: ; is the beam output vector, is the optimal weight of MVDR, is the signal space feature matrix.

[0061] This embodiment uses sub-band decomposition to process different frequency components separately, improving signal processing efficiency and focusing on the relevant frequency range. Phase alignment ensures temporal consistency of signals from different microphones, reducing distortion and improving beamforming. Using MVDR optimal weights effectively reduces noise variance, maintains the target signal, and outputs a clearer signal. The overall process significantly improves the signal-to-noise ratio and beam directionality, achieving excellent performance in complex noisy environments. Furthermore, a modular signal processing approach can improve computational efficiency and reduce power consumption.

[0062] In one embodiment, frequency response equalization and dynamic range compression are performed on the initial beam signal to obtain an enhanced output signal, including: The initial beam signal is subjected to time-frequency transformation, typically using the short-time Fourier transform (STFT) to convert it from the time domain to the frequency domain, yielding a time-frequency analysis signal. The STFT decomposes the signal into multiple time frames by selecting an appropriate window function and window shift. Fourier transforms are then performed within each frame, revealing the signal's distribution at different time points and frequencies.

[0063] Perform energy spectrum smoothing on the time-frequency analysis signal. This step calculates the square of the amplitude of the STFT result to obtain the energy spectrum. A low-pass filter or other smoothing filter is then applied to reduce noise interference and highlight significant spectral components. The goal of smoothing is to eliminate high-frequency noise from the spectrum and make the energy spectrum smoother.

[0064] The time-frequency analysis signal is sub-banded based on the smoothed energy spectrum, dividing it into multiple balanced sub-band components. Sub-band division can be uniform or adaptive based on the distribution of the energy spectrum. Each sub-band covers a specific frequency range, ensuring relatively balanced energy distribution across sub-bands.

[0065] Each equalized subband component is subjected to amplitude processing, such as normalization or scaling, to ensure that the amplitudes of the subbands are consistent. The purpose of this step is to eliminate the amplitude differences between subbands, making subsequent processing more efficient.

[0066] Frequency response compensation coefficients are calculated based on the amplitude sub-band components. These coefficients are used to correct for possible frequency response non-uniformity in the system and are usually determined by comparing the ideal frequency response with the actual frequency response.

[0067] The frequency response compensation coefficient is multiplied by the amplitude sub-band component to obtain the compensation sub-band component. The purpose of this step is to adjust the amplitude of each sub-band to make the overall frequency response flatter.

[0068] Perform dynamic range detection on the compensation sub-band components to obtain dynamic range parameters, such as maximum and minimum amplitudes. The purpose of dynamic range detection is to evaluate the signal amplitude variation range of each sub-band.

[0069] The gain of the compensation sub-band component is adjusted according to the dynamic range parameter to obtain the compressed sub-band component. The gain adjustment strategy can be to compress the larger dynamic range so that the signal amplitude changes tend to be stable and the audibility of the signal is improved.

[0070] Subband synthesis is performed on the compressed subband components, recombining the individual subband signals into a single time-frequency signal. This signal is then converted back to the time domain using an inverse time-frequency transform (e.g., inverse STFT) to produce the enhanced output signal. The goal of subband synthesis and inverse transform is to restore the signal's time domain characteristics, making it suitable for further processing or playback.

[0071] This embodiment significantly improves the clarity and intelligibility of audio signals by implementing frequency response equalization and dynamic range compression. By dividing the signal into subbands and applying customized compensation coefficients, it effectively corrects frequency response unevenness and improves the balance of the sound effects. Dynamic range compression enhances the audibility of soft sounds while preventing overload and distortion of loud signals, making the audio output more consistent and easier to understand. In addition, by focusing on the target sound source and reducing background noise, this method is particularly suitable for speech recognition and audio in complex acoustic environments. Its modular design optimizes computational efficiency, making it suitable for real-time processing applications.

[0072] In one embodiment, performance monitoring and parameter adjustment are performed based on the enhanced output signal to obtain a final real-time beam signal, including: Short-term performance parameters are calculated on the enhanced output signal to obtain performance monitoring parameters. These parameters may include signal-to-noise ratio (SNR), distortion, or other related indicators used to evaluate signal quality and performance. By calculating these short-term parameters, the signal performance status can be monitored in real time.

[0073] Based on these performance monitoring parameters, the enhanced output signal is subjected to spectral analysis to obtain spectral characteristic parameters. This process may involve techniques such as Fourier transform to analyze the signal's frequency components and extract characteristic parameters such as spectral flatness and spectral centroid, thereby further understanding the signal's spectral characteristics.

[0074] Thresholds are determined and statistics are collected for these spectral characteristic parameters to obtain performance evaluation indicators. This step may involve setting thresholds to determine whether the parameters are within an acceptable range and performing statistical analysis, such as averaging over time or across frequency bandwidth, to obtain more stable performance indicators.

[0075] Based on these performance metrics, beamforming weights are calculated to obtain optimized beam weights. This may involve using algorithms such as minimum variance distortionless response (MVDR) to adjust the weights of each microphone signal based on the performance metrics to optimize beamforming. These optimized beam weights are used to reconstruct the enhanced output signal into a beam, thereby obtaining a real-time beam signal.

[0076] This embodiment significantly improves the adaptability and signal quality of the beamforming system through performance monitoring and parameter adjustment. Real-time calculation of performance short-term parameters and spectral characteristic parameters can dynamically evaluate the signal status and ensure stability in different acoustic environments. Optimizing beam weights based on performance evaluation indicators makes beamforming more accurate, effectively enhancing the target signal and suppressing interference noise. The combination of spectrum analysis and threshold judgment further improves the robustness of the system and avoids performance degradation due to environmental changes. Ultimately, the real-time beam signal generated by beam reconstruction has higher clarity and directionality, and is suitable for speech enhancement and noise suppression in complex scenarios. The overall approach optimizes computational efficiency, reduces power consumption, and improves the flexibility and practicality of the system, providing an efficient solution for real-time audio processing.

[0077] Reference Figure 2 As shown, the present invention provides a microphone array beamforming method system, which is applied to any of the above microphone array beamforming methods, including: The acquisition module is used to obtain the original sound signal collected by the omnidirectional microphone, perform synchronous calibration processing on the original sound signal, and obtain the corresponding pre-processed signal; An analysis module is used to perform differential processing on the pre-processed signal to obtain multiple cardioid directional unit signals, and perform phase correction and gain balance optimization to obtain an optimized cardioid unit signal; An association module is used to calculate the signal covariance matrix and steering vector of the optimized cardioid unit signal according to a preset MVDR algorithm to obtain the corresponding MVDR optimal weight; A processing module, the processing module is used to perform beamforming processing on the MVDR optimal weight value and the optimized cardioid unit signal to obtain an initial beam signal; A control module, the control module is used to perform frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; The execution module is used to perform performance monitoring and parameter adjustment based on the enhanced output signal to obtain the final real-time beam signal.

[0078] The present invention provides a microphone array beamforming system that can more accurately capture the spatial information of the target sound source by synchronously calibrating and preprocessing the original sound signal and combining it with the cardioid directional unit signal obtained by differential processing, thereby improving the accuracy of beamforming and providing a reliable foundation for subsequent signal enhancement. By performing phase correction and gain balance optimization on the cardioid unit signal and calculating the covariance matrix based on the MVDR algorithm, refined processing of sound source signals in different directions is achieved, which helps to improve the enhancement effect of the target sound source and reduce environmental noise interference. Beamforming processing based on the optimized cardioid unit signal can ensure that the system can operate stably in different acoustic environments, reduce adverse effects such as reverberation, and improve the overall performance of the system. By performing frequency response equalization and dynamic range compression on the initial beam signal, a more reasonable signal enhancement strategy is formulated, and adaptive optimization of the system is achieved through real-time monitoring and parameter adjustment, thereby effectively improving voice quality. At the same time, by considering the complexity of the actual application environment, the beamforming strategy can be flexibly adjusted according to the characteristics of different scenarios and changes in demand, making the system more adaptable to diverse application scenarios.

[0079] Reference Figure 3 As shown, the present invention also provides a microphone array beamforming method and device, comprising: Memory, used to store programs; The processor is configured to execute a program to implement each step of any one of the above-mentioned microphone array beamforming methods.

[0080] In this embodiment, the processor and memory may be connected via a bus or other means. The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive. The processor may be a general-purpose processor, such as a central processing unit, a digital signal processor, an application-specific integrated circuit, or one or more integrated circuits configured to implement the embodiments of the present invention.

[0081] It should be noted that, those skilled in the art will clearly understand that, for the sake of convenience and brevity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0082] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A microphone array beamforming method, characterized in that: include: Acquire an original sound signal collected by an omnidirectional microphone, perform synchronous calibration processing on the original sound signal, and obtain a corresponding preprocessed signal; Performing differential processing on the preprocessed signal to obtain a plurality of cardioid directional unit signals, and performing phase correction and gain balance optimization to obtain an optimized cardioid unit signal; Calculating the signal covariance matrix and the steering vector of the optimized cardioid unit signal according to a preset MVDR algorithm to obtain the corresponding MVDR optimal weight; Performing beamforming processing on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal; performing frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; Performance monitoring and parameter adjustment are performed based on the enhanced output signal to obtain a final real-time beam signal.

2. The microphone array beamforming method according to claim 1, wherein: The obtaining of the original sound signal collected by the omnidirectional microphone and performing synchronous calibration processing on the original sound signal to obtain a corresponding preprocessed signal includes: Acquire the sound signal collected by the omnidirectional microphone to obtain the original sound signal; Performing analog-to-digital synchronous conversion on the original sound signal to obtain a synchronous signal group; performing a DC offset elimination process on the synchronization signal group to obtain a de-biased signal group; Performing delay compensation filtering on the de-biased signal group to obtain a delay compensation signal group; Resampling the delay compensation signal group according to a preset sampling frequency to obtain a resampled signal group; performing acoustic calibration processing on the resampled signal group to obtain a calibration signal group; Performing frame processing on the calibration signal group according to a preset frame length to obtain a frame signal; Windowing is performed on the framed signal to obtain the preprocessed signal.

3. The microphone array beamforming method according to claim 1, wherein: The differential processing of the pre-processed signal to obtain a plurality of cardioid directional unit signals, and performing phase correction and gain balance optimization to obtain an optimized cardioid unit signal includes: Grouping the preprocessed signals into groups of two adjacent microphones to obtain multiple groups of microphone signal pairs; Performing amplitude difference calculation on each pair of microphone signals to obtain an initial differential signal; Performing Fourier transform on the initial differential signal to obtain a frequency domain differential signal; Calculating the time delay parameters between adjacent microphones based on the frequency domain differential signal to obtain a phase compensation coefficient; Performing a convolution operation on the frequency domain difference signal and the phase compensation coefficient to obtain a phase correction signal; Performing gain modulation on the phase correction signal according to a preset cardioid directivity gain function to obtain a gain modulated signal; Performing an inverse Fourier transform on the gain modulated signal to obtain a time-domain cardioid unit signal; Calculating the average energy distribution according to the time-domain cardioid unit signal to obtain an energy correction coefficient; The time-domain cardioid unit signal and the energy correction coefficient are weighted to obtain the optimized cardioid unit signal.

4. The microphone array beamforming method according to claim 1, wherein: The signal covariance matrix and steering vector calculation of the optimized cardioid unit signal according to the preset MVDR algorithm to obtain the corresponding MVDR optimal weight includes: Performing frame processing on the optimized cardioid unit signal to obtain multiple signal frames; Performing windowing processing on the optimized cardioid unit signal according to the signal frame and a preset time window function to obtain a windowed signal; Performing autocorrelation and cross-correlation operations on the windowed signal to obtain a correlation signal matrix; Regularizing the correlation signal matrix according to the correlation signal matrix and a preset regularization coefficient to obtain a signal covariance matrix; Performing delay calculation on the optimized cardioid unit signal according to the array geometric layout parameters and the target sound source distance parameters to obtain delay information; Constructing a steering vector for the optimized cardioid unit signal according to the delay information and preset reference point parameters to obtain an initial steering vector; Performing weight optimization calculation on the optimized cardioid unit signal according to the signal covariance matrix and the initial steering vector to obtain the MVDR optimal weight; The formula for weight optimization calculation includes: ; is the optimal weight of the MVDR, is the signal covariance matrix, is the initial steering vector, and H is the conjugate transpose, which is used to ensure that the calculation result is a real value.

5. The microphone array beamforming method according to claim 4, wherein: The performing beamforming processing on the MVDR optimal weight value and the optimized cardioid unit signal to obtain an initial beam signal includes: performing sub-band decomposition processing on the optimized cardioid unit signal to obtain multiple frequency band signals; performing phase alignment processing on the plurality of frequency band signals to obtain an alignment signal matrix; Performing frame processing on the aligned signal matrix to obtain multiple signal subframes; Performing windowing processing on the signal subframe to obtain a windowed signal sequence; Performing Fourier transform processing on the signal sequence to obtain a frequency domain signal vector; Performing spatial correlation analysis on the frequency domain signal vector to construct a signal space feature matrix; Performing a complex matrix multiplication operation on the signal space feature matrix and the MVDR optimal weight to obtain a beam output vector; Performing inverse Fourier transform processing on the beam output vector to obtain a windowed signal sequence; Performing sub-band synthesis processing on the de-windowed signal sequence to obtain a reconstructed signal; Performing time domain compensation processing on the reconstructed signal to obtain the initial beam signal; The calculation formula for performing the complex matrix multiplication operation includes: ; is the beam output vector, is the optimal weight of the MVDR, is the signal space feature matrix.

6. The microphone array beamforming method according to claim 1, wherein: The performing frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal includes: Performing time-frequency transformation on the initial beam signal to obtain a time-frequency analysis signal; Performing energy spectrum smoothing analysis on the time-frequency analysis signal to obtain a smoothed energy spectrum; Dividing the time-frequency analysis signal into sub-bands according to the smoothed energy spectrum to obtain a plurality of equalized sub-band components; performing amplitude processing on the multiple equalized sub-band components respectively to obtain amplitude sub-band components; Calculating sub-band coefficients according to the amplitude sub-band components to obtain corresponding frequency response compensation coefficients; Performing a product operation on the frequency response compensation coefficient and the amplitude sub-band component to obtain a compensation sub-band component; Performing dynamic range detection on the compensation sub-band component to obtain a dynamic range parameter; Performing gain adjustment on the compensation sub-band component according to the dynamic range parameter to obtain a compression sub-band component; Performing sub-band synthesis on the compressed sub-band components to obtain a synthesized time-frequency signal; Performing inverse time-frequency transform on the synthesized time-frequency signal to obtain the enhanced output signal.

7. The microphone array beamforming method according to claim 1, wherein: The performing performance monitoring and parameter adjustment based on the enhanced output signal to obtain a final real-time beam signal includes: Performing short-term performance parameter calculation on the enhanced output signal to obtain performance monitoring parameters; Performing spectrum analysis on the enhanced output signal according to the performance monitoring parameters to obtain spectrum characteristic parameters; Performing threshold judgment and data statistics on the spectrum characteristic parameters to obtain performance evaluation indicators; Calculating beamforming weights according to the performance evaluation index to obtain optimized beam weights; The enhanced output signal is beam reconstructed according to the optimized beam weight to obtain the real-time beam signal.

8. A microphone array beamforming method system, characterized in that: The microphone array beamforming method according to any one of claims 1 to 7 comprises: An acquisition module is used to acquire an original sound signal collected by an omnidirectional microphone, perform synchronous calibration processing on the original sound signal, and obtain a corresponding preprocessed signal; An analysis module, configured to perform differential processing on the preprocessed signal to obtain a plurality of cardioid directional unit signals, and perform phase correction and gain balance optimization to obtain an optimized cardioid unit signal; An association module is configured to calculate a signal covariance matrix and a steering vector for the optimized cardioid unit signal according to a preset MVDR algorithm to obtain a corresponding MVDR optimal weight; a processing module, configured to perform beamforming processing on the MVDR optimal weight and the optimized cardioid unit signal to obtain an initial beam signal; a control module, configured to perform frequency response equalization and dynamic range compression on the initial beam signal to obtain an enhanced output signal; An execution module is used to perform performance monitoring and parameter adjustment based on the enhanced output signal to obtain a final real-time beam signal.

9. A microphone array beamforming method and apparatus, characterized in that: include: Memory, used to store programs; A processor is configured to execute the program to implement the steps of the microphone array beamforming method according to any one of claims 1 to 8.