Differential microphone array speech enhancement device and method based on reference point optimization

By selecting the geometric center of the microphone array as the reference point and combining it with the Jacobi-Angel expansion differential beamforming algorithm, the problems of performance degradation in the high-frequency band and white noise gain deterioration in the low-frequency band of traditional differential microphone arrays are solved, and the speech signal clarity and robustness are improved across the entire frequency band.

CN121509862BActive Publication Date: 2026-03-17WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Traditional differential microphone arrays suffer from performance degradation in the high-frequency band and deterioration of white noise gain in the low-frequency band, resulting in insufficient speech signal clarity and robustness, especially when capturing female and children's voices.

Method used

By selecting the geometric center of the microphone array as the reference point and combining it with the Jacobi-Anger expansion differential beamforming algorithm, the reference point position is optimized, a differential beamformer is designed, the approximation accuracy of the Jacobi-Anger expansion is improved, and the stability of the beam pattern shape and white noise gain across the entire frequency band are ensured.

Benefits of technology

It improves the clarity and robustness of speech signals across the entire frequency band, significantly improves beamforming performance in the high-frequency band, reduces white noise gain, and enhances the system's noise immunity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509862B_ABST
    Figure CN121509862B_ABST
Patent Text Reader

Abstract

The application provides a differential microphone array voice enhancement device and method based on reference point optimization, and relates to the technical field of voice signal processing. The application improves the reference point of the microphone array from the "first microphone" in the prior art to the "geometric center" in the application, reduces the maximum value of the distance r m of the microphone in the microphone array to the reference point by half, significantly reduces the high frequency band, and ensures the approximation accuracy of Jacobi-Anger expansion. The application optimizes the reference point, aims to realize the frequency-invariant beam pattern in the wideband range, and jointly optimizes the DF and WNG. While meeting the beam pattern approximation requirement and the distortionless constraint, the white noise gain can be maximized, and the robustness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal processing technology, and specifically to a differential microphone array speech enhancement device and method based on reference point optimization. Background Technology

[0002] Microphone arrays and beamforming technology are core technologies for recovering target speech signals from noise observations and have been widely used in various voice interaction devices. Compared to single-microphone systems, microphone arrays acquire sound field information through spatial sampling and utilize signal processing algorithms to enhance the target direction signal, suppress interference and noise in the spatial domain, significantly improving the quality of the speech signal and the robustness of the system. Microphone array technology has evolved from simple delay-summing beamforming to complex adaptive beamforming, super-directional beamforming, and differential beamforming. Among them, differential microphone arrays (... Because it has high directional gain and a frequency-invariant beam pattern (i.e., the beam shape does not change with frequency, avoiding speech distortion), it is suitable for small-sized arrays (such as those with a diameter of...). It has significant advantages in applications and has become the preferred solution for portable voice devices.

[0003] The core design principle of Directivity Detection (DMA) is to approximate the sound pressure difference (spatial derivative) between the target and interference directions to achieve directivity. Its mathematical model typically involves a series expansion of the complex exponential phase term in the array response function (steering vector) to approximate the spatial gradient of the sound pressure. Early DMAs were based on the McLaurin series, but this series was only applied between microphones. Extremely small (e.g.) The approximation effect is good when the microphone spacing is [good / good]. Increasing the beamwidth will result in severe beam distortion. (The microphone spacing is also affected.) Enlarged to sizes commonly used in practical applications (such as...) When the microphone spacing is high, higher-order term errors accumulate rapidly, leading to severe beam pattern distortion. For example, at microphone spacing... Time frequency (hereinafter referred to as frequency) At that time, the main lobe width of the beam will increase. The above methods cannot meet the requirements for directional sound pickup.

[0004] Jacobi-Anger expansion from least squares error ( From the perspective of [missing information], this is the optimal approximation method for the beam pattern exponential function. It utilizes the series form of the first-kind Bessel function to achieve approximation at larger microphone spacings. Within range (e.g., microphone spacing) Maintaining high approximation accuracy, based on this expansion The design can significantly improve beam performance. However, traditional robustness In the design, to improve robustness against array defects (such as microphone sensitivity mismatch), the minimum norm method is usually adopted and the number of microphones is increased. However, this method suffers from severe performance degradation at high frequencies.

[0005] The root cause of the aforementioned high-frequency degradation lies in the improper selection of the reference point: traditional robust The design uses the first microphone as a reference point, at which point the... The distance from each microphone to the reference point is With the number of microphones Add, remote microphone (such as the first) (number) distances from the reference point This will significantly increase the size, leading to two major problems:

[0006] 1. Decreased approximation accuracy of the Jacobi-Anger expansion: The effectiveness of this expansion depends on the dimensionless parameter " The range of values ​​for "" (usually needs to be specified) To ensure the approximation capability of the Bessel function, and to ensure that the series approximation of the Bessel function is sufficiently accurate), where c is the speed of sound in air, when When increased, the high-frequency band (such as At that time, angular frequency )of This threshold is easily exceeded, causing the approximation error of the exponential function to increase rapidly, and the designed beam pattern to deviate significantly from the ideal shape.

[0007] 2. Bessel functions approaching zero: Bessel functions of the first kind It exhibits an oscillating decay trend as x increases, in the high-frequency range. Larger Approaching zero, these tiny values ​​introduce significant numerical instability and computational errors when calculating beamforming weights, ultimately distorting the weight vector, manifesting as a directional factor. ) decrease and white noise gain ( )deterioration.

[0008] Besides the aforementioned high-frequency issues, traditional differential beamformers, in pursuit of extremely high directivity factors, often operate in the low-frequency range (such as...). The following introduces a severe white noise gain degradation problem. This is because the amplification effect of noise is more significant at low frequencies due to the differentiating operation. An excessively low WNG means the system is extremely sensitive to defects such as microphone noise floor, preamplifier noise, and sensitivity mismatch between microphones, causing low-frequency speech signals to be submerged in amplified noise, severely affecting the overall clarity, intelligibility, and naturalness of the speech. A truly robust system needs to achieve full-band (not just high or low frequencies) performance. and A good balance.

[0009] In existing technologies, although numerous studies have attempted to improve robustness by increasing the number of microphones M, employing more complex regularization methods, or directly constraining the WNG in the optimization objective, none of these methods address the fundamental architectural flaw of inappropriate reference point selection. For example, a typical traditional robust... design( )exist At that time, its actual beam pattern has a relatively low similarity to the ideal second-order supercardioid pattern, and at the same time As low as This means that white noise is amplified thousands of times. Such performance is insufficient for clearly capturing speech rich in high-frequency components (such as female or child speech). Summary of the Invention

[0010] The purpose of this invention is to provide a differential microphone array speech enhancement device and method based on reference point optimization. By selecting the geometric center of the microphone array as the reference point and combining it with the design of a differential beamformer, the performance degradation problem of existing methods in the high-frequency band is effectively solved, ensuring the clarity and robustness of the speech signal across the entire frequency band.

[0011] To achieve the above objectives, in a first aspect, the present invention provides a differential microphone array speech enhancement device based on reference point optimization, comprising a microphone array, a pre-processing module, a reference point selection module, and a signal processing module; the microphone array comprises M microphones arranged linearly at equal intervals, and the M microphones correspond to M channels respectively;

[0012] The pre-processing module is used to pre-process the M analog signals corresponding to the M channels to obtain M digital signals, and then perform frame division and windowing on the M digital signals, and then perform short-time Fourier transform to obtain frequency domain signals.

[0013] The reference point selection module is used to define the steering vector, determine the observed signal vector of the microphone array based on the steering vector and the frequency domain signal, and select the geometric center of the microphone array as the reference point.

[0014] The signal processing module is used for differential beamforming algorithm based on Jacobi-Angel expansion. It designs differential beamformer by combining reference points. The observed signal vector is input into differential beamformer to obtain the estimated value of the observed signal. Then, inverse short-time Fourier transform is performed to obtain the enhanced time domain signal.

[0015] According to the present invention, a differential microphone array voice enhancement device based on reference point optimization performs framing and windowing on M-channel digital signals, including:

[0016] Each digital signal The data is divided into overlapping time frames, and a window function is applied to the digital signal of each frame to obtain the first frame. Frame number The digital signal of the channel is:

[0017]

[0018] Where w[n] represents the window function, R represents the frame shift, and N represents the frame shift. f This represents the frame length, m=1,2,…,M.

[0019] According to the present invention, a differential microphone array speech enhancement device based on reference point optimization is provided, wherein the frequency domain signal is represented as:

[0020]

[0021] in, Indicates the first The m-th channel frequency domain signal at the k-th angular frequency point of the frame. This represents the k-th angular frequency point in the discrete frequency domain.

[0022] According to the present invention, a differential microphone array speech enhancement device based on reference point optimization has a steering vector of:

[0023]

[0024] in, The imaginary unit, = , Angular frequency, For time frequency, For the first The distance from each microphone to the reference point For the first The azimuth angle of each microphone. The speed of sound in air. The angle of incidence of the sound source is denoted by ; the superscript T indicates the transpose operation, and e is the base of the natural logarithm.

[0025] The present invention provides a differential microphone array speech enhancement device based on reference point optimization.

[0026]

[0027]

[0028] Where M represents the number of microphones. This represents the spacing between adjacent microphones.

[0029] According to the present invention, a differential microphone array speech enhancement device based on reference point optimization is provided, wherein the observed signal vector is:

[0030]

[0031] in, Indicates the guide vector. Indicates the relevant source signal, This represents the noise signal vector.

[0032] According to the present invention, a differential microphone array speech enhancement device based on reference point optimization is provided. Based on the Jacobi-Angel expansion differential beamforming algorithm, and combined with reference points, a differential beamformer is designed, comprising:

[0033] The actual beam pattern of the differential beamformer is as follows:

[0034]

[0035] in, This represents the conjugate transpose of a differential beamformer. For the first The complex weighted conjugate of each microphone;

[0036] The desired beam pattern is an Nth-order frequency-invariant pattern, represented as:

[0037]

[0038] in, For real coefficients; the desired beam pattern is written as:

[0039]

[0040] in,

[0041]

[0042] Expand the complex exponential terms in the guiding vector into a series form of a Bessel function of the first kind:

[0043]

[0044] in, It is an nth-order Bessel function of the first kind, and satisfies , The truncation order of the expansion series;

[0045] Substitute the actual beam diagram and limit the number of orders to... The order yields an approximate expression for the actual beam pattern:

[0046]

[0047] in, for vector; for Matrix, its first Line 1 Column (corresponding column index) The elements of ) are ;

[0048] To make the actual beam pattern approximate the desired beam pattern, i.e. ; combination The definition can be further transformed into:

[0049]

[0050] in, The coefficient vector of the desired beam pattern;

[0051] because = and ,therefore If the column vectors satisfy the conjugate symmetry relation, then Corresponding constraints and The corresponding constraints are completely equivalent; therefore, the original The constraints are simplified to Given several independent constraints, we obtain the optimization problem:

[0052]

[0053] in, It is an (N+1)×1 vector; It is an (N+1)×M matrix.

[0054]

[0055] The analytical solution to the optimization problem is:

[0056]

[0057] in, for The conjugate transpose of . It is the inverse matrix of the matrix product.

[0058] According to the present invention, a differential microphone array speech enhancement device based on reference point optimization inputs the observed signal vector into a differential beamformer to obtain an estimated value of the observed signal, comprising:

[0059] The enhanced single-channel frequency domain signal obtained by spatial filtering using a differential beamformer is as follows:

[0060]

[0061] in, Indicates the first The estimated value of the observed signal at the k-th angular frequency point of the m-th channel in the frame, where the superscript H indicates the conjugate transpose of the matrix or vector;

[0062] By iterating through all non-negative angular frequency points and taking conjugate symmetry for the negative angular frequency points, the complete frequency domain output signal for each frame is obtained. That is, the estimated value of the observed signal.

[0063] According to the present invention, a differential microphone array speech enhancement device based on reference point optimization is used to perform an inverse short-time Fourier transform to obtain an enhanced time-domain signal, comprising:

[0064] Output signal in frequency domain for each frame Perform inverse short-time Fourier transform to obtain the time-domain frame signals:

[0065]

[0066] The overlapping addition method is used to synthesize the signals from each time-domain frame into a continuous enhanced time-domain signal:

[0067] .

[0068] In a second aspect, the present invention provides a reference-point optimized differential microphone array speech enhancement method, applied to the reference-point optimized differential microphone array speech enhancement device of the first aspect, the method comprising:

[0069] After preprocessing the M analog signals corresponding to the M channels, M digital signals are obtained. The M digital signals are then framed and windowed, and then short-time Fourier transform is performed to obtain the frequency domain signals.

[0070] Define a steering vector, determine the observed signal vector of the microphone array based on the steering vector and the frequency domain signal, and select the geometric center of the microphone array as the reference point;

[0071] Based on the differential beamforming algorithm of Jacobi-Angel expansion, a differential beamformer is designed in conjunction with a reference point. The observed signal vector is input into the differential beamformer to obtain the estimated value of the observed signal, and then an inverse short-time Fourier transform is performed to obtain the enhanced time-domain signal.

[0072] The present invention has at least the following beneficial effects:

[0073] This invention provides a differential microphone array speech enhancement device and method based on reference point optimization. By improving the reference point of the microphone array from the "first microphone" in existing technologies to the "geometric center" of this invention, rm The maximum value is reduced by half, high frequency band This significantly reduces noise and ensures the approximation accuracy of the Jacobi-Anger expansion. The invention aims to achieve frequency-invariant beammaps over a wide bandwidth and joint optimization of DF and WNG through reference point optimization. While meeting beammap approximation requirements and distortion-free constraints, it maximizes white noise gain and improves robustness. Attached Figure Description

[0074] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0075] In the attached diagram:

[0076] Figure 1 This is a structural block diagram of the differential microphone array speech enhancement device based on reference point optimization according to the present invention;

[0077] Figure 2 This is a flowchart of the differential microphone array speech enhancement method based on reference point optimization according to the present invention;

[0078] Figure 3 (a), (b), (c), and (d) are respectively the ideal second-order supercardioid beam pattern, the beam pattern under conventional DMA (M=3), the beam pattern under conventional robust DMA (M=8), and the beam pattern under robust DMA (M=8) proposed in this invention;

[0079] Figure 4 (a) and (b) are respectively and A schematic diagram showing the relationship between frequency and frequency. Detailed Implementation

[0080] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0081] The following detailed description of some embodiments of the present invention will be provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0082] Please see Figure 1This invention provides a differential microphone array voice enhancement device based on reference point optimization, including a microphone array, a pre-processing module, a reference point selection module, and a signal processing module.

[0083] Microphone arrays, as signal acquisition front-ends, are typically designed with a compact structure to suit embedded devices such as smart speakers and conference terminals. In some embodiments, the microphone array employs a uniform linear array, consisting of... An omnidirectional microphone at intervals The microphone array is composed of uniformly arranged microphones, with a total length of [missing information]. Where M represents the number of microphones, which can be selected according to performance requirements, typically... If higher-order beamforming is required (such as beam pattern order) ) can be extended to ; The spacing between adjacent microphones is typically used to ensure the effectiveness of differential beamforming. And satisfy δ λ min , where λ min For the highest operating frequency f max Corresponding wavelength (e.g.) hour ).

[0084] The microphone array includes M microphones, each corresponding to one of the M channels, thus outputting M analog signals. Each analog signal undergoes pre-processing by a pre-processing module, including the following steps:

[0085] Preamplifier. Composed of a low-noise operational amplifier, it amplifies weak analog signals to a level suitable for sampling by an ADC (analog-to-digital converter). The gain of the operational amplifier is adjustable to adapt to different sound pressure level environments;

[0086] Anti-aliasing filtering. A low-pass filter (typically Butterworth or Chebyshev type) is used, with a cutoff frequency slightly higher than the highest operating frequency f. max (e.g., 8kHz) to filter out high-frequency noise and prevent sampling aliasing;

[0087] Analog-to-digital conversion. A multi-channel synchronous sampling ADC is used, operating at the same sampling rate f. s M analog signals (e.g., 16kHz or 48kHz) are synchronously digitized to obtain M digital signals in discrete time series form. m=1,2,…,M; then the M digital signals are preprocessed (usually by digital signal processing (DSP) or central processing unit (CPU)), for example, DC offset removal, subtracting the DC component of the digital signal;

[0088] Pre-emphasis filtering. A first-order high-pass filter is applied to enhance the high-frequency components of the speech, balance the spectrum, and facilitate subsequent processing;

[0089] Framing and windowing. Speech signals are non-stationary, but within a short timeframe (…). Within this range, the signal can be considered quasi-stationary. Therefore, each digital signal is divided into overlapping short-time frames. The frame length is... The corresponding time length is usually 1000. (For example in) hour, Point); Frame shift : Usually the length of a frame or To achieve a smooth transition; window function w[n]: Apply a window function (such as Hamming window or Hanning window) to each frame of digital signal to reduce spectral leakage caused by frame truncation. This yields the... Frame number The digital signal of the channel is:

[0090]

[0091] Time-frequency domain transformation. A short-time Fourier transform is performed on the time-domain signal of the m-th channel in the τ-th frame to obtain the frequency-domain signal, represented as:

[0092]

[0093] in, Indicates the first The m-th channel frequency domain signal at the k-th frequency point of the frame. This represents the k-th angular frequency point in the discrete frequency domain, i.e., the frequency point.

[0094] Converting the speech signal from the time domain to the frequency domain makes it easier to perform beamforming processing independently at each frequency point.

[0095] like Figure 1 As shown, the existing method uses the first microphone as the reference point ( Figure 1 Point A in the middle). This invention uses the geometric center of the microphone array as the reference point (…). Figure 1 By using point O in the middle, the maximum distance from each microphone to the reference point will be significantly reduced, which will help improve the approximation accuracy of the Jacobi-Anger expansion and thus improve high-frequency performance.

[0096] For the signal model at the reference point of existing methods, consider a model composed of... A uniform linear array of omnidirectional microphones, with an adjacent microphone spacing of [missing information]. Under the far-field plane wave assumption, taking the first microphone as the reference point, for a uniform linear array located on the x-axis, assuming the sound wave propagates in a plane, the incident angle of the sound source is... Defined as the microphone array normal ( If the angle between the axes is 0 and 1, then the guiding vector can be expressed as:

[0097]

[0098] in, The value is the speed of sound, and the superscript T indicates the transpose operation. Angular frequency, For time frequency.

[0099] The spatial distance from the m-th microphone to the reference point is r. m =(m 1) If δ is , then the above formula can be written as:

[0100]

[0101] Therefore, we can conclude that when Larger, and angular frequency When higher, the exponential term The value of is very large, which not only leads to the aforementioned Jacobi-Anger expansion approximation problem, but also poses challenges in numerical computation.

[0102] Based on this, the reference point selection module is used to redefine the guide vector, obtain the observed signal vector of the microphone array based on the guide vector, and select the geometric center of the microphone array as the reference point. This invention optimizes the reference point position, allowing the reference point to be located at any position within the geometric range of the microphone array. Therefore, the guide vector can be reformulated as:

[0103]

[0104] in, The imaginary unit, = , It is angular frequency. It is time frequency. It is the first The distance from each microphone to the reference point It is the first The azimuth angle of each microphone. It is the speed of sound in air, usually assumed to be... . The angle of incidence of the sound source is denoted by ; the superscript T indicates the transpose operation, and e is the base of the natural logarithm.

[0105] Assuming the sound source travels to the microphone array from any direction, the observed signal vector of the microphone array can be represented as:

[0106]

[0107] in, The steering vector, which is the core of beamforming technology, describes the angle of incidence of a sound wave from the sound source. The relative phase delay generated by each microphone in the microphone array. Indicates the relevant source signal, Represents a noise signal vector; Indicates the first The frequency domain signal of the m-th channel at the k-th frequency point of the frame, where m = 1, 2, ..., M. The superscript H indicates the conjugate transpose of the matrix or vector.

[0108] Specifically, the signal processing module is used to design a differential beamforming algorithm based on Jacobi-Angel expansion, combined with a reference point; the observed signal vector is input into the differential beamforming to obtain the estimated value of the observed signal, and then an inverse short-time Fourier transform is performed to obtain the enhanced time-domain signal.

[0109] It should be noted that traditional differential beamformers typically use the first microphone as a reference point, i.e. , The distance between adjacent microphones, azimuth angle In this case, as the number of microphones in the microphone array increases, r m The increase in size leads to a performance degradation of Jacobi-Anger-based beamforming in the high-frequency band.

[0110] This invention uses a reference point selection module to set the reference point at the geometric center of the microphone array. For a uniformly linear microphone array, selecting the geometric center of the microphone array as the reference point results in:

[0111]

[0112]

[0113] Therefore, if M is odd, the microphone at the geometric center is used as the reference point; if M is even, the midpoint between the two microphones closest to the center is chosen as the reference point. This minimizes the maximum distance from all microphones to the reference point, resulting in a higher-performance steering vector. This improves the approximation accuracy of the Jacobi-Anger expansion.

[0114] In beamforming design, the goal is to target each frequency point and the expected angle of incidence of the sound source Calculate a differential beamformer, i.e., a complex weight vector. , indicating a length of Differential beamformer, i.e. The spatial filter coefficients of each microphone enable the output of the differential beamformer to enhance the desired sound source incident angle. While providing directional signals, it also suppresses noise and interference to the greatest extent possible.

[0115] First, let's define a few key metrics:

[0116] Directivity factor (DF): measures the ability of a beamformer to suppress isotropic diffused noise (such as room reverberation).

[0117]

[0118] in It is the spatial coherence matrix of the diffuse noise field. The larger the value, the stronger the resistance to diffused noise.

[0119] White noise gain (WNG): measures the sensitivity of a beamformer to unrelated noise (such as microphone background noise, circuit noise, etc.).

[0120]

[0121] The larger the value (or the smaller the negative decibel value), the better the system's robustness to its own defects. Too low is traditional The main problem.

[0122] Beamformation: Describes the complex gain of a beamformer for plane waves incident at different sound sources at angles θ.

[0123] Among them, a differential beamforming algorithm based on Jacobi-Angel expansion is used, combined with a reference point, to design a differential beamformer, including:

[0124] The actual beam pattern of the differential beamformer is as follows:

[0125]

[0126] in, This represents the conjugate transpose of a differential beamformer. For the first The complex weighted conjugate of each microphone, For the first The distance from each microphone to the reference point Its azimuth angle, For the speed of sound, It is the imaginary unit.

[0127] The desired beam pattern is an Nth-order frequency-invariant pattern (a second-order supercardioid is used in this embodiment), and its general form can be expressed as:

[0128]

[0129] in These are real coefficients that determine the directivity shape of the beam. To establish a connection with the Jacobi-Anger expansion, the desired beam pattern is rewritten as:

[0130]

[0131] The coefficients satisfy a symmetric relationship:

[0132]

[0133] Using the Jacobi-Anger expansion (the best least-squares approximation of the beammap exponential function), the complex exponential term in the steering vector can be expanded into a series form of a Bessel function of the first kind:

[0134]

[0135] in It is an nth-order Bessel function of the first kind, and satisfies , The truncation order of the expansion series.

[0136] Substituting the above expansion into the actual beam diagram and limiting the number of orders to... By determining the order, we can obtain an approximate expression for the actual beam pattern:

[0137]

[0138] in, for vector; for Matrix, its first Line 1 Column (corresponding column index) The elements of ) are .

[0139] The core objective of beamforming is to make the actual beam pattern approximate the desired beam pattern, that is... Combining The definition can be further transformed into:

[0140]

[0141] in This represents the desired beam pattern coefficient vector.

[0142] because = (Symmetric coefficients) and (Since the Bessel function is symmetric), we can obtain The column vectors satisfy a conjugate symmetry relation, therefore Corresponding constraints and The corresponding constraints are completely equivalent. Based on this, the original... The constraints are simplified to With one independent constraint, the above equation can be simplified to an optimization problem:

[0143]

[0144] in, It is an (N+1)×1 vector containing the independent coefficients of the desired beam pattern. It is an (N+1)×M matrix.

[0145]

[0146] The analytical solution to this optimization problem is:

[0147]

[0148] in, for The conjugate transpose of . Let be the inverse matrix of the matrix product. This solution, while satisfying the beammap approximation requirement and the distortion-free constraint, maximizes the white noise gain. This improves system robustness.

[0149] Specifically, the observed signal vector is input into the differential beamformer to obtain an estimated value of the observed signal, which is then subjected to an inverse short-time Fourier transform to obtain the enhanced time-domain signal, including:

[0150] The enhanced frequency domain signal obtained by spatial filtering using a differential beamformer is:

[0151]

[0152] in, Indicates the first The estimated value of the observed signal at the k-th angular frequency point of the m-th channel in the frame, where the superscript H indicates the conjugate transpose of the matrix or vector.

[0153] Traverse all non-negative angular frequency points ( By taking conjugate symmetry around the negative angular frequency points, a complete frame of frequency domain output signal is obtained. .

[0154] Output signal in frequency domain for each frame Perform inverse short-time Fourier transform to obtain the time-domain frame signal. :

[0155]

[0156] Then, the overlapping addition method is used to synthesize the signals from each time-domain frame into a continuous enhanced time-domain signal. :

[0157]

[0158] Even after beamforming, the enhanced speech signal may still contain residual noise, echoes, and non-stationary interference. In practical applications, a certain post-processing step is required before it can be used in downstream tasks. Therefore, the differential microphone array speech enhancement device of this invention may further include a post-processing module. This module performs echo cancellation, noise suppression, and gain control on the enhanced time-domain signal, outputting the desired enhanced speech signal. Echo cancellation ensures that speech quality will not degrade or cause false wake-ups due to echo interference during application. Noise suppression effectively addresses background noise and transient interference in the environment while maintaining the naturalness of the speech. Gain control avoids auditory discomfort caused by sudden gain changes, ensuring stable signal amplitude received by downstream tasks. Of course, the specific implementation of the post-processing module is already available in existing technology, and will not be elaborated upon here.

[0159] Based on the same inventive concept, another embodiment of the present invention provides a reference-point optimized differential microphone array speech enhancement method, applied to the reference-point optimized differential microphone array speech enhancement device of the aforementioned embodiment, such as... Figure 2 As shown, the method includes:

[0160] Step 1: Perform preprocessing on the M analog signals corresponding to the M channels to obtain M digital signals, then perform frame division and windowing on the M digital signals, and then perform short-time Fourier transform to obtain frequency domain signals.

[0161] Step 2: Define the steering vector, determine the observed signal vector of the microphone array based on the steering vector and the frequency domain signal, and select the geometric center of the microphone array as the reference point;

[0162] Step 3: Based on the Jacobi-Angel expansion differential beamforming algorithm, and combined with the reference point, design a differential beamformer; input the observed signal vector into the differential beamformer to obtain the estimated value of the observed signal, and then perform inverse short-time Fourier transform to obtain the enhanced time-domain signal.

[0163] This invention presents a speech enhancement method based on reference point optimization, which demonstrates significant performance advantages in simulations. For example... Figure 3 As shown, in conditions, Figure 3(a), (b), (c), and (d) represent the ideal beam pattern, the beam pattern under conventional DMA (M=3), the beam pattern under conventional robust DMA (M=8), and the beam pattern under the robust DMA (M=8) proposed in this invention, respectively. Conventional DMA refers to M=3, with the first microphone as the reference point, based on simple delay summation or McLaurin series design, resulting in extremely low WNG (less than or equal to -40dB) and severe noise amplification. Conventional robust DMA refers to M=5 or 8, with the first microphone as the reference point, improving robustness based on the minimum norm method, but with decreased DF and deteriorated WNG at high frequencies (WNG≤-30dB at 3kHz). The beam pattern of this invention highly matches the ideal second-order supercardioid pattern, while the conventional robust DMA... The beam pattern has become significantly distorted. The advantages of this invention compared to the two traditional methods are as follows:

[0164] DF: The solution of this invention (8dB) > Traditional robust DMA (3dB) > Traditional DMA (2dB);

[0165] WNG: The solution of this invention (-10dB) > conventional robust DMA (-30dB) > conventional DMA (-40dB).

[0166] Therefore, the beam pattern of this invention has a high similarity to the ideal second-order supercardioid pattern, and is significantly more robust than traditional beam patterns. The similarity to an ideal second-order supercardioid pattern demonstrates that the core advantage of this invention lies in the high-frequency range (above 2kHz), where traditional robust DMA performance collapses, while the invention remains stable.

[0167] like Figure 4 As shown in (a), the present invention Da Compared to traditional methods promote times; such as Figure 4 As shown in (b), the present invention Da Compared to traditional methods promote Therefore, this invention is... and It outperforms both traditional methods in two key metrics, especially in The above high-frequency bands, Improvement This effectively solves the performance degradation problem of traditional methods in the high-frequency band, ensuring the clarity and robustness of voice signals across the entire frequency band.

[0168] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that the invention is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A differential microphone array speech enhancement device based on reference point optimization, characterized in that, The device comprises a microphone array, a pre-processing module, a reference point selection module and a signal processing module; the microphone array comprises M linearly equidistantly arranged microphones, and the M microphones correspond to M channels respectively; The pre-processing module is used for pre-processing M analog signals corresponding to the M channels respectively to obtain M digital signals, and performing framing and windowing on the M digital signals, and then performing short-time Fourier transform to obtain frequency domain signals; The reference point selection module is used for defining a steering vector, determining an observation signal vector of the microphone array based on the steering vector and the frequency domain signals, and selecting a geometric center of the microphone array as a reference point; The signal processing module is used for designing a differential beamformer based on a differential beamforming algorithm based on Jacobi-Anger expansion in combination with the reference point, inputting the observation signal vector into the differential beamformer to obtain an observation signal estimate, and then performing inverse short-time Fourier transform to obtain an enhanced time domain signal.

2. The reference point optimization based differential microphone array speech enhancement apparatus according to claim 1, wherein, The framing and windowing on the M digital signals comprises: Each digital signal The data is divided into overlapping time frames, and a window function is applied to the digital signal of each frame to obtain the first frame. Frame number The digital signal of the channel is: where w[n] represents a window function, R represents a frame shift, N f represents a frame length, m = 1, 2, …, M.

3. The reference-point-based optimization differential microphone array speech enhancement apparatus according to claim 2, wherein, The frequency domain signal is expressed as: wherein, represents the mth channel of the frame at the kth angular frequency point, represents the mth channel of the frame at the kth angular frequency point, represents the kth angular frequency point of the discrete frequency domain.

4. The reference-point-based optimization differential microphone array speech enhancement apparatus according to claim 3, wherein, The steering vector is: wherein is the imaginary unit, = 1, , m = 1, 2,..., M, is the angular frequency, is the time frequency, is the distance from the th microphone to the reference point, are the azimuth angles of the 1st, 2nd, Mth microphones, respectively, is the sound speed in air, is the sound source incidence angle; the superscript T represents the transpose operation, and e is the base of the natural logarithm.

5. The differential microphone array speech enhancement device based on reference point optimization according to claim 4, characterized in that, wherein M is the number of microphones, is the distance between adjacent microphones.

6. The reference-point-based optimization differential microphone array speech enhancement apparatus according to claim 5, wherein, The observation signal vector is: wherein denotes a steering vector, denotes a related source signal, denotes a noise signal vector.

7. The reference-point-based optimization differential microphone array speech enhancement apparatus according to claim 6, wherein, The differential beamformer is designed based on the differential beamforming algorithm based on Jacobi-Anger expansion in combination with the reference point, and comprises: The actual beam pattern of the differential beamformer is: wherein, denotes the conjugate transpose of the difference beamformer, is the complex weight conjugate for the th microphone. The desired beam pattern is an N-order frequency-invariant directivity pattern, and is expressed as: wherein are real coefficients; write the desired beam pattern as: wherein, The complex exponential term in the steering vector is expanded into a series form of the first Bessel function: wherein is the n-th order first kind Bessel function and satisfies , is the truncation order of the expansion series; Substitute the actual beam pattern and limit the series to order, the approximate expression of the actual beam pattern is obtained as in, for vector; for Matrix, its first Line number The elements of the column are Column index corresponding ; approximating the actual beam pattern to the desired beam pattern, i.e. ; in combination with the definition of further translates into: wherein, is a coefficient vector of the desired beam pattern; Since = and , therefore the column vector of satisfies the conjugate symmetric relation, then the corresponding constraint is completely equivalent to the corresponding constraint; thus the original constraints are simplified to independent constraints, and the optimization problem is obtained: wherein is an (N + 1) x 1 vector; is an (N + 1) x M matrix, The analytical solution of the optimization problem is: wherein is the conjugate transpose of is the inverse matrix of the matrix product.

8. The reference-point-based optimization differential microphone array speech enhancement apparatus according to claim 7, wherein, The observation signal vector is input into the differential beamformer to obtain an observation signal estimate, and comprises: The enhanced single-channel frequency domain signal is obtained by spatial filtering through the differential beamformer, and is: wherein represents the mth Ym(k) represents the observed signal estimate at the kth angular frequency point of the mth channel of the frame, and the superscript H represents the conjugate transpose of a matrix or vector; The complete frequency domain output signal of each frame is obtained by traversing all non-negative angular frequency points and taking the conjugate symmetry for negative angular frequency points i.e. the observation signal estimate.

9. The reference-point-based optimization differential microphone array speech enhancement apparatus according to claim 8, wherein, The inverse short-time Fourier transform is performed to obtain an enhanced time domain signal, and comprises: for each frame of the frequency domain output signal performing an inverse short-time Fourier transform to obtain a respective time domain frame signal: The overlap-add method is used to synthesize continuous enhanced time domain signals from the time domain frame signals: 。 10. A method for differential microphone array speech enhancement based on reference point optimization, the method comprising: The method is applied to the differential microphone array speech enhancement device based on reference point optimization as claimed in any one of claims 1-9, and the method comprises: M analog signals corresponding to M channels are pre-processed respectively to obtain M digital signals, and framing and windowing are performed on the M digital signals, and then short-time Fourier transform is performed to obtain frequency domain signals; A steering vector is defined, an observation signal vector of the microphone array is determined based on the steering vector and the frequency domain signals, and a geometric center of the microphone array is selected as a reference point; A differential beamformer is designed based on a differential beamforming algorithm based on Jacobi-Anger expansion in combination with the reference point, the observation signal vector is input into the differential beamformer to obtain an observation signal estimate, and then inverse short-time Fourier transform is performed to obtain an enhanced time domain signal.

Citation Information

Patent Citations

  • Audio signal processing method and device and difference beam forming method and device

    CN104464739A

  • Pickup method of microphone array, electronic equipment and storage medium

    CN117037830A