Microphone combination array beam forming method, system and equipment

Through the multi-objective optimization and improved Jacobiang function algorithm in the microphone composite array, the problem of unstable beamforming performance of traditional microphone arrays in complex acoustic environments is solved, and high performance and robust beamforming is achieved, combining high white noise gain and high directionality for the full frequency.

CN120091248APending Publication Date: 2025-06-03YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510189986.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional microphone arrays have unstable beamforming performance in complex or dynamic acoustic environments, insufficient white noise gain, limited beam direction, and existing beamforming algorithms have problems such as inverse proportional white noise gain and directionality, uneven beam width, and inability to cover the complete pitch angle range.

Method used

The geometric distribution of the circular microphone array is determined through multi-objective optimization, combined with the improved Jacobian function to generate beam coefficients in the desired direction, and the beam output is obtained by weighting summing.

Benefits of technology

It realizes a high-performance microphone array with adjustable constant beam width, combining full-frequency high directionality and white noise gain, improving beam robustness, avoiding low-frequency noise amplification problems, and accurately covering the entire pitch angle range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091248A_ABST
    Figure CN120091248A_ABST
Patent Text Reader

Abstract

The invention discloses a microphone combination array beam forming method, system and device. The method comprises the following steps: determining geometric distribution of an annular microphone array; performing multi-objective optimization on the microphone array to obtain a plurality of groups of microphone arrays of a first combination, each combination in the microphone arrays of the first combination comprising combinations of microphone numbers and radiuses of different concentric rings; performing iterative optimization on the microphone array of the first combination to obtain a microphone array of a second combination; acquiring a sound source signal by using the microphone array of the second combination, and acquiring actual azimuth information of the sound source signal; generating a beam coefficient in an expected direction from the actual azimuth information of the sound source signal through an improved Jacobian function; and carrying out weighted summation on the beam coefficient and the collected microphone array sound source signal to obtain beam output. The high-performance microphone array with the adjustable constant beam width can be realized, and the directivity, the white noise gain and the beam robustness are improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of microphone beamforming, and in particular, to a microphone combined array beamforming method, system and device. Background Art

[0002] Microphone array technology has been widely used in various scenarios, such as teleconferences, speech recognition systems, and environmental sound monitoring. In these application scenarios, to achieve good results, the key lies in being able to capture and process audio signals from different directions, and effectively suppress background noise, reverberation, and various interference signals while enhancing the signals in the desired direction. One of the key technologies in microphone array signal processing is beamforming. The beamforming technology creates an output pattern that enhances the signal in the specified direction, called a "beam", by assigning specific weight coefficients to multiple microphone signals arranged at specific positions and performing product and summation operations. By adjusting the weight coefficients of each microphone in the array, beamforming can direct the focus of the array to the desired signal source, thereby reducing interference from unwanted directions. Usually, the white noise gain and the directivity factor (directivity) are two important indicators used to measure the performance of the beam. Among them, the white noise gain determines the ability of the beam to suppress the omnidirectional background noise floor, and the directivity factor determines the ability of the beam to suppress the directional noise. Therefore, how to select a suitable array configuration and reasonably allocate the weight coefficients of different microphones in combination with the beamforming algorithm to obtain a beam with both higher white noise gain and directivity factor has always been a major challenge in beamforming optimization technology.

[0003] However, in traditional microphone array configurations, such as single-ring uniform circular arrays and uniform linear arrays, there are problems of insufficient white noise gain, limited beam directivity, and non-uniformity of the beam width, which will lead to unstable beamforming performance in complex or dynamic acoustic environments; and in the existing beamforming algorithms, there are limitations such as the white noise gain and directivity being inversely proportional, uneven beam width, inability to cover the complete elevation angle range (0° to 90°), and easy occurrence of noise amplification problems at lower frequencies, which limit the application of this technology in full-space beamforming. The above problems need to be solved. Summary of the Invention

[0004] The main objective of this application is to overcome the drawbacks and deficiencies of the prior art, and provide a microphone combined array beamforming method, system and device, which can achieve a high-performance microphone array with adjustable and constant beam width, and at the same time has full-frequency high directivity and white noise gain, and improves the robustness of the beam, avoids the problem of low-frequency noise amplification, and enables it to accurately cover the entire pitch angle range to meet the application requirements of beamforming and sound source localization in three-dimensional space. It ensures that the microphone array has good beam response in the full frequency band (0 - 8 kHz), achieving stronger beam performance and better pick-up quality.

[0005] To achieve the above objective, this application adopts the following technical solutions:

[0006] In the first aspect, this application provides a microphone combined array beamforming method, including the following steps:

[0007] Determine the geometric distribution of the positions of several microphones in the circular microphone array according to the preset constraint conditions;

[0008] Perform multi-objective optimization on the geometric distribution of the positions of several microphones in the microphone array to obtain several groups of microphone arrays of the first combination that meet the constraint conditions, where each combination in the microphone array of the first combination includes a combination of the number of microphones and the radius of different concentric circular rings;

[0009] Perform iterative optimization on the microphone array of the first combination to obtain a microphone array of the second combination;

[0010] Use the microphone array of the second combination to collect the sound source signal, and obtain the actual azimuth information of the sound source signal through sound source localization;

[0011] Generate the beam coefficient in the desired direction by passing the actual azimuth information of the sound source signal through the improved Jacobian function;

[0012] Perform weighted summation on the beam coefficient and the collected sound source signal of the microphone array to obtain the beam output.

[0013] As a preferred technical solution, the preset constraint conditions include the upper limit of the number of microphones, the total number of concentric circular rings, the range of the number of microphones per ring, and the range of the radius of the concentric circular rings.

[0014] As a preferred technical solution, the multi-objective optimization includes maximizing the white noise gain and directivity factor of the microphone array beam, minimizing the spatial aliasing effect and sidelobe level of the microphone array beam, and limiting the main beam width of the microphone array within the desired range.

[0015] As a preferred technical solution, the combination of the number of microphones and the radius of different concentric circular rings included in the second combined microphone array is as follows:

[0016] The radii of different concentric circular rings are, from the inside to the outside: 0 m, 0.021 m, 0.044 m, 0.065 m, 0.090 m, 0.130 m, 0.170 m, 0.255 m;

[0017] Among them, the number of microphones corresponding to the concentric circular ring with a radius of 0 m is 1, the number of microphones corresponding to the concentric circular ring with a radius of 0.021 m is 12, and the number of microphones corresponding to the concentric circular rings with radii of 0.044 m, 0.065 m, 0.090 m, 0.130 m, 0.170 m, and 0.255 m is 19 each.

[0018] As a preferred technical solution, generating the beam coefficient in the desired direction by using the improved Jacobi-Anger function for the actual azimuth information of the sound source signal includes:

[0019] Defining the order of the Bessel function, where the order of the Bessel function is used to represent the wave propagation characteristics in the circular microphone array configuration;

[0020] Calculating the phase change of the sound wave when propagating in the microphone array for each microphone element in the microphone array at different frequencies, different positions, and different sound source elevation angles to obtain a phase matrix;

[0021] Based on the phase matrix, calculating the Bessel function values of each Bessel order at each frequency by using the first kind of Bessel function;

[0022] Based on the order of the Bessel function and the angle of the microphone element, calculating the phase shift of each microphone element;

[0023] According to the phase shift of each microphone element, constructing a comprehensive phase shift matrix;

[0024] Based on the comprehensive phase shift matrix, calculating the beam coefficient in the desired direction.

[0025] As a preferred technical solution, the value range of the sound source elevation angle is 0 to 90°.

[0026] As a preferred technical solution, calculating the phase change of the sound wave when propagating in the microphone array for each microphone element in the microphone array at different frequencies, different positions, and different sound source elevation angles includes:

[0027] Obtaining the sampling frequency of the sound wave signal collected by the microphone array;

[0028] Based on the sampling frequency, calculate the maximum frequency;

[0029] Obtain a scaling factor according to the ratio of the sampling frequency to the maximum frequency;

[0030] Based on the maximum frequency and the speed of sound, calculate the shortest wavelength when the sound source signal received by the microphone array;

[0031] Obtain the normalized radius of each concentric ring in the microphone array;

[0032] Set the number of Fourier transform points, and calculate multiple frequency components based on the number of Fourier transform points and the sampling frequency;

[0033] Calculate the phase change when the sound wave propagates in the microphone array according to the scaling factor, the shortest wavelength, the normalized radius and the frequency components.

[0034] As a preferred technical solution, the calculating the beam coefficient in the desired direction based on the comprehensive phase shift matrix includes:

[0035] Define an identity matrix having the same number of rows as the comprehensive phase shift matrix;

[0036] Multiply the comprehensive phase shift matrix and the conjugate transpose of the comprehensive phase shift matrix to obtain a new matrix;

[0037] Based on the new matrix and the identity matrix scaled by the regularization factor, obtain a regularization matrix;

[0038] Perform a pseudo-inverse calculation on the regularization matrix to obtain the beam coefficient in the desired direction.

[0039] In a second aspect, the present application provides a microphone combined array beamforming system, which is applied to the microphone combined array beamforming method described above, and includes a determining array distribution module, a target optimization module, an iterative optimization module, an obtaining azimuth information module, a generating beam coefficient module, and a beam output module;

[0040] The determining array distribution module is configured to determine the geometric distribution of the positions of several microphones in the circular microphone array according to preset constraint conditions;

[0041] The target optimization module is configured to perform multi-objective optimization on the geometric distribution of the positions of several microphones in the microphone array to obtain several sets of the first combined microphone arrays that satisfy the constraint conditions, wherein each combination in the first combined microphone arrays includes a combination of the number of microphones and the radius of different concentric rings;

[0042] The iterative optimization module is used to iteratively optimize the microphone array of the first combination to obtain the microphone array of the second combination;

[0043] The azimuth information acquisition module is used to collect the sound source signal by using the microphone array of the second combination and obtain the actual azimuth information of the sound source signal through sound source localization;

[0044] The beam coefficient generation module is used to generate the beam coefficient in the desired direction by passing the actual azimuth information of the sound source signal through an improved Jacobi-Anger function;

[0045] The beam output module is used to perform weighted summation on the beam coefficient and the collected sound source signal of the microphone array to obtain the beam output.

[0046] In a third aspect, the present application provides an electronic device, which includes:

[0047] At least one processor; and a memory communicatively connected to the at least one processor;

[0048] Wherein, the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the method for beamforming of a microphone combined array.

[0049] In summary, compared with the prior art, the effective effects brought by the technical solution provided by the present application at least include:

[0050] The present application proposes a beamforming method for a microphone combined array, which determines the geometric distribution of the positions of several microphones in an annular microphone array according to preset constraint conditions; performs multi-objective optimization on the geometric distribution of the positions of several microphones in the microphone array to obtain several groups of microphone arrays of the first combination that meet the constraint conditions, where each combination in the microphone array of the first combination includes a combination of the number of microphones and the radius of different concentric circular rings; performs iterative optimization on the microphone array of the first combination to obtain the microphone array of the second combination; uses the microphone array of the second combination to collect the sound source signal, and obtains the actual azimuth information of the sound source signal through sound source localization; generates the beam coefficient in the desired direction by passing the actual azimuth information of the sound source signal through an improved Jacobi-Anger function; performs weighted summation on the beam coefficient and the collected microphone array signal to obtain the beam output. By combining the combined optimization of the concentric circular array and the improved Jacobi-Anger function beamforming algorithm, the present application can significantly improve the directivity factor of the microphone array beam and enhance the white noise gain, and maintain an approximately consistent beam width within the entire frequency range. The beam width can be freely adjusted between 20° and 60°, and the robustness of the beam is improved, avoiding the problem of low-frequency noise amplification, and enabling it to accurately cover the entire pitch angle range to meet the application requirements of beamforming and sound source localization in three-dimensional space, achieving stronger beam performance and better sound pickup quality. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0052] Figure 1 It is a flowchart of a beamforming method for a microphone combined array provided by an embodiment of the present application;

[0053] Figure 2 It is a schematic diagram of the geometric distribution of a microphone combined array provided by an embodiment of the present application;

[0054] Figure 3 It is a block diagram of a beamforming system for a microphone combined array provided by an embodiment of the present application. Detailed Embodiments

[0055] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0056] The mention of "embodiment" in this application means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0057] As an important technology in the field of audio signal processing, microphone array technology has received extensive attention and applications in various application scenarios such as teleconferences, speech recognition systems, and environmental sound monitoring in recent years. These application scenarios require the system to be able to capture and process audio signals from different directions, while effectively suppressing background noise, reverberation, and various interference signals, and enhancing the signals in the desired direction to achieve high-quality audio acquisition and processing effects.

[0058] Beamforming technology is one of the key means to achieve the above goals. It assigns specific weight coefficients to the signals of multiple microphones arranged at specific positions and performs product and summation operations, thereby creating an output pattern that enhances the signal in the specified direction, namely the "beam". By adjusting the weight coefficients of each microphone in the array, beamforming technology can direct the focus of the array to the desired signal source, thereby reducing interference from unwanted directions.

[0059] However, traditional microphone array configurations, such as single-ring uniform circular arrays and uniform linear arrays, have several significant drawbacks. First, the white noise gain is insufficient, which limits the ability to extract clear speech signals in low signal-to-noise ratio environments; second, the beam directivity is limited, making their suppression effect on interference directions poor when focusing on sound sources from specific directions, especially in the low-frequency part; third, the non-uniformity of the beam width, the beam response of traditional arrays usually shows a trend that the beam width shrinks as the frequency increases, which will lead to unstable beamforming performance in complex or dynamic acoustic environments.

[0060] To overcome the deficiencies of traditional arrays, the exploration of multi-loop combined arrays, especially concentric circular arrays, has begun. Compared with linear arrays, concentric circular arrays possess 360° omnidirectional uniformity, which can effectively solve the problems existing in traditional arrays. For example, using multiple concentric circles (increasing the number of microphones) can further enhance the white noise gain of the beam, thereby improving the clarity of signals in low signal-to-noise ratio environments. In addition, the concentric circle array configuration combination can also significantly improve the directivity of the beam. By optimally designing and combining array microphones with different beam characteristics at different radii, a higher directivity factor than that of traditional single loops can be obtained, thus more effectively suppressing or even completely shielding interference signals. Another key advantage of multi-loop combined arrays is the ability to control a constant beam width within the broadband speech frequency range (0 - 8 kHz). This characteristic effectively solves the problem of non-uniform traditional beam widths, making the pick-up effect in complex scenarios more stable. Although multi-loop combined arrays have many advantages, the realization of these advantages depends on a reasonable array configuration design combination and the cooperation of beamforming algorithms. Improper design may lead to problems such as uneven beam widths, sidelobe spatial aliasing, and deteriorated white noise gain. Therefore, how to select an appropriate array configuration and reasonably allocate the weight coefficients of different microphones in combination with beamforming algorithms to obtain a beam with both higher white noise gain and directivity factor has always been a major challenge in beamforming optimization technology. In terms of beamforming algorithms, classical methods such as the delay-and-sum method or the minimum variance distortionless response method (MVDR) have limitations, such as the problem of inverse proportionality between white noise gain and directivity and the problem of uneven beam widths. In recent years, researchers have proposed more advanced beamforming techniques, among which the Jacobi-Anger functions are adopted. These functions are solutions of Bessel differential equations, which can more efficiently model the directional patterns of arrays and have the advantage of a constant beam width, especially suitable for accurately modeling concentric circular arrays. However, existing Jacobi-Anger functions also have limitations, such as being unable to cover the complete elevation angle range (0° to 90°) and being prone to noise amplification problems at lower frequencies, which limits their application in full-space beamforming.

[0061] Therefore, to solve the above problems, the present application proposes a beamforming method for a microphone combined array, aiming to achieve a high-performance microphone array with an adjustable constant beam width through a reasonable array configuration and an improved Jacobi-Anger function beamforming algorithm, realizing a beam output with both high white noise gain and high directivity factor, as well as improving the robustness of the beam, avoiding the low-frequency noise amplification problem, and enabling it to accurately cover the entire elevation angle range, meeting the application requirements of beamforming and sound source localization in three-dimensional space, so as to achieve stronger beam performance and better pick-up quality compared with traditional microphone arrays.

[0062] The technical solutions provided by the embodiments in the present application will be described in detail below with reference to the accompanying drawings.

[0063] Please refer to Figure 1 , in an embodiment of the present application, a microphone combination array beamforming method is provided, including the following steps:

[0064] S1. Determine the geometric distribution of the positions of several microphones in the circular microphone array according to the preset constraint conditions.

[0065] Furthermore, the preset constraint conditions include the upper limit of the number of microphones, the total number of concentric circles, the range of the number of microphones per circle, and the range of the radii of the concentric circles;

[0066] Among them, considering the actual application of the ceiling microphone array (for example, the side length of the ceiling square is about 0.6 m), in this embodiment, the maximum radius of the array is constrained to 0.3 m (meter), and the radius range of all circles is controlled within 0 - 0.3 m; considering the upper limit problem of the number of microphones for synchronous control, the maximum number of microphones is set to 128; considering the number of cooperative circles of the concentric circle array combination and the white noise gain problem, the total number of circles is set to 5 - 10 circles; except for the central microphone, the number of microphones per circle is 10 - 20.

[0067] Since circular arrays with different radii have different beam characteristics, these combinations will have different effects in the entire frequency range (0 - 8 kHz) or elevation angle range (0 - 90°). If the positions of the microphones and the array combination are not properly selected and optimized, the combined effect may have extremely poor beam responses in certain directions or certain frequency bands (for example, high-frequency spatial aliasing sidelobe bulges resulting in poor suppression effects of direction interference signals and low-frequency voiceprint distortion and decreased speech clarity, etc.), and ultimately, the expected perfect beamforming performance cannot be achieved in the entire frequency spectrum range. Therefore, it is also necessary to optimize the combination of the concentric circular microphone arrays.

[0068] S2. Perform multi-objective optimization on the geometric distribution of the positions of several microphones in the microphone array to obtain several groups of microphone arrays of the first combination that meet the constraint conditions, where each combination in the microphone array of the first combination includes a combination of the number of microphones and the radius of different concentric circles.

[0069] In the embodiment of the present application, the selection of the radius and number of the microphone array mainly involves the requirement of optimizing the beam response in the full frequency band. The purpose of multi-objective optimization is to improve the white noise gain and directivity of the combined beam (i.e., maximize the white noise gain and directivity factor), while reducing the spatial aliasing effect of the combined beam and minimizing the sidelobe level, and limiting the main beam width within the desired size range.

[0070] Based on the above preset constraint conditions, in the embodiments of the present application, a swarm intelligence optimization algorithm (such as a genetic algorithm or a particle swarm algorithm) is first used to perform multi-objective optimization on the microphone array. To improve the multi-objective optimization efficiency, the interval accuracy of the radius is set to 0.02 m. Therefore, 6 groups of microphone arrays of the first combination that meet the preset constraint conditions can be obtained. Among them, each combination in the microphone arrays of the first combination includes a combination of the number of microphones and the radius of different concentric circles.

[0071] S3. Iteratively optimize the microphone arrays of the first combination to obtain the microphone arrays of the second combination.

[0072] After obtaining 6 groups of microphone arrays of the first combination that meet the preset constraint conditions, a convex optimization algorithm is used to perform high-precision iterative optimization with a radius of 0.001 m for these 6 sub-optimal solution schemes (the microphone arrays of the second combination) respectively. Finally, the combined solution of the microphone arrays of the second combination with the highest average white noise gain and the best directivity factor performance is selected, which is the geometric distribution structure of the above microphone combined array.

[0073] Based on the above method to obtain the optimal combined configuration of the microphone array, please refer to Figure 2 For the geometric distribution design of the microphone combined array of the present application, the microphones are arranged in concentric circles. The inner circle corresponds to a smaller radius, and the outer circle corresponds to gradually increasing radii. The microphones in each circle are evenly arranged at a specific radius interval and show a regular increase from the inner circle to the outer circle. Among them, the optimal combined configuration of the microphone array obtained in the present application is: the radii (unit: m) from the inside to the outside are: 0, 0.021, 0.044, 0.065, 0.090, 0.130, 0.170, 0.255, and the corresponding numbers of microphones are: 1, 12, 19, 19, 19, 19, 19, 19, totaling 127.

[0074] S4. Use the microphone arrays of the second combination to collect the sound source signal, and obtain the actual azimuth information of the sound source signal through sound source localization.

[0075] S5. Generate the beam coefficient in the desired direction from the actual azimuth information of the sound source signal through an improved Jacobian function.

[0076] Since the Jacobi Anger function is a solution of the Bessel differential equation, it can not only model the directivity pattern of the array more efficiently, but also has the advantage of a constant beam width, especially suitable for accurately modeling concentric circular arrays. However, past research has shown that the existing Jacobi Anger function cannot cover the full elevation angle range (0° to 90°), and is prone to noise amplification problems at lower frequencies, which limits the application of this technology in full-space beamforming. Therefore, the embodiment of this application improves the Jacobi Anger function and reasonably allocates the beam coefficients of different microphones based on the improved Jacobi Anger function.

[0077] Furthermore, the actual azimuth information of the sound source signal is passed through the improved Jacobi Anger function to generate the beam coefficients in the desired direction, and the steps include:

[0078] S51. Define the order of the Bessel function, denoted as n, and the order of the Bessel function is used to represent the wave propagation characteristics in the circular microphone array configuration;

[0079] where n = [-N:N]; N is the order of the Bessel function; the order N of the Bessel function is selected according to the resolution of the required beamforming system; the index of n represents the radial mode of the circular microphone.

[0080] S52. Calculate the phase change of the sound wave when propagating in the microphone array for each microphone element in the microphone array at different frequencies, different positions, and different sound source elevation angles to obtain a phase matrix;

[0081] Specifically, calculating the phase change of the sound wave when propagating in the microphone array for each microphone element in the microphone array at different frequencies, different positions, and different sound source elevation angles includes:

[0082] S52.1. Obtain the sampling frequency F of the sound wave signal collected by the microphone array S ;

[0083] F S = 1 / Ts; where Ts is the sampling period.

[0084] S52.2. Based on the sampling frequency, calculate the maximum frequency fmax;

[0085] fmax = F S / 2;

[0086] S52.3. According to the ratio of the sampling frequency to the maximum frequency, obtain the scaling factor alpha;

[0087] alpha = F S / fmax;

[0088] S52.4. Calculate the shortest wavelength lambda when the sound source signal received by the microphone array is obtained based on the maximum frequency and the speed of sound min ;

[0089] lambda min = C / fmax; where C is the speed of sound in air.

[0090] S52.5. Obtain the normalized radius of each concentric ring in the microphone array;

[0091] S52.6. Set the number of Fourier transform points N fft , and based on the number of Fourier transform points N fft and the sampling frequency F S , calculate multiple frequency components f. For example, if N fft is 320, then the total number of divided frequency points is N fft / 2 + 1 = 161 frequency points. In the full frequency range of 0 - 8 kHz, each frequency point F is 0 Hz, 50 Hz, 100 Hz, 150 Hz,..., 8000 Hz in turn. Then the frequency component f = F / F S ;

[0092] S52.7. Calculate the phase change when the sound wave propagates in the microphone array according to the scaling factor, the shortest wavelength, the normalized radius, and the frequency component.

[0093] Furthermore, first define the normalized array radius r_hat of the p-th ring p , which is expressed as:

[0094] r_hat p = ring_radius / lambda min ;

[0095] where ring_radius is the radius of the p-th microphone ring;

[0096] For each frequency f, calculate the phase change eta when the sound wave propagates in the microphone array. The formula is:

[0097] eta = 2 * π * f * r_hat p * alpha * sin(theta_d * π / 180).

[0098] Among them, f represents multiple frequency components in the sound source signal received by the microphone array; theta_d is the elevation angle of the sound source, and its value range is 0 to 90°. The traditional planar circular array is exactly one subset (theta_d = 90°), and here it is extended to half of the space to meet the application mode of the ceiling microphone array.

[0099] In this application, the elevation angle extension is performed on the phase term eta of the Bessel function, so that the beamformer can effectively cover the entire elevation angle range of 0 to 90°.

[0100] S53. Based on the phase matrix, use the first kind of Bessel function to calculate the Bessel function values of each Bessel order at each frequency;

[0101] Calculate the Bessel function values of each Bessel order n at each frequency component f, which is completed by the first kind of Bessel function besselj, and the formula is:

[0102] bessel_val(idx_bessel_order,:) = besselj(n(idx_bessel_order), eta);

[0103] Among them, idx_bessel_order = 1:length(n) is the index of the Bessel order, and eta is the phase term calculated before.

[0104] S54. Based on the order of the Bessel function and the angles of the microphone array elements, calculate the phase offset of each microphone array element;

[0105] The phase offset is determined by the angles phi_p_m of the microphone array elements in the p-th ring, and the angles are given in the cell array phi_p_m; the phase offset psi_p is calculated as:

[0106] psi_p = exp(-i * n * sensors_angles);

[0107] Among them, n is the Bessel order, and sensors_angles is the angle array corresponding to the positions of the microphone array elements on the p-th ring.

[0108] S55. According to the phase offset of each microphone array element, construct a comprehensive phase offset matrix, and its formula is:

[0109] bar_psi_p(:,:,idx_f) = (bessel_val(:,idx_f) * ones(1,length(m))) * psi_p;

[0110] Among them, \(m = [-K_p + 1:K_p]\), and \(K_p = M_p / 2\) is half the number of microphone array elements \(M_p\) on the \(p\)-th ring; the matrix \(\bar{\psi}_p\) is the comprehensive matrix of the phase offset of the microphone array elements, and the three dimensions respectively represent the Bessel function values corresponding to the \(p\)-th ring, the number of microphone array elements, and the frequency; ":" represents all items included in this dimension in the current matrix; \(\bar{\psi}_p(:,:,idx_f)\) represents the two-dimensional matrix corresponding to the frequency \(f\) at the \(idx_f\) index; \(bessel\_val(:,idx_f)\) represents the one-dimensional matrix corresponding to the frequency \(f\) at the \(idx_f\) index.

[0111] Finally, by storing the Psi matrix of each individual ring, the total Psi matrix of all rings is assembled:

[0112] \(\bar{\Psi}\{p\} = \bar{\psi}_p\);

[0113] Among them, \(p\) is the index of each ring, and then by combining the matrices of all rings, the complete Psi matrix is constructed.

[0114] S56. Based on the comprehensive phase offset matrix, calculate the beam coefficient in the desired direction.

[0115] The formula for calculating the beam coefficient \(h_f\) in the desired direction is:

[0116] \(h_f=\bar{\Psi}_f^H*pinv(\bar{\Psi}_f*\bar{\Psi}_f^H)*conj(J)*conj(\Gamma)*b_{2N}\);

[0117] Among them, \(\bar{\Psi}_f=\bar{\Psi}(:,:,idx_f)\) is the Psi matrix corresponding to the frequency \(f\); \(J = diag((1i)^{-n})\) is the diagonal matrix of the phase offset caused by the Bessel order; \(\Gamma = diag(exp(-i*n*(\phi_d*\pi / 180)))\) is the steering vector along the azimuth angle \(\phi_d\); \(b_{2N}\) is a control coefficient with an adjustable constant beam width; \(\bar{\Psi}_f^H\) represents the conjugate transpose of \(\bar{\Psi}_f\).

[0118] The beam coefficient \(h_f\) is used to direct the microphone array to focus on the desired sound source while suppressing interference from other directions.

[0119] Specifically, based on the comprehensive phase offset matrix, calculating the beam coefficient in the desired direction includes:

[0120] S56.1. Define an identity matrix \(I\) with the same number of rows as the comprehensive phase offset matrix;

[0121] I = eye(size(bar_Psi_f, 1));

[0122] Among them, eye() creates an identity matrix of appropriate size.

[0123] S56.2. Multiply the combined phase shift matrix and the conjugate transpose of the combined phase shift matrix to obtain a new matrix bar_Psi_f1;

[0124] bar_Psi_f1 = bar_Psi_f * bar_Psi_f';

[0125] Among them, bar_Psi_f' is the conjugate transpose of the combined phase shift matrix.

[0126] S56.3. Based on the new matrix and the identity matrix scaled by the regularization factor, obtain the regularization matrix;

[0127] regularized_matrix = bar_Psi_f1' * bar_Psi_f1 + lambda * I;

[0128] Among them, lambda is a regularization factor. Its addition can solve the problem of low-frequency white noise amplification. The value range of lambda is 0.0001 to 1. According to the beam characteristics of different frequency points and different elevation angles, selecting an appropriate regularization factor can effectively control the balance between white noise gain and directivity. The closer the value of lambda is to 0.0001, the greater the directivity of the beam can be increased but the white noise gain is reduced at the same time. The closer the value of lambda is to 1, the greater the white noise gain of the beam can be increased but the directivity is reduced at the same time; bar_Psi_f1' represents the conjugate transpose of bar_Psi_f1.

[0129] S56.4. Perform a pseudo-inverse calculation on the regularization matrix to obtain the beam coefficient in the desired direction.

[0130] The beam coefficient h_f_regularized of the regularized beamformer is calculated by applying the pseudo-inverse of the regularization matrix:

[0131] h_f_regularized = bar_Psi_f' * pinv_bar_Psi_f * conj(J) * conj(Gamma) * b_{2N};

[0132] Among them, pinv_bar_Psi_f is the pseudo-inverse of the regularization matrix:

[0133] pinv_bar_Psi_f = pinv(regularized_matrix) * bar_Psi_f1';

[0134] Note that the additional bar_Psi_f1' term in pinv_bar_Psi_f here is also crucial to ensure the regularization takes effect.

[0135] This application makes a regularization improvement to the Psi matrix. For different pitch angles and different frequency points, by selecting an appropriate regularization factor, the contradiction between the beam white noise gain and the directivity can be better balanced, and the noise amplification phenomenon that occurs in the Bessel function at low frequencies is solved, thereby increasing the beam robustness of the Jacobi-Anger function.

[0136] S6. Weightedly sum the beam coefficients and the sound source signals of the collected microphone array to obtain a beam output.

[0137] This application proposes an improved Jacobi-Anger expression, which overcomes the limitations in traditional microphone array beamforming. By introducing a novel regularization technique, the problems of insufficient pitch angle coverage and low-frequency white noise amplification are effectively solved. Coupled with its inherent constant beamwidth characteristic, it can be more attractive than traditional beam schemes in beamforming and sound source localization applications in three-dimensional space; it can improve the beam robustness, avoid the low-frequency noise amplification problem, and enable it to accurately cover the entire pitch angle range (0 to 90°) to meet the application requirements of beamforming and sound source localization in three-dimensional space. It should be noted that the improved Jacobi-Anger expression in this application is not limited to the combined array of this application and is also practical in the system application of a single-loop array.

[0138] In summary, this application is based on the optimal combination design of the beamforming technology based on the improved Jacobi-Anger function and the uniform concentric circular array. This scheme has powerful beam performance and supports full-space beamforming, with an azimuth angle range from 0 to 360° and a pitch angle range from 0 to 90°. It is very suitable for the sound pickup and localization applications of ceiling microphone arrays. Even in an environment with a low signal-to-noise ratio, it can provide reliable noise suppression capabilities and excellent sound pickup quality. The main performance and technical effects of this application scheme are as follows: 1) High directivity factor: The average sidelobe suppression depth is achieved at 25 dB, and the maximum suppression depth exceeds 40 dB. This high level of sidelobe suppression can effectively reduce directional interference noise, minimize vocal reverberation, and greatly enhance the clarity of the target sound source. 2) High white noise gain: The white noise gain improvement amount in different frequency bands is achieved at 6 to 20 dB. This enhancement makes the vocal performance clearer in a low signal-to-noise ratio environment and improves the speech intelligibility in a complex acoustic environment. 3) Constant beamwidth: The beamwidth remains nearly consistent throughout the entire frequency range, and the beamwidth can be freely adjusted between 20° and 60°. This flexibility allows the beam to be configured according to specific application requirements, ensuring that this beamforming solution can adapt to various acoustic environments and usage scenarios.

[0139] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously.

[0140] Based on the same idea as a microphone combined array beamforming method in the above embodiments, the present application also provides a microphone combined array beamforming system, which can be used to execute the above microphone combined array beamforming method. For the sake of convenience of description, in the structural schematic diagram of an embodiment of a microphone combined array beamforming system, only the parts related to the embodiments of the present application are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the system, and it may include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.

[0141] Please refer to Figure 3 , in another embodiment of the present application, a microphone combined array beamforming system is provided. The system includes a determining array distribution module 101, a target optimization module 102, an iterative optimization module 103, an obtaining azimuth information module 104, a generating beam coefficient module 105, and a beam output module 106;

[0142] The determining array distribution module 101 is configured to determine the geometric distribution of the positions of several microphones in the circular microphone array according to preset constraint conditions;

[0143] The target optimization module 102 is configured to perform multi-objective optimization on the geometric distribution of the positions of several microphones in the microphone array to obtain several groups of microphone arrays of a first combination that meet the constraint conditions, where each combination in the microphone arrays of the first combination includes a combination of the number of microphones and the radius of different concentric circular rings;

[0144] The iterative optimization module 103 is configured to perform iterative optimization on the microphone arrays of the first combination to obtain microphone arrays of a second combination;

[0145] The obtaining azimuth information module 104 is configured to collect a sound source signal by using the microphone arrays of the second combination and obtain the actual azimuth information of the sound source signal through sound source localization;

[0146] The generating beam coefficient module 105 is configured to generate beam coefficients in the desired direction by using the actual azimuth information of the sound source signal through an improved Jacobi-Anger function;

[0147] The beam output module 106 is configured to perform weighted summation on the beam coefficients and the collected microphone array sound source signals to obtain a beam output.

[0148] It should be noted that a microphone combination array beamforming system of the present application corresponds one-to-one with a microphone combination array beamforming method of the present application. The technical features and their beneficial effects described in the embodiments of the above-mentioned microphone combination array beamforming method are applicable to the embodiments of a microphone combination array beamforming system. For specific content, reference can be made to the description in the method embodiments of the present application, which will not be elaborated here. This is hereby declared.

[0149] In addition, in the implementation manner of the microphone combination array beamforming system in the above embodiment, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the microphone combination array beamforming system is divided into different program modules to complete all or part of the functions described above.

[0150] In another embodiment, an electronic device for implementing a microphone combination array beamforming method is provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, a microphone combination array beamforming method of any embodiment of the present application is implemented.

[0151] Exemplarily, in this embodiment, the computer program can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present application. The one or more module elements can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the device.

[0152] The device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include, but is not limited to, a processor and a memory.

[0153] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the device, and connects various parts of the entire device using various interfaces and lines.

[0154] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0155] Correspondingly, the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute a microphone combination array beamforming method according to any one of the above embodiments.

[0156] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0158] The above embodiments are preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present application should be equivalent replacement methods and are all included in the protection scope of the present application.

Claims

1. A microphone combination array beamforming method, characterized in that: The steps include: Determining the geometric distribution of a plurality of microphone positions in the circular microphone array according to preset constraints; Performing multi-objective optimization on the geometric distribution of the positions of a plurality of microphones in the microphone array to obtain a plurality of first combination microphone arrays satisfying the constraint conditions, wherein each combination in the first combination microphone arrays includes a combination of the number and radius of microphones in different concentric circular rings; Iteratively optimizing the microphone array of the first combination to obtain a microphone array of the second combination; Using the microphone array of the second combination to collect sound source signals, and obtaining actual position information of the sound source signals through sound source localization; The actual azimuth information of the sound source signal is passed through an improved Jacobian function to generate a beam coefficient in a desired direction; The beam coefficient is weighted and summed with the collected microphone array sound source signal to obtain a beam output.

2. The microphone combination array beamforming method according to claim 1, characterized in that: The preset constraints include an upper limit on the number of microphones, a total number of concentric circles, a range of the number of microphones in each circle, and a range of the radius of the concentric circles.

3. The microphone combination array beamforming method according to claim 1, characterized in that: The multi-objective optimization includes maximizing the white noise gain and directivity factor of the microphone array beam, minimizing the spatial aliasing effect of the microphone array beam and minimizing the sidelobe level, and limiting the main beamwidth of the microphone array to a desired range.

4. The microphone combination array beamforming method according to claim 1, characterized in that: The combination of the number and radius of microphones in different concentric circles included in the microphone array of the second combination is: The radii of different concentric circles from inside to outside are: 0m, 0.021m, 0.044m, 0.065m, 0.090m, 0.130m, 0.170m, 0.255m; Among them, the number of microphones corresponding to the concentric circle ring with a radius of 0 meter is 1, the number of microphones corresponding to the concentric circle ring with a radius of 0.021 meter is 12, and the number of microphones corresponding to the concentric circle rings with radii of 0.044 meter, 0.065 meter, 0.090 meter, 0.130 meter, 0.170 meter and 0.255 meter are all 19.

5. The microphone combination array beamforming method according to claim 1, characterized in that: The actual azimuth information of the sound source signal is passed through an improved Jacobian function to generate a beam coefficient in a desired direction, comprising: Defining the order of a Bessel function, wherein the order of the Bessel function is used to represent wave propagation characteristics in a donut-shaped microphone array configuration; Calculate the phase change of each microphone element in the microphone array when the sound wave propagates in the microphone array at different frequencies, different positions and different sound source pitch angles to obtain a phase matrix; Based on the phase matrix, using Bessel first kind function to calculate the Bessel function value of each Bessel order at each frequency; Based on the order of the Bessel function and the angle of the microphone array element, a phase shift of each microphone array element is calculated; According to the phase shift of each microphone array element, a comprehensive phase shift matrix is ​​constructed; Based on the integrated phase offset matrix, a beam coefficient in a desired direction is calculated.

6. The microphone combination array beamforming method according to claim 5, characterized in that: The value range of the sound source pitch angle is 0 to 90 degrees.

7. The microphone combination array beamforming method according to claim 5, characterized in that: The calculating of the phase change of each microphone array element in the microphone array when the sound wave propagates in the microphone array at different frequencies, different positions and different sound source pitch angles includes: Obtain the sampling frequency of the microphone array to collect sound wave signals; Based on the sampling frequency, a maximum frequency is calculated; Obtaining a scaling factor according to a ratio of the sampling frequency to the maximum frequency; Based on the maximum frequency and the speed of sound, the shortest wavelength of the sound source signal received by the microphone array is calculated; Get the normalized radius of each concentric ring in the microphone array; Setting the number of Fourier transform points, and calculating a plurality of frequency components based on the number of Fourier transform points and the sampling frequency; The phase change of the sound wave when propagating in the microphone array is calculated according to the scaling factor, the shortest wavelength, the normalized radius and the frequency component.

8. The microphone combination array beamforming method according to claim 5, characterized in that: The step of calculating a beam coefficient in a desired direction based on the integrated phase offset matrix includes: defining an identity matrix having the same number of rows as the integrated phase offset matrix; Multiplying the integrated phase offset matrix and the conjugate transpose of the integrated phase offset matrix to obtain a new matrix; Obtaining a regularized matrix based on the new matrix and the identity matrix scaled by the regularization factor; A pseudo-inverse calculation is performed on the regularization matrix to obtain the beam coefficient in the desired direction.

9. A microphone array beamforming system, characterized in that: A microphone combination array beamforming method applied to any one of claims 1-8, comprising an array distribution determination module, a target optimization module, an iterative optimization module, an orientation information acquisition module, a beam coefficient generation module, and a beam output module; The array distribution determination module is used to determine the geometric distribution of a plurality of microphone positions in the circular microphone array according to preset constraints; The target optimization module is used to perform multi-objective optimization on the geometric distribution of the positions of a plurality of microphones in the microphone array to obtain a plurality of first combination microphone arrays satisfying the constraint conditions, wherein each combination of the first combination microphone arrays includes a combination of the number and radius of microphones of different concentric circular rings; The iterative optimization module is used to iteratively optimize the microphone array of the first combination to obtain the microphone array of the second combination; The module for obtaining position information is used to collect sound source signals by using the microphone array of the second combination, and obtain actual position information of the sound source signals by sound source localization; The module for generating beam coefficients is used to generate beam coefficients in a desired direction by using an improved Jacobian function to obtain actual azimuth information of the sound source signal; The beam output module is used to perform weighted summation on the beam coefficient and the collected microphone array sound source signal to obtain a beam output.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute a microphone combination array beamforming method as described in any one of claims 1-8.