Robust beamforming method, device, equipment and medium of a microphone array

By combining short-time Fourier transform and robust super-directional beamformer in the microphone array, the problem of limited applicability of microphone array beamforming methods is solved, and robust beamforming of microphone arrays with arbitrary structures is realized, improving beam steering capability and noise suppression effect.

CN119603606BActive Publication Date: 2025-12-05WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411725964.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-12-05
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Existing microphone array beamforming methods have limited applicability and cannot be effectively applied to microphone arrays with arbitrary planar geometry, resulting in limited beam steering capability and low-frequency white noise amplification problems.

Method used

A microphone array based on arbitrary planar geometry is used. The time-domain signal is converted into a frequency-domain signal through short-time Fourier transform technology. A robust super-directional beamformer is used for spatial filtering. An optimization problem is constructed and solved by the Lagrange multiplier method to constrain the white noise gain and form a robust beamformer.

Benefits of technology

It achieves consistent performance for microphone arrays with arbitrary structures, effectively suppresses white noise gain within a reasonable range, and improves the robustness and flexibility of beamforming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119603606B_ABST
    Figure CN119603606B_ABST
Patent Text Reader

Abstract

The application provides a robust beamforming method, device and equipment of a microphone array and a medium, wherein the robust beamforming method of the microphone array comprises: setting a microphone array and acquiring audio information of the microphone array; converting time domain signals of the microphone array into frequency domain signals through a short-time Fourier transform technology; determining an optimization problem by using the microphone array, and solving a robust super directivity beamformer for the optimization problem; performing spatial domain filtering on the frequency domain signals of the microphone array through the robust super directivity beamformer to obtain frequency domain output signals, and inversely transforming the frequency domain output signals into time domain output signals. Through the application, the problem that the array structure range applicable to the beamforming method in the prior art is limited is solved, and the ability of flexibly controlling the beam steering direction and the robustness is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech signal processing, and particularly relates to a robust beamforming method, device and equipment of a microphone array and a medium. BACKGROUND

[0002] Microphone array beamforming technology is widely used in sensor array systems to improve the quality of target speech by suppressing interference and noise. Many algorithms have been proposed in this field, including delay-and-sum beamforming, superdirective beamforming and differential beamforming. Among them, differential beamforming is particularly attractive due to its ability to provide high directional gain and broadband frequency-invariant beam patterns, making it very suitable for small-size microphone array applications to achieve high-fidelity speech signal acquisition. For example, in the application of voice assistants, remote meetings and smart speakers, microphone array technology uses multiple microphones to achieve long-distance sound pickup and reduce background noise and echo, providing clearer and more accurate speech signals.

[0003] However, differential microphone arrays currently face two major challenges: 1. Achieving arbitrary angle pointing within the plane; 2. Reducing or avoiding low-frequency white noise amplification. For challenge one, current research has shown that ring or concentric ring microphone arrays can achieve arbitrary angle pointing within the plane, but in practical applications, irregular array structures are often encountered. Therefore, the design of a robust differential beamformer for arbitrary geometric microphone arrays in the plane needs to be addressed.

[0004] There is currently no effective solution to the problem of limited application scope of existing beamforming methods. SUMMARY

[0005] The present application provides a robust beamforming method, device and equipment of a microphone array to solve the defect of limited application scope of the beamforming method in the prior art, and realizes the ability to flexibly control the beam steering direction and robustness.

[0006] Compared with a single microphone, the advantage of a microphone array lies in its multi-channel configuration, which provides more abundant information and higher flexibility for implementing various acoustic signal processing tasks. The sound signals received by microphones at different positions include redundancy, which helps to improve the robustness of the signals. By processing the received multiple signals, noise can be more effectively suppressed, and sound source estimation can be calculated to achieve certain goals, such as speech noise reduction / enhancement, dereverberation, sound source localization and tracking. Moreover, the microphone array can be communicatively coupled to a processing device (such as a digital signal processor (DSP) or a central processing unit (CPU)), which contains circuitry that can be programmed to implement a beamformer to calculate the sound source estimation.

[0007] Beamformer is a kind of spatial filter, which mainly forms a beam of a specific direction by weighting and phase adjustment on the sound signals received by a multi-channel microphone array, which can selectively enhance or suppress sound signals in space, so that the system can more effectively capture sound sources in a specific direction. However, the existing array structure is generally a linear array, a ring array and other regular arrays, and the corresponding beamforming method cannot be simply extended to a microphone array with arbitrary geometry. When facing a microphone array with irregular distribution, the performance will be severely degraded, and the beam steering capability will also be limited. To solve this practical problem, the application develops a differential beamforming method based on a microphone array with arbitrary geometry on a plane, and designs it robustly.

[0008] In a first aspect, the application provides a robust beamforming method for a microphone array, comprising:

[0009] Setting up a microphone array and obtaining audio information of the microphone array;

[0010] Converting time-domain signals of the microphone array into frequency-domain signals by using short-time Fourier transform technology;

[0011] Using the microphone array, determining an optimization problem, and solving a robust superdirectivity beamformer for the optimization problem;

[0012] Performing spatial filtering on the frequency-domain signals of the microphone array by using the robust superdirectivity beamformer to obtain frequency-domain output signals, and inversely transforming the frequency-domain output signals into time-domain output signals.

[0013] According to the robust beamforming method for a microphone array provided by the application, the time-domain signals of the microphone array are converted into frequency-domain signals by using short-time Fourier transform technology, comprising:

[0014] Given a multi-channel input signal, the multi-channel input signal is converted from a time-domain signal into a frequency-domain signal by using short-time Fourier transform technology, and the size of the frequency-domain signal is determined.

[0015] According to the robust beamforming method for a microphone array provided by the application, the microphone array is used to determine an optimization problem, comprising:

[0016] From the perspective of minimum beam mean square error, the white noise gain is constrained to satisfy the minimum threshold, and the desired direction signal is distortionless, the initial optimization problem is determined, and the initial beamformer for the initial optimization problem is designed.

[0017] The initial beamformer is decomposed into a form of sum of a delay-and-sum beamformer and a reduced-rank beamformer, and is brought into the initial optimization problem to generate a transformed optimization problem.

[0018] According to the robust beamforming method of the microphone array provided by the application, the initial optimization problem is to minimize the error between the ideal beam pattern and the actual beam pattern while constraining the white noise gain in the ideal range, and the signal in the desired direction can be transmitted without distortion.

[0019] According to the robust beamforming method of the microphone array provided by the application, the robust super-directive beamformer is generated for the optimization problem, which comprises:

[0020] The cost function is constructed based on the optimization problem by the Lagrange multiplier method, and the optimization problem is transformed into a standard form;

[0021] Based on the standard form of the optimization problem, the closed-form solution of the robust super-directive beamformer is obtained.

[0022] According to the robust beamforming method of the microphone array provided by the application, the frequency domain signal of the microphone array is spatially filtered by the robust super-directive beamformer to obtain a frequency domain output signal, which comprises:

[0023] For each frame of frequency domain signal of the microphone array, for each frame of frequency domain signal, each frequency point is traversed, and the target signal is intercepted;

[0024] The target signal is weighted and filtered by the robust super-directive beamformer, and the corresponding frequency domain output signal is obtained, until the weighted filtering operation is completed for each frequency point of each frame.

[0025] According to the robust beamforming method of the microphone array provided by the application, the frequency domain output signal is inversely transformed into a time domain output signal, which comprises:

[0026] The frequency domain output signal is converted into a time domain output signal by inverse short-time Fourier transform technology.

[0027] In a second aspect, the application also provides a robust beamforming device of a microphone array, characterized in that it comprises:

[0028] The acquisition module is used for setting a microphone array and acquiring audio information of the microphone array;

[0029] The conversion module is used for converting the time domain signal of the microphone array into a frequency domain signal by short-time Fourier transform technology;

[0030] an optimization module configured to determine an optimization problem using the microphone array and to solve a robust superdirectivity beamformer for the optimization problem;

[0031] a processing module configured to spatially filter a frequency domain signal of the microphone array by the robust superdirectivity beamformer to obtain a frequency domain output signal and to inverse transform the frequency domain output signal to a time domain output signal.

[0032] In a third aspect, the present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable in the processor, wherein the processor implements the robust beamforming method of the microphone array according to the first aspect when executing the program.

[0033] In a fourth aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the robust beamforming method of the microphone array according to the first aspect.

[0034] In a fifth aspect, the present application also provides a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the robust beamforming method of the microphone array according to the first aspect.

[0035] Compared with the prior art, the present application has the following beneficial effects:

[0036] The robust beamforming method of the microphone array provided by the present application can approximate an ideal beam pattern by using Jacobi-Anger expansion to achieve consistent performance for microphone arrays of any structure, while constraining white noise gain within a reasonable range, thereby constructing an optimization problem, and then converting the optimization problem into a quadratic eigenvalue problem to construct a robust superdirectivity beamformer, and performing spatial filtering through the robust superdirectivity beamformer, thereby solving the problem that the existing beamforming method has a limited scope of application. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0038] Figure 1 is a flowchart of the robust beamforming method of the microphone array provided by the present application;

[0039] Figure 2Fig. 1 is a schematic diagram of a planar arbitrary geometry microphone array in an embodiment of the present application;

[0040] Figure 3 Fig. 2 is a schematic diagram of the variation of the directivity factor DF and the white noise gain WNG of a PDMA (planar differential microphone array) as a function of frequency in an embodiment of the present application;

[0041] Figure 4 Fig. 3 is a schematic diagram of the variation of the directivity factor DF and the white noise gain WNG as a function of frequency under different spread orders of a PDMA in an embodiment of the present application;

[0042] Figure 5 Fig. 4 is a beam pattern of a PDMA pointing at -30° in an embodiment of the present application;

[0043] Figure 6 Fig. 5 is a beam pattern of a PDMA pointing at -120° in an embodiment of the present application;

[0044] Figure 7 Fig. 6 is a beam pattern of a PDMA pointing at 120° in an embodiment of the present application;

[0045] Figure 8 Fig. 7 is a beam pattern of a PDMA pointing at 30° in an embodiment of the present application;

[0046] Figure 9 Fig. 8 is a structure block diagram of a robust beamforming device of a microphone array provided by the present application;

[0047] Figure 10 Fig. 9 is a structure schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0049] The performance of a microphone array can be reflected by a beam pattern, and a beam pattern Figure 1An aspect can be quantified by a directivity factor (DF), which is the ability of a beam pattern to maximize its sensitivity in an observation direction over its average sensitivity in all directions. The observation direction is the angle of incidence of the sound signal with the maximum sensitivity. However, the microphone array is very sensitive to the noise generated by the hardware elements of each microphone itself. This effect is called white noise gain (WNG). The design of a microphone array beamformer can focus on finding the optimal beamshaping filter with a given array geometry under certain conditions. However, this is usually a traditional method designed based on regular microphone arrays, and for irregular planar arrays with arbitrary geometry, the beamformer is designed by first expanding the Jacobi-Anger series, and then approximating the actual beam to an ideal beam pattern. However, the beamformer designed in this way has the problem of poor robustness, which is manifested in the serious amplification of white noise, especially in the low frequency band.

[0050] Therefore, the present application provides a robust beamforming method for a microphone array based on a planar differential array (PDMA), Figure 1 is a flowchart of the robust beamforming method for a microphone array provided by the present application, as Figure 1 shown, the method comprises the following steps:

[0051] Step S101, setting up a microphone array and obtaining audio information of the microphone array.

[0052] Step S102, converting the time-domain signal of the microphone array into a frequency-domain signal by a short-time Fourier transform technique.

[0053] Step S103, determining an optimization problem using the microphone array, and solving a robust superdirectivity beamformer for the optimization problem.

[0054] Step S104, spatially filtering the frequency-domain signal of the microphone array by the robust superdirectivity beamformer to obtain a frequency-domain output signal, and inversely transforming the frequency-domain output signal into a time-domain output signal.

[0055] In the method, first, in an actual acoustic scene (such as a classroom, a conference room, etc.), a designed differential microphone array with arbitrary geometry is placed on a table, and audio information is recorded. Then, the time-domain signal of the microphone array is converted into a frequency-domain signal by a short-time Fourier transform (STFT) technique, the multi-channel number is denoted as , the frequency point number is denoted as K , and the frame length is denoted asN , the size of the frequency domain signal is Then, for the above designed microphone array and the final optimization target, an optimization problem is determined, and a corresponding robust superdirectivity beamformer is constructed. For example, in order to reduce the influence of white noise in the beamforming process, the optimization problem is to control the white noise gain within a reasonable range, and correspondingly, the constructed robust superdirectivity beamformer can control the influence of white noise gain. Finally, the frequency domain output signal is obtained by spatial filtering the frequency domain signal of the microphone array through the robust superdirectivity beamformer, and the time domain output signal is obtained by inverse transforming the frequency domain output signal. In the above process, the ideal beam pattern is approximated by using Jacobi-Anger expansion to achieve consistent performance for microphone arrays of any structure, while the white noise gain is constrained within a reasonable range to construct an optimization problem, and then the optimization problem is converted into a quadratic eigenvalue problem to construct a robust superdirectivity beamformer, and spatial filtering is performed through the robust superdirectivity beamformer, solving the problem that the existing beamforming method is limited in scope.

[0056] In some embodiments, the time domain signal of the microphone array is converted into a frequency domain signal by a short-time Fourier transform technique, including: given a multi-channel input signal, converting the multi-channel input signal from a time domain signal to a frequency domain signal using a short-time Fourier transform technique, and determining the size of the frequency domain signal.

[0057] An exemplary Figure 2 is a schematic diagram of a planar microphone array of any geometric structure in the embodiments of the present application, as Figure 2 shown, for the convenience of description and without limitation, in the plane, the positions of the subsequent microphones are measured with the origin as the reference point. Therefore, the position of the microphone sensor in space can be described according to the distance of the first microphone sensor to the reference position , and the included angle of the first microphone relative to the reference point. The direction of the source signal to the array can be parameterized by the azimuth angle . The steering vector represents the relative phase shift of the incident far-field wave form in the array. As described above, the steering vector of the array can be defined as:

[0058]

[0059] wherein is an imaginary unit, , is an angular frequency, is a time frequency, is the first ​the distance of the microphone to the origin, is the azimuth angle of the th microphone. is the speed of sound in air, usually assumed to be 340 m / s.

[0060] The developed array has flexible steering properties in space, so it can be assumed that the sound source is transmitted from any azimuth to the array. Therefore, the observed signal vector of the array can be expressed as:

[0061] ,

[0062] where, represents the received signal of the m th microphone, represents the propagation vector of the signal, is the relevant source signal, is the noise signal vector, which is defined similarly to . The output of the observed signal through the beamformer can be expressed as:

[0063]

[0064] where is the estimation of the desired signal, and the superscript H represents the conjugate transpose of a matrix or a vector,

[0065] ,

[0066] is the beamformer with length , which represents the spatial filter coefficients of the M th microphone.

[0067] The post-processor can convert the estimated value of (for each of a plurality of frequency bands) to the time domain to provide an estimated sound source represented as .

[0068] This PDMA can be associated with a beam pattern that reflects the sensitivity of the beamformer to a plane wave incident on the array from a specific angular direction . In the above PDMA, the beam pattern of the plane wave incident from the angle can be defined as:

[0069]

[0070] where, represents the beam pattern, which reflects the filter's ability to enhance or suppress signals coming from different directions in space. DF represents the beamformer's ability to suppress spatial noise coming from directions other than the observation direction, which, as mentioned above, can be written for the PDMA as follows:

[0071]

[0072] where, is the pseudo-coherent matrix of noise signals in the diffuse noise field (of size ). The (i, j)th element of i , j can be expressed as: , and is the distance between the i th microphone sensor and the j th microphone sensor. i j WNG is used to evaluate the beamformer's sensitivity to some of its own imperfections (e.g., noise from its own hardware). As mentioned above, the WNG associated with the PDMA can be written as:

[0073]

[0074]

[0075] To simplify the expression, in the following we omit the angular frequency .

[0076] In some embodiments among them, using a microphone array, an optimization problem is determined, including: for the microphone array, from the perspective of minimizing the beam mean square error, constraints are added that the white noise gain meets a minimum threshold, and the desired direction signal is distortionless, an initial optimization problem is determined, and an initial beamformer is designed for the initial optimization problem; the initial beamformer is decomposed into the form of a sum of a delay-and-sum beamformer and a reduced-rank beamformer, and is brought into the initial optimization problem to generate a transformed optimization problem.

[0077] Specifically, the initial optimization problem is to minimize the error between the ideal beam pattern and the actual beam pattern while constraining the white noise gain within the ideal range and ensuring that the desired direction signal is transmitted without distortion.

[0078] The specific design is as follows. First, based on the ideal beam pattern, then the steering vector is expanded in Jacobi-Anger series, and then the beam pattern of the beamformer is approximated to the ideal beam pattern while constraining the white noise gain to be not lower than the minimum threshold , and ensuring that the desired direction signal is transmitted without distortion, an optimization problem is constructed. Based on the ideal beam pattern, we can get:

[0079] ​​

[0080] where:

[0081]

[0082]

[0083]

[0084] The actual beam pattern can be expanded in Jacobi-Anger series as:

[0085]

[0086] where:

[0087]

[0088] .

[0089] From the perspective of minimum mean square beam error, the constraint of minimum error between the desired beam and the actual beam, the robustness constraint and the constraint of no distortion in the desired direction are added to the optimization process, and the following optimization problem can be obtained:

[0090]

[0091] Exemplarily, in order to solve the initial optimization problem, we express the sum of a delay-and-sum beamformer and a reduced-rank beamformer, that is:

[0092] ,

[0093] where, is the delay-and-sum beamformer, the matrix size is is the null space of is a filter with a length of and .

[0094] Then, the above formula is substituted into the optimization problem, and the initial optimization problem can be rewritten as:

[0095]

[0096] where:

[0097]

[0098] ​On this basis, the robust superdirectivity beamformer is generated for the optimization problem, including: constructing a cost function based on the optimization problem by using the Lagrange multiplier method, and converting the optimization problem into a standard form; based on the standard form of the optimization problem, a closed-form solution of the robust superdirectivity beamformer is obtained.

[0099] Exemplarily, a cost function is constructed by using the Lagrange multiplier method:

[0100]

[0101] Wherein, is the Lagrange multiplier. Then the cost function can be derived, and then the optimization problem is converted into a quadratic eigenvalue problem (QEP) problem in a standard form, and finally a closed-form solution of the robust beamformer for any PDMA is obtained.

[0102] Figure 3 is a schematic diagram of the variation curve of the directivity factor DF and the white noise gain WNG of the PDMA as a function of frequency in the embodiment of the application. As can be seen from the curve, the beamformer realizes the constraint on the WNG value. Figure 4 is a curve diagram of the directivity factor DF and the white noise gain WNG as a function of frequency under the condition of different expansion orders of the PDMA in the embodiment of the application. As can be seen from the curve, all the beamformers satisfy the preset WNG constraint at different expansion orders. Figures 5-8 is a beam pattern of the PDMA at different pointing angles in the embodiment of the application. As can be seen from the diagram, the beamformer can realize the consistency of the beam pattern at any angle.

[0103] In some embodiments, the frequency domain signal of the microphone array is spatially filtered by the robust superdirectivity beamformer to obtain a frequency domain output signal, including: traversing each frame of the frequency domain signal of the microphone array, for each frame of the frequency domain signal, traversing each frequency point, and intercepting a target signal; the target signal is weighted and filtered by using the robust superdirectivity beamformer, and a corresponding frequency domain output signal is obtained, until the weighted filtering operation is completed for each frequency point of each frame. On this basis, the frequency domain output signal is inversely transformed into a time domain output signal, including: converting the frequency domain output signal into a time domain output signal by using an inverse short-time Fourier transform technology.

[0104] Exemplarily, from to traversing each frame of the frequency domain signal, when the current frame is executed, from to traverse each frequency point of the converted domain signal in the current frame, and intercept a size of the signal is weighted and filtered by the obtained robust super-directivity beamformer weighting vector, and the filtered frequency domain output signal is stored; the above process is repeated until the conversion frequency point after which the current round of iteration ends; the above process is repeated until the frame number after which the current round of iteration ends. Finally, the frequency domain output signal is inversely transformed into a time domain output signal by using inverse STFT technology.

[0105] In summary, the present application proposes a robust beamforming method based on a differential microphone array with an arbitrary geometric structure on a plane, which minimizes the error between an ideal beam pattern and an actual beam pattern while constraining the white noise gain at an ideal level and the expected direction without distortion transmission, thereby establishing an optimization problem. By constructing a standard form QEP, the optimization problem is solved to obtain a robust differential beamformer.

[0106] The present application also provides a robust beamforming device for a microphone array. The robust beamforming device for a microphone array provided by the present application is described below, and the robust beamforming device for a microphone array described below can be mutually corresponding with reference to the robust beamforming method for a microphone array described above. Figure 9 is a structural block diagram of the robust beamforming device for a microphone array provided by the present application, as Figure 9 shown, the device comprises:

[0107] The acquisition module 901 is configured to set a microphone array and acquire audio information of the microphone array.

[0108] The conversion module 902 is configured to convert a time domain signal of the microphone array into a frequency domain signal by using a short-time Fourier transform technology.

[0109] The optimization module 903 is configured to determine an optimization problem by using the microphone array, and solve a robust super-directivity beamformer for the optimization problem.

[0110] The processing module 904 is configured to perform spatial domain filtering on the frequency domain signal of the microphone array by using the robust super-directivity beamformer to obtain a frequency domain output signal, and inversely transform the frequency domain output signal into a time domain output signal.

[0111] In use, the acquisition module 901 places the designed differential microphone array with an arbitrary geometric structure on a table in an actual acoustic scene (such as a classroom, a conference room, etc.), and records the audio information. Then, the conversion module 902 converts the time domain signal of the microphone array into a frequency domain signal by using a short-time Fourier transform (STFT) technology, the number of channels is denoted as , the number of frequency points is denoted as K , and the frame length is denoted asN , then the size of the frequency domain signal is Then, the optimization module 903 determines an optimization problem for the designed microphone array and the final optimization target, and constructs a corresponding robust superdirectivity beamformer. For example, in order to reduce the influence of white noise amplification in the beamforming process, the optimization problem is to control the white noise gain within a reasonable range, and correspondingly, the constructed robust superdirectivity beamformer can control the value of the white noise gain. Finally, the processing module 904 performs spatial filtering on the frequency domain signal of the microphone array through the robust superdirectivity beamformer to obtain a frequency domain output signal, and inversely transforms the frequency domain output signal into a time domain output signal. In the above process, the ideal beam pattern is approximated by using the Jacobi-Anger expansion to achieve consistent performance for microphone arrays of any structure, while the white noise gain is constrained within a reasonable range to construct an optimization problem, and then the optimization problem is converted into a quadratic eigenvalue problem to construct a robust superdirectivity beamformer, and spatial filtering is performed through the robust superdirectivity beamformer, solving the problem that the existing beamforming method is limited in scope.

[0112] Figure 10 An example of a schematic diagram of the physical structure of an electronic device is shown in Figure 10 The electronic device can include a processor 1001, a communications interface 1002, a memory 1003, and a communications bus 1004, wherein the processor 1001, the communications interface 1002, and the memory 1003 communicate with each other through the communications bus 1004. The processor 1001 can invoke logical instructions in the memory 1003 to execute a robust beamforming method for a microphone array, the method comprising:

[0113] setting up a microphone array and obtaining audio information of the microphone array;

[0114] converting time domain signals of the microphone array into frequency domain signals through a short-time Fourier transform technique;

[0115] determining an optimization problem using the microphone array, and solving a robust superdirectivity beamformer for the optimization problem;

[0116] performing spatial filtering on the frequency domain signals of the microphone array through the robust superdirectivity beamformer to obtain a frequency domain output signal, and inversely transforming the frequency domain output signal into a time domain output signal.

[0117] Further, the logic instructions in the memory 1003 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods according to the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0118] In another aspect, the present application also provides a computer program product, the computer program product comprising a computer program, the computer program being stored in a non-transitory computer readable storage medium, and the computer program being executable by a processor to cause a computer to execute the robust beamforming method of a microphone array provided by the above-mentioned methods, the method comprising:

[0119] setting a microphone array and obtaining audio information of the microphone array;

[0120] converting time-domain signals of the microphone array into frequency-domain signals by using a short-time Fourier transform technology;

[0121] determining an optimization problem by using the microphone array, and solving a robust superdirectivity beamformer for the optimization problem;

[0122] performing spatial filtering on the frequency-domain signals of the microphone array by using the robust superdirectivity beamformer to obtain frequency-domain output signals, and inversely transforming the frequency-domain output signals into time-domain output signals.

[0123] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the robust beamforming method of a microphone array provided by the above-mentioned methods, the method comprising:

[0124] setting a microphone array and obtaining audio information of the microphone array;

[0125] converting time-domain signals of the microphone array into frequency-domain signals by using a short-time Fourier transform technology;

[0126] determining an optimization problem by using the microphone array, and solving a robust superdirectivity beamformer for the optimization problem;

[0127] The frequency domain output signal is obtained by spatially filtering the frequency domain signals of the microphone array through a robust super-directive beamformer, and the frequency domain output signal is inversely transformed into a time domain output signal.

[0128] The apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0129] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0130] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A robust beamforming method of a microphone array, characterized in that, The method comprises: setting up a microphone array and obtaining audio information of the microphone array; converting time-domain signals of the microphone array into frequency-domain signals by using a short-time Fourier transform technique; determining an optimization problem by using the microphone array, and solving a robust super-directivity beamformer for the optimization problem; performing spatial filtering on the frequency-domain signals of the microphone array by using the robust super-directivity beamformer to obtain frequency-domain output signals, and inversely transforming the frequency-domain output signals into time-domain output signals; determining an optimization problem by using the microphone array, comprising: from the perspective of minimum beam mean square error, constraining the white noise gain to satisfy a minimum threshold and the desired direction signal to be distortionless for the microphone array, determining an initial optimization problem, and designing an initial beamformer for the initial optimization problem; decomposing the initial beamformer into a form of a sum of a delay-and-sum beamformer and a reduced-rank beamformer, and bringing the initial beamformer into the initial optimization problem to generate a transformed optimization problem; the initial optimization problem is to minimize the error between an ideal beam pattern and an actual beam pattern while constraining the white noise gain to be within an ideal range and enabling the desired direction signal to be transmitted without distortion.

2. The method of claim 1, wherein, converting time-domain signals of the microphone array into frequency-domain signals by using a short-time Fourier transform technique, comprising: given a multi-channel input signal, converting the multi-channel input signal from a time-domain signal into a frequency-domain signal by using a short-time Fourier transform technique, and determining the size of the frequency-domain signal.

3. The method of claim 1, wherein, generating a robust super-directivity beamformer for the optimization problem, comprising: constructing a cost function based on the optimization problem by using a Lagrange multiplier method, and converting the optimization problem into a standard form; obtaining a closed-form solution of the robust super-directivity beamformer based on the standard form of the optimization problem.

4. The method of claim 1, wherein, performing spatial filtering on the frequency-domain signals of the microphone array by using the robust super-directivity beamformer to obtain frequency-domain output signals, comprising: traversing each frame of the frequency-domain signals of the microphone array, for each frame of the frequency-domain signals, traversing each frequency point, and intercepting a target signal; performing weighted filtering on the target signal by using the robust super-directivity beamformer, and obtaining a corresponding frequency-domain output signal, until the weighted filtering operation is completed for each frequency point of each frame.

5. The method of claim 1, wherein, inversely transforming the frequency-domain output signals into time-domain output signals, comprising: converting the frequency-domain output signals into time-domain output signals by using an inverse short-time Fourier transform technique.

6. A robust beamforming apparatus of a microphone array, configured to implement the robust beamforming method of any one of claims 1-5, characterized in that, The method comprises: an acquisition module configured to set up a microphone array and obtain audio information of the microphone array; a conversion module configured to convert time-domain signals of the microphone array into frequency-domain signals by using a short-time Fourier transform technique; an optimization module configured to determine an optimization problem by using the microphone array, and solve a robust super-directivity beamformer for the optimization problem; a processing module configured to perform spatial filtering on the frequency-domain signals of the microphone array by using the robust super-directivity beamformer to obtain frequency-domain output signals, and inversely transform the frequency-domain output signals into time-domain output signals.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor, when executing the program, implements the robust beamforming method of the microphone array according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the robust beamforming method of the microphone array according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Microphone array voice enhancement method and related equipment

    CN113035216A

  • Robust adaptive beam forming directional pickup method based on subarray division

    CN113593596A