Beamforming post-filtering method, device, storage medium and equipment
The frequency domain signal of the microphone array is obtained through sound source positioning and Fourier transform, the direction power of the sound source is calculated, and the interference components are filtered out using a post-filtering algorithm, which solves the problem of difficult to suppress directional interference after beam formation, and achieves the effect of improving the quality of voice signals.
Patent Information
- Application Number
- CN202111402346.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The prior art is difficult to further suppress directional interference after beam formation, affecting the quality of the voice signal.
The sound source direction of the audio input signal is obtained through the sound source positioning algorithm, Fourier transforms to obtain the frequency domain signal, calculate the sound source direction power at each frequency point, and filter out the interference components in the direction of the interfering sound source through the post-filtering algorithm.
Effectively suppress directional interference and improve voice signal quality.
Smart Images

Figure CN114255773B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing, and in particular to a beamforming post-filtering method, device, storage medium and equipment. Background Art
[0002] At present, microphone arrays are widely used in the research of speech signal processing. Using a beamformer for microphone array processing can make the microphone array spatially selective, that is, it allows the sound source signal in a specified direction to pass through, while suppressing the sound source signals in other directions. The ideal beamformer has a gain of 1 for the desired direction signal and a gain of 0 for other directions. However, in real life, due to errors caused by various factors such as noise, it is difficult to completely eliminate the interfering sound sources in other directions by relying solely on a beamformer, that is, it is difficult to obtain an ideal beamformer. Therefore, a post-filtering scheme is proposed, that is, a post-filtering algorithm is added after beamforming to adjust the gain of each frequency point according to the first-order and second-order statistical characteristics of the desired sound source and the interfering sound source. However, this scheme cannot further suppress directional interference after beamforming, thereby affecting the quality of speech signals. Summary of the invention
[0003] The purpose of the embodiments of the present invention is to provide a beamforming post-filtering method, apparatus, storage medium and device, which can further suppress directional interference after beamforming and improve the quality of speech signals.
[0004] To achieve the above object, an embodiment of the present invention provides a beamforming post-filtering method, comprising:
[0005] Obtaining M audio input signals received by M array elements in a microphone array;
[0006] Using a sound source localization algorithm to localize the sound sources of the M audio input signals, and obtaining N sound source directions, wherein the N sound source directions include a desired sound source direction and N-1 interference sound source directions, N>1, M>1;
[0007] Performing Fourier transform on the M audio input signals to obtain each frequency domain signal of the microphone array at each frequency point;
[0008] At each frequency point, a beamforming algorithm is used to perform beamformer calculation on the N sound source directions to obtain each N beamformer expression of each frequency point;
[0009] Calculating the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expressions;
[0010] According to the gain of each frequency point, each of the N beamformer expressions and each frequency domain signal, construct each set of equations about the beamformer expression and the gain of each frequency point, and obtain each of the N powers of each frequency point by solving the set of equations;
[0011] Calculate the power ratio of each frequency point according to each of the N powers to obtain a gain parameter of each post-filter of each frequency point;
[0012] According to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction, every N-1 interference components in every N-1 interference sound source direction are filtered out.
[0013] As an improvement of the above solution, the equation group is specifically:
[0014] H H (θ1,w)*X(w)*X H (w)*H(θ1,w)=a1*P(θ1,θ1,f w )+a2*P(θ1,θ2,f w )+...+a N *P(θ1,θ N ,f w )
[0015] H H (θ2,w)*X(w)*X H (w)*H(θ2,w)=a1*P(θ2,θ1,f w )+a2*P(θ2,θ2,f w )+...+a N *P(θ2,θ N ,f w ) ......
[0016] H H (θ N ,w)*X(w)*X H (w)*H(θ N ,w)=a1*P(θ N ,θ1,f w )+a2*P(θ N ,θ2,f w )+...+a N *P(θ N ,θ N ,f w )
[0017] The superscript H represents the conjugate transpose, H(θ i ,w) indicates the wth frequency point and the pointing direction is θi The beamformer expression is: X(w) represents the frequency domain signal of the wth frequency point, θ i Denotes the beamformer H(θ i ,w) pointing direction,θ j represents the direction of the jth sound source, P(θ i ,θ j ,f w ) means that when the beamformer H(θ i ,w) points to θ i Direction, for θ j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<i≤N, 0<j≤N, 0<n≤N.
[0018] As an improvement of the above solution, the power ratio calculation for each of the frequency points is performed according to each of the N powers to obtain the gain parameter of each post-filter of each frequency point, specifically including:
[0019] The power ratio of each frequency point is calculated according to the following formula to obtain the gain parameter of each post-filter of each frequency point:
[0020]
[0021] Among them, P(θ1,θ j ,f w ) means that when the beamformer H(θ1,w) points to the desired sound source direction θ1, j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<j≤N, 0<n≤N.
[0022] As an improvement of the above solution, filtering out every N-1 interference components in every N-1 interference sound source direction according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction includes:
[0023] According to the following formula, every N-1 interference components in the direction of every N-1 interference sound source are filtered out:
[0024] Y(w)=gain(w)*X H (w)*H(θ1,w)
[0025] Wherein, the superscript H represents the conjugate transpose, X(w) represents the frequency domain signal at the w-th frequency point, gain(w) represents the gain parameter of the post-filter at the w-th frequency point, H(θ1,w) represents the beamformer expression of the w-th frequency point and the direction of the desired sound source direction θ1, and Y(w) is the audio output signal of the microphone array at the w-th frequency point.
[0026] As an improvement of the above solution, the beamforming algorithm includes at least one of the following: a minimum variance distortionless response algorithm, a delay-and-sum algorithm, or a generalized sidelobe cancellation algorithm.
[0027] As an improvement of the above solution, the sound source localization algorithm includes at least one of the following: a sound source localization algorithm based on controllable beamforming, a sound source localization algorithm based on high-resolution spectrum estimation, or a sound source localization algorithm based on arrival time difference.
[0028] To achieve the above object, an embodiment of the present invention further provides a beamforming post-filtering device, comprising:
[0029] An audio receiving module, used to obtain M audio input signals received by M array elements in a microphone array;
[0030] A sound source localization module, configured to use a sound source localization algorithm to perform sound source localization on the M audio input signals to obtain N sound source directions, wherein the N sound source directions include a desired sound source direction and N-1 interference sound source directions, N>1, M>1;
[0031] A Fourier transform module, used for performing Fourier transform on the M audio input signals to obtain each frequency domain signal of the microphone array at each frequency point;
[0032] A beamformer calculation module, configured to perform beamformer calculation on the N sound source directions at each frequency point using a beamforming algorithm to obtain each N beamformer expression at each frequency point;
[0033] A gain calculation module, configured to calculate the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expressions;
[0034] an equation group construction module, configured to construct each equation group about the beamformer expression and the gain of each frequency point according to the gain of each frequency point, each N beamformer expression and each frequency domain signal, and obtain each N power of each frequency point by solving the equation group;
[0035] A power proportion calculation module, used to calculate the power proportion of each of the frequency points according to each of the N powers, and obtain a gain parameter of each post-filter of each of the frequency points;
[0036] The filtering module is used to filter out every N-1 interference components in every N-1 interference sound source directions according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction.
[0037] To achieve the above-mentioned purpose, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to perform the post-filtering method of beamforming as described above.
[0038] To achieve the above objectives, an embodiment of the present invention further provides a beamforming post-filtering device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the beamforming post-filtering method as described above when executing the computer program.
[0039] Compared with the prior art, the embodiments of the present invention provide a beamforming post-filtering method, device, storage medium and equipment, which obtain N sound source directions of an audio input signal through a sound source localization algorithm, and can accurately obtain the desired sound source direction of the audio input signal. By performing Fourier transform on the audio input signal, the frequency domain signal of each frequency point of the microphone array can be obtained, and the power of each sound source direction of each frequency point can be calculated, so that the interference component of the interfering sound source direction can be filtered out through post-filtering, thereby further suppressing directional interference and improving the quality of the voice signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flow chart of a beamforming post-filtering method provided by an embodiment of the present invention;
[0041] Figure 2 is a structural block diagram of a beamforming post-filtering device provided by an embodiment of the present invention;
[0042] Figure 3 It is a structural block diagram of a beamforming post-filtering device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] See also Figure 1, Figure 1 1 is a flowchart of a beamforming post-filtering method provided by an embodiment of the present invention, wherein the beamforming post-filtering method comprises:
[0045] S1, obtaining M audio input signals received by M array elements in a microphone array;
[0046] S2. Using a sound source localization algorithm to localize the sound sources of the M audio input signals, and obtaining N sound source directions, wherein the N sound source directions include a desired sound source direction and N-1 interference sound source directions, N>1, M>1;
[0047] S3, performing Fourier transform on the M audio input signals to obtain each frequency domain signal of the microphone array at each frequency point;
[0048] S4. At each frequency point, a beamforming algorithm is used to perform beamformer calculation on the N sound source directions to obtain each N beamformer expression of each frequency point;
[0049] S5. Calculate the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expression;
[0050] S6. Constructing a set of equations about the beamformer expression and the gain of each frequency point according to the gain of each frequency point, each of the N beamformer expressions and each frequency domain signal, and obtaining each of the N powers of each frequency point by solving the set of equations;
[0051] S7, calculating the power ratio of each of the frequency points according to each of the N powers, to obtain a gain parameter of each post-filter of each of the frequency points;
[0052] S8. Filter out every N-1 interference components in every N-1 interference sound source directions according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction.
[0053] In an embodiment of the present invention, N sound source directions of an audio input signal are obtained through a sound source localization algorithm, and the desired sound source direction of the audio input signal can be accurately obtained. By performing Fourier transform on the audio input signal, the frequency domain signal of each frequency point of the microphone array can be obtained, and the power of each sound source direction of each frequency point can be calculated. Thus, the interference component of the interfering sound source direction can be filtered out through post-filtering, thereby further suppressing directional interference and improving the quality of the voice signal.
[0054] Optionally, in step S2, the sound source localization algorithm includes at least one of the following: a sound source localization algorithm based on controllable beamforming, a sound source localization algorithm based on high-resolution spectrum estimation, or a sound source localization algorithm based on arrival time difference.
[0055] It should be noted that the sound source localization algorithm is used to localize the sound source of the M audio input signals to obtain N sound source directions, which are recorded as θ1, θ2, ..., θ N , for the convenience of description, the desired sound source direction is recorded as θ1, and the other sound source directions are interference sound source directions; N can be any integer, and is not specifically limited here. It can be understood that when there are less than 5 sound source directions in the actual audio input signal, for example, there are only 3 sound source directions, then the signal output power of the remaining two sound source directions is actually very low, and has no effect on the subsequent steps;
[0056] In the embodiment of the present invention, the sound source direction of the audio input signal can be accurately located by using a sound source localization algorithm to obtain the desired sound source direction and the interference sound source direction.
[0057] Specifically, in step S3, each frequency domain signal includes each frequency and each amplitude;
[0058] Then, perform an n-point discrete Fourier transform on the audio input signal to obtain the frequency of the w-th frequency point:
[0059] f w =wF / n
[0060] Where F is the sampling rate of the audio input signal.
[0061] Optionally, in step S5, the beamforming algorithm includes at least one of the following: a minimum variance distortionless response algorithm, a delay-sum algorithm or a generalized sidelobe cancellation algorithm.
[0062] It can be understood that, assuming that the microphone array is composed of M array elements (microphones), for the wth frequency point of the audio input signal, the signal spectrum received by the mth (1<=m<=M) array element is X m (w), the transfer function of the beamformer in the mth channel is H m (w), therefore, under the action of the beamformer H(w), the output signal Y(w) can be obtained from the audio input signal X(w). The mathematical expression of the beamformer is as follows:
[0063]
[0064] H(w)=[H1(w),H2(w),...,H m (w),...,H M (w)] T
[0065] Common beamformers include minimum variance distortionless response algorithm, delay-sum algorithm, and generalized sidelobe cancellation algorithm. The difference between these algorithms lies in the different calculation methods of H(w), which has no effect on subsequent steps. Therefore, the embodiments of the present invention can adopt any beamformer.
[0066] Specifically, in step S5, calculating the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expression includes:
[0067] The gain of each frequency point is calculated according to the following formula:
[0068] P(θ N ,θ N ,f w )=H H (θ N ,w)*p(θ N ,θ N ,f w )*p H (θ N ,θ N ,f w )*H(θ N ,w)
[0069] In the formula, the superscript H represents the conjugate transpose, H(θ i ,w) indicates the wth frequency point and the pointing direction is θ i The beamformer expression for p(θ i ,θ j ,f w ) means that when the beamformer H(θ i ,w) points to θ i Direction, for θ j direction and frequency f w The matrix response vector of the signal, P(θ i ,θ j ,f w ) means that when the beamformer H(θ i ,w) points to θ i Direction, for θ j direction and frequency f w The gain of the signal;
[0070] It should be noted that the function of the array response vector is to test the response of a beamformer to a specific direction. It is calculated based on the sound source direction, frequency and array geometry, and will not be described in detail here.
[0071] It is understandable that since p(θ i ,θ j,f w ) is 1, so according to the above formula, we can calculate the power of the beamformer H(θ i ,w) points to θ i When the microphone array is in the direction of θ j direction and frequency f w The signal gain P(θ i ,θ j ,f w ).
[0072] Specifically, in step S5, the equation group is:
[0073] H H (θ1,w)*X(w)*X H (w)*H(θ1,w)=a1*P(θ1,θ1,f w )+a2*P(θ1,θ2,f w )+...+a N *P(θ1,θ N ,f w )
[0074] H H (θ2,w)*X(w)*X H (w)*H(θ2,w)=a1*P(θ2,θ1,f w )+a2*P(θ2,θ2,f w )+...+a N *P(θ2,θ N ,f w ) ......
[0075] H H (θ N ,w)*X(w)*X H (w)*H(θ N ,w)=a1*P(θ N ,θ1,f w )+a2*P(θ N ,θ2,f w )+...+a N *P(θ N ,θ N ,f w )
[0076] The superscript H represents the conjugate transpose, H(θ i ,w) indicates the wth frequency point and the pointing direction is θ i The beamformer expression is: X(w) represents the frequency domain signal of the wth frequency point, θ i Denotes the beamformer H(θi ,w) pointing direction,θ j represents the direction of the jth sound source, P(θ i ,θ j ,f w ) means that when the beamformer H(θ i ,w) points to θ i Direction, for θ j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<i≤N, 0<j≤N, 0<n≤N.
[0077] In the embodiment of the present invention, the power of each sound source direction at each frequency point is calculated to obtain the interference component of the array output signal, which is then filtered out by post-value filtering.
[0078] Specifically, in step S7, the power proportion calculation is performed on each of the frequency points according to each of the N powers to obtain the gain parameter of each post-filter of each frequency point, specifically including:
[0079] The power ratio of each frequency point is calculated according to the following formula to obtain the gain parameter of each post-filter of each frequency point:
[0080]
[0081] Among them, P(θ1,θ j ,f w ) means that when the beamformer H(θ1,w) points to the desired sound source direction θ1, j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<j≤N, 0<n≤N.
[0082] Specifically, in step S8, filtering out every N-1 interference components in every N-1 interference sound source direction according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction includes:
[0083] According to the following formula, every N-1 interference components in the direction of every N-1 interference sound source are filtered out:
[0084] Y(w)=gain(w)*X H (w)*H(θ1,w)
[0085] Wherein, the superscript H represents the conjugate transpose, X(w) represents the frequency domain signal at the w-th frequency point, gain(w) represents the gain parameter of the post-filter at the w-th frequency point, H(θ1,w) represents the beamformer expression of the w-th frequency point and the direction of the desired sound source direction θ1, and Y(w) is the audio output signal of the microphone array at the w-th frequency point.
[0086] A beamforming post-filtering method provided in an embodiment of the present invention obtains N sound source directions of an audio input signal through a sound source localization algorithm, and can accurately obtain the desired sound source direction of the audio input signal. By performing Fourier transform on the audio input signal, the frequency domain signal of each frequency point of the microphone array can be obtained, and the power in all sound source directions of each frequency point can be calculated, so that the interference component in the direction of the interfering sound source can be filtered out through post-filtering, thereby further suppressing directional interference and improving the quality of the voice signal.
[0087] See also Figure 2 , Figure 2 1 is a structural block diagram of a beamforming post-filtering device 10 provided in an embodiment of the present invention, wherein the beamforming post-filtering device comprises:
[0088] The audio receiving module 11 is used to obtain M audio input signals received by M array elements in the microphone array;
[0089] A sound source localization module 12 is used to use a sound source localization algorithm to perform sound source localization on the M audio input signals to obtain N sound source directions, wherein the N sound source directions include a desired sound source direction and N-1 interference sound source directions, N>1, M>1;
[0090] A Fourier transform module 13, configured to perform Fourier transform on the M audio input signals to obtain each frequency domain signal of the microphone array at each frequency point;
[0091] A beamformer calculation module 14 is used to perform beamformer calculation on the N sound source directions at each frequency point using a beamforming algorithm to obtain each N beamformer expression at each frequency point;
[0092] A gain calculation module 15, configured to calculate the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expressions;
[0093] an equation group construction module 16, configured to construct each equation group about the beamformer expression and the gain of each frequency point according to the gain of each frequency point, each N beamformer expression and each frequency domain signal, and obtain each N power of each frequency point by solving the equation group;
[0094] A power proportion calculation module 17, configured to calculate the power proportion of each of the frequency points according to each of the N powers, and obtain a gain parameter of each post-filter of each of the frequency points;
[0095] The filtering module 18 is used to filter out every N-1 interference components in every N-1 interference sound source directions according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction.
[0096] Specifically, the equation group is:
[0097] H H (θ1,w)*X(w)*X H (w)*H(θ1,w)=a1*P(θ1,θ1,f w )+a2*P(θ1,θ2,f w )+...+a N *P(θ1,θ N ,f w )
[0098] H H (θ2,w)*X(w)*X H (w)*H(θ2,w)=a1*P(θ2,θ1,f w )+a2*P(θ2,θ2,f w )+...+a N *P(θ2,θ N ,f w ) ......
[0099] H H (θ N ,w)*X(w)*X H (w)*H(θ N ,w)=a1*P(θ N ,θ1,f w )+a2*P(θ N ,θ2,f w )+...+a N *P(θ N ,θ N ,f w )
[0100] The superscript H represents the conjugate transpose, H(θ i ,w) indicates the wth frequency point and the pointing direction is θ i The beamformer expression is: X(w) represents the frequency domain signal of the wth frequency point, θ i Denotes the beamformer H(θ i ,w) pointing direction,θ jrepresents the direction of the jth sound source, P(θ i ,θ j ,f w ) means that when the beamformer H(θ i ,w) points to θ i Direction, for θ j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<i≤N, 0<j≤N, 0<n≤N.
[0101] Specifically, the calculating the power ratio of each frequency point according to each N power points to obtain the gain parameter of each post-filter of each frequency point specifically includes:
[0102] The power ratio of each frequency point is calculated according to the following formula to obtain the gain parameter of each post-filter of each frequency point:
[0103]
[0104] Among them, P(θ1,θ j ,f w ) means that when the beamformer H(θ1,w) points to the desired sound source direction θ1, j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<j≤N, 0<n≤N.
[0105] Specifically, filtering out every N-1 interference components in every N-1 interference sound source direction according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction includes:
[0106] According to the following formula, every N-1 interference components in the direction of every N-1 interference sound source are filtered out:
[0107] Y(w)=gain(w)*X H (w)*H(θ1,w)
[0108] Wherein, the superscript H represents the conjugate transpose, X(w) represents the frequency domain signal at the w-th frequency point, gain(w) represents the gain parameter of the post-filter at the w-th frequency point, H(θ1,w) represents the beamformer expression of the w-th frequency point and the direction of the desired sound source direction θ1, and Y(w) is the audio output signal of the microphone array at the w-th frequency point.
[0109] Preferably, the beamforming algorithm comprises at least one of the following: a minimum variance distortionless response algorithm, a delay-sum algorithm or a generalized sidelobe cancellation algorithm.
[0110] Preferably, the sound source localization algorithm includes at least one of the following: a sound source localization algorithm based on controllable beamforming, a sound source localization algorithm based on high-resolution spectrum estimation, or a sound source localization algorithm based on time difference of arrival.
[0111] It is worth noting that the working process of each module in the beamforming post-filtering device 10 described in the embodiment of the present invention can refer to the working process of the beamforming post-filtering method described in the above embodiment, which will not be repeated here.
[0112] A beamforming post-filtering device 10 provided in an embodiment of the present invention obtains N sound source directions of an audio input signal through a sound source localization algorithm, and can accurately obtain the desired sound source direction of the audio input signal. By performing Fourier transform on the audio input signal, the frequency domain signal of each frequency point of the microphone array can be obtained, and the power in all sound source directions of each frequency point can be calculated, so that the interference component in the direction of the interfering sound source can be filtered out through post-filtering, thereby further suppressing directional interference and improving the quality of the voice signal.
[0113] An embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to perform the post-filtering method of beamforming as described in the above embodiment.
[0114] See also Figure 3 , Figure 3 1 is a structural block diagram of a beamforming post-filtering device 20 provided in an embodiment of the present invention, and the beamforming post-filtering device 20 includes: a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above-mentioned beamforming post-filtering method embodiment are implemented. Alternatively, when the processor 21 executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0115] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the beamforming post-filtering device 20.
[0116] The beamforming post-filtering device 20 may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The beamforming post-filtering device 20 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art may understand that the schematic diagram is merely an example of the beamforming post-filtering device 20 and does not constitute a limitation on the beamforming post-filtering device 20. The beamforming post-filtering device 20 may include more or fewer components than shown in the figure, or may combine certain components, or different components. For example, the beamforming post-filtering device 20 may also include input and output devices, network access devices, buses, etc.
[0117] The processor 21 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor 21 is the control center of the beamforming post-filtering device 20, and uses various interfaces and lines to connect various parts of the entire beamforming post-filtering device 20.
[0118] The memory 22 can be used to store the computer program and / or module. The processor 21 realizes various functions of the beamforming post-filter device 20 by running or executing the computer program and / or module stored in the memory 22 and calling the data stored in the memory 22. The memory 22 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0119] Wherein, if the module / unit integrated in the post-filter device 20 of the beamforming is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-mentioned various method embodiments can be implemented. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0120] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.
[0121] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A post-filtering method for beamforming, characterized in that: include: Obtaining M audio input signals received by M array elements in a microphone array; Using a sound source localization algorithm to localize the sound sources of the M audio input signals, and obtaining N sound source directions, wherein the N sound source directions include a desired sound source direction and N-1 interference sound source directions, N>1, M>1; Performing Fourier transform on the M audio input signals to obtain each frequency domain signal of the microphone array at each frequency point; At each frequency point, a beamforming algorithm is used to perform beamformer calculation on the N sound source directions to obtain each N beamformer expression of each frequency point; Calculating the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expressions; According to the gain of each frequency point, each of the N beamformer expressions and each frequency domain signal, construct each set of equations about the beamformer expression and the gain of each frequency point, and obtain each of the N powers of each frequency point by solving the set of equations; Calculate the power ratio of each frequency point according to each of the N powers to obtain a gain parameter of each post-filter of each frequency point; According to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction, every N-1 interference components in every N-1 interference sound source direction are filtered out.
2. The post-filtering method for beamforming according to claim 1, characterized in that: The equation group is specifically: H H (θ1,w)*X(w)*X H (w)*H(θ1,w)=a1*P(θ1,θ1,f w )+a2*P(θ1,θ2,f w )+...+a N *P(θ1,θ N ,f w ) H H (θ2,w)*X(w)*X H (w)*H(θ2,w)=a1*P(θ2,θ1,f w )+a2*P(θ2,θ2,f w )+...+a N *P(θ2,θ N ,f w ) ...... H H (θ N ,w)*X(w)*X H (w)*H(θ N ,w)=a1*P(θ N ,θ1,f w )+a2*P(θ N ,θ2,f w )+...+a N *P(θ N ,θ N ,f w ) The superscript H represents the conjugate transpose, H(θ i ,w) represents the wth frequency point and the pointing direction is θ i The beamformer expression is: X(w) represents the frequency domain signal of the wth frequency point, θ i Denotes the beamformer H(θ i ,w) pointing direction,θ j represents the direction of the jth sound source, P(θ i ,θ j ,f w ) means that when the beamformer H(θ i ,w) points to θ i Direction, for θ j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<i≤N, 0<j≤N, 0<n≤N.
3. The post-filtering method for beamforming according to claim 1, characterized in that: The calculating the power ratio of each frequency point according to each N power points to obtain the gain parameter of each post-filter of each frequency point specifically includes: The power ratio of each frequency point is calculated according to the following formula to obtain the gain parameter of each post-filter of each frequency point: Among them, P(θ1,θ j ,f w ) means that when the beamformer H(θ1,w) points to the desired sound source direction θ1, j direction and frequency f w The gain of the signal, a n Represents the power in the direction of the nth sound source, 0<j≤N, 0<n≤N.
4. The post-filtering method for beamforming according to claim 1, characterized in that: The filtering out every N-1 interference components in every N-1 interference sound source directions according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction comprises: According to the following formula, every N-1 interference components in the direction of every N-1 interference sound source are filtered out: Y(w)=gain(w)*X H (w)*H(θ1,w) Wherein, the superscript H represents the conjugate transpose, X(w) represents the frequency domain signal at the w-th frequency point, gain(w) represents the gain parameter of the post-filter at the w-th frequency point, H(θ1,w) represents the beamformer expression of the w-th frequency point and the direction of the desired sound source direction θ1, and Y(w) is the audio output signal of the microphone array at the w-th frequency point.
5. The post-filtering method for beamforming according to claim 1, characterized in that: The beamforming algorithm includes at least one of the following: a minimum variance distortionless response algorithm, a delay-sum algorithm, or a generalized sidelobe cancellation algorithm.
6. The post-filtering method for beamforming according to claim 1, characterized in that: The sound source localization algorithm includes at least one of the following: a sound source localization algorithm based on controllable beamforming, a sound source localization algorithm based on high-resolution spectrum estimation, or a sound source localization algorithm based on time difference of arrival.
7. A beamforming post-filtering device, characterized in that: include: An audio receiving module, used to obtain M audio input signals received by M array elements in a microphone array; A sound source localization module, configured to use a sound source localization algorithm to perform sound source localization on the M audio input signals to obtain N sound source directions, wherein the N sound source directions include a desired sound source direction and N-1 interference sound source directions, N>1, M>1; A Fourier transform module, used for performing Fourier transform on the M audio input signals to obtain each frequency domain signal of the microphone array at each frequency point; A beamformer calculation module, configured to perform beamformer calculation on the N sound source directions at each frequency point using a beamforming algorithm to obtain each N beamformer expression at each frequency point; A gain calculation module, configured to calculate the gain of each frequency point according to each N array response vectors of each frequency point and each N beamformer expressions; an equation group construction module, configured to construct each equation group about the beamformer expression and the gain of each frequency point according to the gain of each frequency point, each N beamformer expression and each frequency domain signal, and obtain each N power of each frequency point by solving the equation group; A power proportion calculation module, used to calculate the power proportion of each of the frequency points according to each of the N powers, and obtain a gain parameter of each post-filter of each of the frequency points; The filtering module is used to filter out every N-1 interference components in every N-1 interference sound source directions according to the gain parameter of each post-filter and the beamformer expression of each desired sound source direction.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises a stored computer program; wherein, when the computer program is run, the device where the computer-readable storage medium is located is controlled to perform the post-filtering method for beamforming according to any one of claims 1 to 6.
9. A beamforming post-filtering device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the post-filtering method of beamforming according to any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Audio input system used in home environment based on microphone array
CN102164328A
Two-channel beam forming speech enhancement method based on noise mixed coherence
CN105869651A