Microphone signal beamforming processing method, electronic device and storage medium

By performing time-frequency transformation and cross-mode analysis on microphone signals, the problem of beam performance attenuation in traditional beamforming methods is solved, and beamforming processing with narrower beams and higher signal-to-noise ratio is achieved.

WO2025208671A1PCT designated stage Publication Date: 2025-10-09AAC TECHNOLOGIES PTE LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/088930
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2024-04-19
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Traditional beamforming methods suffer from beam performance attenuation, cannot achieve narrow beams, and cannot effectively suppress sidelobes and background noise.

Method used

The output signals of multiple microphones are transformed in time-frequency mode, and multiple groups of beamforming preprocessing and cross-mode analysis are performed to obtain positive weighting coefficients. The combined coefficients are multiplied with the frequency domain signals, and inverse time-frequency transformation is performed to generate weighted spectral components.

Benefits of technology

It achieves beamforming with narrower beams, improves the signal-to-noise ratio, and effectively suppresses sidelobe noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024088930_09102025_PF_FP_ABST
    Figure CN2024088930_09102025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present invention relate to the technical field of microphone signal processing. Disclosed is a microphone signal beamforming processing method. In the present invention, the method comprises: performing time-frequency transformation on first output signals corresponding to a plurality of microphones, the number of which is no less than three, to obtain frequency-domain signals, and performing a plurality of different sets of beamforming preprocessing on the frequency-domain signals of the plurality of microphones to obtain a plurality of different beam signals; performing a plurality of different sets of cross-mode analysis on the plurality of beam signals to obtain a plurality of positive weighting coefficients; multiplying the plurality of positive weighting coefficients to obtain a combined coefficient, and multiplying the combined coefficient by a frequency-domain signal of any microphone to obtain a weighted spectral component; and performing inverse time-frequency transformation on the weighted spectral component to obtain second output signals corresponding to the plurality of microphones. Compared with conventional methods, combined cross-mode analysis used in the present invention achieves better channel separation, thereby obtaining narrower beams having great sidelobe suppression and a higher signal-to-noise ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Microphone signal beamforming processing method, electronic device and storage medium

[0001] This application is based on the U.S. patent application with application number "US18625253" and application date of April 3, 2024, and claims priority of the above-mentioned U.S. patent application. The entire content of the above-mentioned U.S. patent application is hereby incorporated into this application by introduction. Technical Field

[0002] Embodiments of the present invention relate to the technical field of microphone signal processing, and in particular to a microphone signal beamforming processing method, electronic equipment, and storage medium. Background Art

[0003] The most widely used beamformers are delay and difference beamformers, which can be implemented using fixed or adaptive polar patterns. More advanced beamformer groups use these as a starting point but add a postfilter, typically implemented as a frequency-domain subband filter, to further suppress sidelobes, reverberation, and background noise.

[0004] Conventional beamforming approaches suffer from performance limitations in terms of system size, dynamic range (especially the noise gain due to beamforming), sidelobe suppression, and polar pattern frequency independence. Post-filtering schemes used to date employ simple processing that limits beam steering flexibility and prevents narrow beams from being achieved. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a microphone signal beamforming processing method, electronic device and storage medium, which solves the problem that traditional beamforming methods have attenuated beam performance and cannot achieve a narrow beam.

[0006] To solve the above technical problems, an embodiment of the present invention provides a microphone signal beamforming processing method, comprising:

[0007] Performing time-frequency transformation on first output signals corresponding to a plurality of microphones (number not less than three) to obtain frequency domain signals, and performing multiple groups of different beamforming preprocessing on the frequency domain signals of the plurality of microphones to obtain multiple different beam signals;

[0008] performing multiple groups of different cross-mode analyses on the multiple beam signals to obtain multiple positive weighting coefficients, wherein the positive weighting coefficients are used to evaluate similarities between the multiple beam signals;

[0009] Multiplying the multiple positive weighting coefficients to obtain a combination coefficient; multiplying the combination coefficient with the frequency domain signal of any microphone to obtain a weighted spectral component;

[0010] Performing inverse time-frequency transformation on the weighted spectral components to obtain second output signals corresponding to the multiple microphones.

[0011] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a microphone signal beamforming processing method.

[0012] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements a microphone signal beamforming processing method when executed by a processor.

[0013] Compared to existing technologies, the microphone signal beamforming processing method of the present invention performs time-frequency transformation on the output signals of multiple microphones to obtain frequency domain signals. It then performs beamforming processing on the frequency domain signals to obtain beam signals. Specifically, it evaluates the similarity of different beam elements based on multiple sets of cross-mode analysis to obtain multiple positive weighting coefficients. These positive weighting coefficients are then multiplied by the frequency domain signals of any microphone to obtain weighted spectral components. This combined cross-mode analysis achieves better channel separation than traditional methods, resulting in narrower beams with excellent sidelobe suppression and a higher signal-to-noise ratio.

[0014] In addition, the distance between every two microphones in the plurality of microphones is less than half of the highest frequency wavelength in the target application scenario.

[0015] In addition, the time-frequency transformation of the first output signals corresponding to the multiple microphones (number not less than three) to obtain the frequency domain signal includes: using multiple time-frequency transformation modules corresponding one-to-one to the multiple microphones to perform time-frequency transformation on the first output signals of the corresponding microphones to obtain the frequency domain signal of the microphone.

[0016] In addition, performing multiple groups of different beamforming preprocessing on the frequency domain signals of the multiple microphones to obtain multiple different beam signals includes: using multiple beamformers with a number of not less than two to perform multiple groups of different beamforming preprocessing on the frequency domain signals of the multiple microphones; wherein each beamformer performs a group of beamforming preprocessing on all frequency domain signals output by at least three given time-frequency transformation modules to obtain one beam signal.

[0017] In addition, the beamformer is a steerable beamformer, and beam signals formed by each of the at least two beamformers have different widths and / or directions.

[0018] In addition, performing different groups of cross-mode analyses on the multiple beam signals to obtain multiple positive weighting coefficients includes: using two cross-analysis modules to measure and calculate the correlation and / or coherence between the multiple beam signals to obtain the positive weighting coefficients.

[0019] In addition, before multiplying the combination coefficient with the frequency domain signal of any of the microphones, it includes: performing gain normalization processing on the combination coefficient based on a gain normalization factor and a bottom value, so as to selectively attenuate the input in the direction of the cross-mode similarity below a predetermined threshold, so as to obtain the desired gain in the main lobe direction of the beam generated after the multiplication.

[0020] In addition, each of the frequency domain signals is divided into multiple parts of multiple frequency windows, wherein the frequency window is determined according to the sampling frequency and the size of the time-frequency conversion module; different multiple groups of beamforming preprocessing are performed on the frequency domain signals of the multiple microphones to obtain different multiple beam signals, including: in each frequency window, different multiple groups of beamforming preprocessing are performed on the frequency domain signals of the multiple microphones belonging to the current frequency window to obtain different multiple beam signals belonging to the current frequency window; different multiple groups of cross-mode analysis are performed on the multiple beam signals to obtain multiple positive weighting coefficients, including: in each frequency window, different multiple groups of cross-mode analysis are performed on the different multiple beam signals belonging to the current frequency window to obtain to multiple positive weighted coefficients belonging to the current frequency window; multiplying the multiple positive weighted coefficients to obtain a combination coefficient, multiplying the combination coefficient with the frequency domain signal of any microphone to obtain a weighted spectral component, including: in each frequency window, multiplying the multiple positive weighted coefficients belonging to the current frequency window to obtain a combination coefficient, multiplying the combination coefficient with the frequency domain signal of any microphone belonging to the current frequency window to obtain a weighted spectral component belonging to the current frequency window; performing an inverse time-frequency transform on the weighted spectral component to obtain second output signals corresponding to the multiple microphones, including: combining the weighted spectral components belonging to each frequency window and performing an inverse time-frequency transform to obtain second output signals corresponding to the multiple microphones. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0022] FIG1 is a flow chart of a microphone signal beamforming processing method provided by an embodiment of the present invention;

[0023] 2 is a flow chart of a microphone signal beamforming processing method provided by an embodiment of the present invention applied to a single microphone group;

[0024] FIG3 is an example diagram of a process for processing positive weighted coefficients and combination coefficients provided by an embodiment of the present invention;

[0025] FIG4 is a function diagram of positive weighting coefficients and combination coefficients provided by an embodiment of the present invention;

[0026] FIG5 is another flow chart of a microphone signal beamforming processing method provided by an embodiment of the present invention;

[0027] 6 is a flow chart of a microphone signal beamforming processing method provided by an embodiment of the present invention applied to a multi-microphone group;

[0028] FIG7 is a schematic diagram of the structure of a multi-layer cross-mode analysis provided by an embodiment of the present invention;

[0029] FIG8 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Modes for Carrying Out the Invention

[0030] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, each embodiment of the present invention will be described in detail below with reference to the accompanying drawings. However, it will be understood by those skilled in the art that in each embodiment of the present invention, many technical details are provided to enable the reader to better understand the present invention. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present invention can be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with each other and referenced to each other under the premise that there is no contradiction.

[0031] In the embodiments of the present invention, "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent the following three situations: A exists alone; A and B exist simultaneously; and B exists alone. A and B can be singular or plural.

[0032] In the embodiments of the present invention, the symbol " / " can indicate that the preceding and following objects are in an "or" relationship. In addition, the symbol " / " can also represent a division sign, that is, performing a division operation. For example, A / B can mean A divided by B.

[0033] In the embodiments of the present invention, the symbol "*", "·", or "×" can represent a multiplication sign, that is, performing a multiplication operation. For example, A*B, A·B, or A×B can represent A multiplied by B.

[0034] The technical solutions, beneficial effects, and related concepts involved in the embodiments of the present invention are described in detail below.

[0035] It should be noted that embodiments of the present invention provide a method for beamforming microphone signals, which can include any microphone device or other audio signal sensor. In some possible implementations, the method is applied to a single, closely spaced group of audio signal sensors, including: portable devices such as mobile phones and tablets; a single device with multiple microphone units for teleconferencing systems (with or without embedded speakers) or networked smart speakers; "single-point" stereo or surround recording microphones (such as camera accessory microphones); and the like. In other possible implementations, the method is used to implement beamforming when using multiple independent microphone groups. Application examples include: spatially separated microphone groups on the same device, such as AR / VR / XR / Telepresence glasses with two-sided microphone groups, wearable wireless headphones, hearing devices (hearing aids), or other augmented hearing systems; combining multiple nearby teleconferencing devices; and combining signals from multiple microphone groups in a car cockpit. Among other possible implementations, the method could be applied to adjusting the focus of an acoustic system; combining acoustic beam steering with visual target recognition or control via an image-based user interface; and using eye tracking to control acoustic beams, particularly in wearable devices.

[0036] One embodiment of the present invention relates to a microphone signal beamforming processing method. The core of this embodiment is to perform time-frequency transformation on first output signals corresponding to a plurality of microphones (not less than three) to obtain frequency domain signals, and perform multiple groups of different beamforming preprocessing on the frequency domain signals of the multiple microphones to obtain multiple different beam signals; perform multiple groups of different cross-mode analysis on the multiple beam signals to obtain multiple positive weighting coefficients, and the positive weighting coefficients are used to evaluate the similarity between the multiple beam signals; multiply the multiple positive weighting coefficients to obtain a combination coefficient; multiply the combination coefficient with the frequency domain signal of any of the microphones to obtain a weighted spectral component; and perform an inverse time-frequency transformation on the weighted spectral component to obtain a second output signal corresponding to the multiple microphones.

[0037] Compared with the prior art, the microphone signal beamforming processing method of this embodiment performs time-frequency transformation on the output signals of multiple microphones to obtain frequency domain signals, making the frequency distribution of the signals clearly visible, thereby more accurately analyzing the spectral characteristics of the signals; further beamforming processing is performed on the frequency domain signals to obtain beam signals, which can independently process sounds from different directions and achieve the effect of multi-channel sound processing; then, based on multiple groups of cross-mode analysis, the similarity of different beam elements is evaluated to obtain multiple positive weighting coefficients, which are then multiplied by the positive weighting coefficients with the frequency domain signals of any microphone to obtain weighted spectral components.

[0038] The microphone signal beamforming processing method of this embodiment achieves better channel separation than traditional methods by using combined cross-mode analysis. Because this embodiment combines cross-mode analysis with weighted processing, it can enhance the sound signal in a specific direction, thereby improving the signal-to-noise ratio, thereby obtaining a narrower beam with excellent sidelobe suppression and a higher signal-to-noise ratio.

[0039] The implementation details of the microphone signal beamforming processing method of this embodiment are described in detail below. The following content is only provided for easy understanding and is not necessary for implementing this solution.

[0040] Please refer to Figures 1 and 2. Figure 1 is a flow chart of the microphone signal beamforming processing method of this embodiment, and Figure 2 is a flow chart of the microphone signal beamforming processing method of this embodiment applied to a single microphone group. A microphone signal beamforming processing method in this embodiment specifically includes:

[0041] S101: Processing signals collected by multiple microphones into frequency domain signals.

[0042] Specifically, the multiple microphones comprise a microphone group of at least three microphones, and a time-frequency transform is performed on the first output signals corresponding to all of the microphones to obtain a frequency domain signal. The distance between each two microphones in the multiple microphones is less than half the wavelength of the highest frequency in the target application scenario. By setting the distance between the microphones, it is ensured that the distance between the microphones does not cause phase differences or signal superposition during the sound wave collection process.

[0043] To facilitate understanding, some examples of the application scenarios and frequency wavelengths described above are provided here. In some examples, for voice communication application scenarios, a sound signal frequency of up to 8 kHz is sufficient. In other examples, for music and general recording application scenarios, a frequency of at least 10 kHz or 20 kHz must be processed. Those skilled in the art will understand that the "frequency wavelength" mentioned above refers to sound waves at these frequencies. For example, at 10 kHz, half the wavelength is 17 mm, meaning the distance between each two microphones should be less than 17 mm.

[0044] Specifically, referring to Figure 2, this embodiment utilizes multiple time-frequency transform modules (TFT) corresponding one-to-one to the multiple microphones (Mic) to perform time-frequency transform on the first output signal of the corresponding microphone to obtain the frequency domain signal of the microphone. The frequency domain signal is further divided into multiple frequency bins (Frequency Bins 1, 2, …, K). Within each frequency bin, the frequency domain signal processed by each TFT module and belonging to the current frequency bin is sent to a beamformer (Steerable beamformer).

[0045] S102: Perform beamforming preprocessing on the frequency domain signal to obtain different beam signals.

[0046] Specifically, in this embodiment, S101, each frequency domain signal is divided into multiple parts within multiple frequency windows, where the frequency windows are determined based on the sampling frequency and the size of the time-frequency transformation module. Within each frequency window, the frequency domain signals of the multiple microphones belonging to the current frequency window are subjected to different sets of beamforming preprocessing to obtain different beam signals belonging to the current frequency window. For ease of explanation, the subsequent steps are performed within the same frequency window, and each step performs corresponding operations within each frequency window.

[0047] Specifically, in this embodiment, different multiple groups of beamforming preprocessing are performed on the frequency domain signals of the multiple microphones to obtain different multiple beam signals, including: using multiple beamformers with a number of not less than two to perform different multiple groups of beamforming preprocessing on the frequency domain signals of the multiple microphones; wherein each beamformer performs a group of beamforming preprocessing on all frequency domain signals output by at least three given time-frequency conversion modules to obtain one beam signal.

[0048] Here, each beamformer is a steerable beamformer, and the beam signals formed by each beamformer have different widths and / or directions, and the beam signals vary according to the direction of the incident sound captured by the microphones, and the microphones are spatially distanced.

[0049] Specifically, referring to FIG. 2 , in this embodiment, in the same frequency bin, each beamformer receives a frequency domain signal processed by a time-frequency transform module to obtain a beam signal.

[0050] Those skilled in the art will appreciate that the beamformer and the time-frequency transform are both linear functions, so the order in which they are executed can be changed, allowing each steerable beamformer to be connected to any two or more microphones. In actual applications, it has been found that implementing the beamformer in the frequency domain may be computationally more efficient, and the output of the time-frequency transform can be shared across multiple instances of the beamforming algorithm. Therefore, in this embodiment, the time-frequency transform in step S101 is performed first, followed by the beamforming process in step S102.

[0051] S103: Obtain a positive weighting coefficient based on the cross-mode analysis, and obtain a combination coefficient based on the positive weighting coefficient.

[0052] Specifically, different groups of cross-mode analyses are performed on the multiple beam signals to obtain multiple positive weighting coefficients. The positive weighting coefficients are used to evaluate the similarity between the multiple beam signals. The function of cross-mode analysis is described in detail in the prior art, particularly in US Pat. No. 9,681,220 B2. In this embodiment, the similarity between the multiple beam signals is analyzed based on cross-mode analysis, and the similarity analysis methods used include but are not limited to coherence or correlation, phase similarity, etc.

[0053] Specifically, referring to FIG. 2 , in this embodiment, two independent cross-pattern analyses are used to measure and calculate the similarity between the multiple beam signals. The similarity analysis methods used include, but are not limited to, coherence or correlation, phase similarity, and the like. The two independent cross-pattern analyses each yield different positive weighting coefficients G1 and G2.

[0054] Specifically, referring to FIG. 2 , in this embodiment, two positive weighted coefficients G1 and G2 are obtained and subjected to a simple scalar multiplication operation (Coefficient multiplication) to obtain a combined coefficient G0.

[0055] For ease of understanding, the embodiment of the present invention provides an example of a process of processing the positive weighting coefficient and the combination coefficient given in S103 , please refer to FIG. 3 .

[0056] Specifically, in Figure 3, the horizontal and vertical axes of each graph should be interpreted as the spatial components of the directional vector in the plane passing through the microphone's acoustic inlet, and the distance from the origin on the curve represents the amplitude. The leftmost subgraph in Figure 3 represents the two beam signals BF1 and BF2 obtained from the beamformer output. BF1 and BF2 are then input into two independent cross-pattern analyses, resulting in two different positive weighting coefficients G1 and G2, as shown in the second left subgraph in Figure 3. The two obtained positive weighting coefficients G1 and G2 are plotted on the same graph, as shown in the second right subgraph in Figure 3, G1 & G2. A simple scalar multiplication is then performed on these two positive weighting coefficients G1 and G2. After this multiplication, the combined coefficient G0 is obtained, which is represented by the gray portion of the Overlapping Pattern G0 in the rightmost subgraph in Figure 3, resulting in a narrower polarization pattern for the control signal.

[0057] Please refer to FIG. 4 , which is a function diagram of the positive weighting coefficients G1 and G2 and the combination coefficient G0 in this embodiment.

[0058] Specifically, in Figure 4, the horizontal axis of each subgraph represents the incident sound direction (angle θ), and the vertical axis represents the weight function values ​​of each coefficient (see the figure caption for details). The left subgraph of Figure 4 shows the weight functions of the weight coefficients G1 and G2 obtained from the initial cross-mode analysis. As shown in the middle subgraph G1 & G2, the two positive weight coefficients G1 and G2 are plotted on the same graph. A simple scalar multiplication of the two positive weight coefficients G1 and G2 yields the weight function of the combined coefficient G0, as shown in the right subgraph of Figure 4.

[0059] S104: Multiply the combination coefficient and the frequency domain signal to obtain a weighted spectrum component.

[0060] Specifically, after obtaining the combination coefficient G0 by performing a simple scalar multiplication operation on the two positive weight coefficients G1 and G2, the combination coefficient is normalized based on the gain normalization factor and bottom value A gain normalization process is performed to selectively attenuate inputs in directions with cross-mode similarities lower than a predetermined threshold, so as to obtain a desired gain in a main lobe direction of the beam generated after the multiplication.

[0061] Please refer to Figure 2. As shown in the Gain normalization part of Figure 2, based on the gain normalization factor and bottom value After completing the gain normalization process, the combination coefficient G is obtained 0,norm Then the combination coefficient G 0,norm A simple scalar multiplication operation is performed with the frequency domain signal obtained by processing any microphone signal, such as the signal post-filtering multiplication in Figure 2, to obtain a weighted spectral component.

[0062] It should be clarified that the signal multiplication operation part (Signal post-filtering multiplication) shown in Figure 2 provided by this embodiment uses the frequency domain signal from Mic 3 and Time-Frequency transform 3 as a multiplier for the multiplication operation. This is only an example. The key is to connect the frequency domain signal output by one and only one time-frequency transform module (Time-Frequency transform) to the signal multiplier controlled by the coefficient G0,norm to complete the multiplication operation, thereby obtaining the weighted spectral component.

[0063] S105: Combine the weighted spectral components and perform inverse time-frequency transform to obtain an output signal.

[0064] Specifically, referring to FIG2 , the weighted spectral components obtained in each frequency window are obtained, and the weighted spectral components belonging to each frequency window are combined and then subjected to inverse time-frequency transform to obtain the second output signals corresponding to the multiple microphones.

[0065] In this embodiment, a time-frequency transform is performed on the output signals of multiple microphones to obtain a frequency domain signal. Beamforming is then performed on the frequency domain signal to obtain a beam signal. Specifically, the similarity of different beam elements is evaluated based on multiple sets of cross-mode analysis to obtain multiple positive weighting coefficients. The positive weighting coefficients are then multiplied by the frequency domain signal of any microphone to obtain weighted spectral components, thereby producing a narrower beam with excellent sidelobe suppression and a high signal-to-noise ratio. By using an additional beamformer and cross-mode analysis capabilities, this embodiment can simultaneously generate multiple beams pointing in different directions using the same physical microphone array. In these use cases, the combined cross-mode analysis achieves better channel separation than traditional methods.

[0066] Furthermore, those skilled in the art will appreciate that the scope of use of the invention can be extended to situations where there are two or more physically separate microphone groups in a system, which are located in different parts of the same device (such as a computer or augmented reality / virtual reality glasses) or in physically separate units (such as two sides of a headphone system, connected by an electrical connection or by a wireless interface), in which case the beamforming processing can be distributed.

[0067] Based on this, the present invention provides another embodiment, which relates to a method for microphone signal beamforming. This embodiment is substantially similar to the previous embodiment, with the primary difference being that, in the previous embodiment, the microphone signal beamforming method is applied to a single, closely spaced audio signal sensor group consisting of multiple microphones. In this embodiment, however, the method is applied to two spatially separated microphone groups or to multiple independent microphone groups.

[0068] Please refer to Figures 5 and 6. Figure 5 is another flow chart of the microphone signal beamforming processing method provided in this embodiment, and Figure 6 is a flow chart of the microphone signal beamforming processing method provided in this embodiment applied to a multi-microphone group. A microphone signal beamforming processing method in this embodiment specifically includes:

[0069] S201: Processing signals collected by multiple microphone groups into frequency domain signals.

[0070] Specifically, the multiple microphones are a microphone group consisting of no fewer than three microphones, and a time-frequency transform is performed on the first output signals corresponding to the multiple microphones in the microphone group to obtain a frequency domain signal. The distance between each two microphones in the multiple microphones is less than half the wavelength of the highest frequency in the target application scenario. By setting the distance between the microphones, it is ensured that the distance between the microphones does not cause phase differences or signal superposition during the sound wave collection process.

[0071] Specifically, referring to Figure 6 , microphone groups A and B are provided, each containing three microphones. Note that the microphone groups shown in Figure 6 are only examples, and this embodiment can process more microphone groups simultaneously, and each microphone group can also contain more than three microphones.

[0072] Specifically, referring to Figure 6, for each microphone group, this embodiment employs multiple time-frequency transform modules (Time-Frequency Transform) corresponding to the multiple microphones (Mic) to perform time-frequency transform on the first output signal of the corresponding microphone to obtain the frequency domain signal of the microphone. The frequency domain signal is further divided into multiple frequency bins (Frequency Bins 1, 2, …, K). Within each frequency bin, the frequency domain signal for the current frequency bin, obtained by processing each time-frequency transform module of each microphone group, is sent to all beamformers (Steerable Beamformers) for subsequent processing.

[0073] S202: Perform beamforming preprocessing on the frequency domain signal to obtain different beam signals.

[0074] Specifically, in S201 of this embodiment, each frequency domain signal is divided into multiple parts and input into multiple frequency windows, where the frequency windows are determined based on the sampling frequency and the size of the time-frequency transformation module. Within each frequency window, different groups of beamforming preprocessing are performed on the frequency domain signals of the multiple microphones belonging to the current frequency window, thereby obtaining multiple different beam signals belonging to the current frequency window.

[0075] For ease of explanation, the subsequent steps are all performed in the same frequency window, and each step performs the same operation in each frequency window.

[0076] Specifically, in this embodiment, each microphone group uses at least two beamformers to perform multiple different groups of beamforming preprocessing on the frequency domain signals of the multiple microphones included in the microphone group; wherein each of the at least two beamformers only performs one group of beamforming preprocessing on all the frequency domain signals output by the given at least three time-frequency transformation modules to obtain one beam signal. Here, each beamformer is a steerable beamformer, and the beam signals formed by each beamformer have different widths and / or directions. The beam signals vary according to the direction of the incident sound captured by the microphone, and the microphones are spatially spaced apart.

[0077] Specifically, please refer to Figure 6. In this embodiment, in the same frequency bin, each beamformer receives the frequency domain signals processed by all time-frequency transform modules of the corresponding microphone group to obtain a beam signal.

[0078] Those skilled in the art will appreciate that the beamformer and time-frequency transform are both linear functions, so the order in which they are executed can be changed, allowing each steerable beamformer to be connected to any two or more microphones. In actual applications, it has been found that implementing the beamformer in the frequency domain may be computationally more efficient, and the output of the time-frequency transform can be shared across multiple instances of the beamforming algorithm. Therefore, in this embodiment, the time-frequency transform in S201 is performed first, followed by the beamforming process in S202.

[0079] S203: Obtain a positive weighted coefficient based on the cross-mode analysis, and obtain a combination coefficient based on the positive weighted coefficient.

[0080] Specifically, different groups of cross-mode analyses are performed on the multiple beam signals to obtain multiple positive weighting coefficients. The positive weighting coefficients are used to evaluate the similarity between the multiple beam signals. The function of cross-mode analysis is described in detail in the prior art, particularly in US Pat. No. 9,681,220 B2. In this embodiment, the similarity between the multiple beam signals is analyzed based on cross-mode analysis, and the similarity analysis methods used include but are not limited to coherence or correlation, phase similarity, etc.

[0081] Specifically, referring to FIG6 , in this embodiment, two independent cross-pattern analyses are used to measure and calculate the similarity between multiple beam signals of each microphone group. The similarity analysis methods used include, but are not limited to, coherence or correlation, phase similarity, and the like. The two independent cross-pattern analyses each yield different positive weighting coefficients G1 and G2.

[0082] Specifically, referring to FIG. 6 , in this embodiment, two positive weighted coefficients G1 and G2 are obtained and subjected to a simple scalar multiplication operation (Coefficient multiplication) to obtain a combined coefficient G0.

[0083] Please refer to FIG. 3 , which is a diagram illustrating an example of the processing of step S103 .

[0084] In Figure 3, the horizontal and vertical axes of each graph should be interpreted as the spatial components of the direction vector in the plane passing through the acoustic inlet of the microphone, and the distance from the origin of the curve represents the amplitude.

[0085] Specifically, the leftmost subgraph represents the two beam signals BF1 and BF2 obtained from the beamformer output. BF1 and BF2 are then input into two independent cross-pattern analyses, resulting in two different positive weighting coefficients G1 and G2, as shown in the second left subgraph in Figure 3. The two obtained positive weighting coefficients G1 and G2 are plotted on the same graph, as shown in the second right subgraph G1&G2. A simple scalar multiplication is then performed on these two positive weighting coefficients G1 and G2. This multiplication yields the combined coefficient G0, which is represented by the gray portion of Overlapping Pattern G0 in the rightmost subgraph in Figure 3, resulting in a narrower polarization pattern for the control signal.

[0086] Please refer to FIG. 4 , which is a functional expression of the positive weighting coefficients G1 and G2 and the combination coefficient G0 in this embodiment.

[0087] In Figure 4, the horizontal axis of each sub-graph is the incident sound direction (angle θ), and the vertical axis is the weight function value of each coefficient (please refer to the figure title for details).

[0088] Specifically, the left sub-graph of FIG4 is the weight function of the weight coefficients G1 and G2 from the initial cross-mode analysis; as shown in the middle sub-graph G1&G2, the two positive weight coefficients G1 and G2 are plotted on the same graph; the two positive weight coefficients G1 and G2 are subjected to a simple scalar multiplication operation (Coefficient multiplication) to obtain the weight function of the combination coefficient G0, as shown in the right sub-graph of FIG4.

[0089] If more microphone groups are needed or more positive weighting coefficients need to be calculated, additional cross-mode analysis layers can be added, as shown in Figure 7. This further reduces the beamwidth and creates a wider null region. The advantage of this approach over using a higher-order beamformer in the preprocessing stage is that the artifacts of a simple higher-order beamformer do not affect the final result, although some noise gain may occur. Figure 7 illustrates the architecture of multi-layer cross-mode analysis processing.

[0090] Specifically, when multiple layers of cross-mode analysis are used, the number of cross-mode analysis modules in each layer should be an integer multiple of 2, and each pair of independent cross-mode analysis modules constitutes a group. The initial input to the multi-layer cross-mode analysis is the beam signal processed by the beamformer. Each cross-mode analysis module outputs a positive weighted coefficient. These two positive weighted coefficients are then combined by a simple scalar multiplication to obtain the combined coefficients. The combined coefficients are then used as input for the next layer of cross-mode analysis. This process is repeated to obtain the final combined coefficients.

[0091] S204: Multiply the combination coefficient and the frequency domain signal to obtain a weighted spectrum component.

[0092] Specifically, after obtaining the combination coefficient G0 by performing a simple scalar multiplication operation on the two positive weight coefficients G1 and G2, the combination coefficient is normalized based on the gain normalization factor and bottom value A gain normalization process is performed to selectively attenuate inputs in directions with cross-mode similarities lower than a predetermined threshold, so as to obtain a desired gain in a main lobe direction of the beam generated after the multiplication.

[0093] Please refer to Figure 6. As shown in Figure 6, the gain normalization process (Gain normalization) is based on the gain normalization factor and bottom value After completing the gain normalization process, the combination coefficient G is obtained 0,norm Then the combination coefficient G 0,norm A simple scalar multiplication operation is performed on the frequency domain signal obtained by processing the signal of any microphone included in any microphone group, such as the signal multiplication operation part (Signal post-filtering multiplication) in Figure 6, to obtain a weighted spectral component.

[0094] It should be clarified that the signal multiplication operation part (Signal post-filtering multiplication) shown in Figure 6 provided by this embodiment uses the frequency domain signal of Mic 3 and Time-Frequency transform 3 from microphone group A as a multiplier of the multiplication operation. This is only an example. The key is to connect the frequency domain signal output by one and only one time-frequency transform module (Time-Frequency transform) to the signal multiplier controlled by the coefficient G0,norm to complete the multiplication operation, thereby obtaining the weighted spectral component.

[0095] S205: Combine the weighted spectral components and perform inverse time-frequency transform to obtain an output signal.

[0096] Specifically, referring to FIG6 , the weighted spectral components obtained in each frequency window are obtained, and the weighted spectral components belonging to each frequency window are combined and then subjected to inverse time-frequency transform to obtain the second output signals corresponding to the multiple microphones.

[0097] In this embodiment, the output signals of multiple microphones within multiple microphone groups are transformed in time-frequency to obtain frequency domain signals. Beamforming processing is then performed on the frequency domain signals to obtain beam signals. In particular, the similarity of different beam elements is evaluated based on multiple sets of cross-mode analysis to obtain multiple positive weighting coefficients. The positive weighting coefficients are then multiplied by the frequency domain signal of any microphone to obtain weighted spectral components, thereby producing a narrower beam with excellent sidelobe suppression and a high signal-to-noise ratio. In this embodiment, by using additional beamformers and cross-mode analysis capabilities, multiple beams pointing in different directions can be simultaneously generated using the same physical microphone array, for example, for multi-channel (surround sound) recording applications. In these use cases, the combined cross-mode analysis achieves better channel separation than traditional methods.

[0098] The steps of the various methods above are divided only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0099] Another embodiment of the present invention relates to an electronic device, as shown in Figure 8, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program or instructions executable by the at least one processor, and the computer program or instructions are executed by the at least one processor to enable the at least one processor to perform the microphone signal beamforming processing method as described above.

[0100] The memory and processor are connected using a bus, which can include any number of interconnected buses and bridges. The bus connects various circuits of one or more processors and memories. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and are therefore not described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over a wireless medium via an antenna. Furthermore, the antenna receives data and transmits it to the processor.

[0101] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.

[0102] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program, which implements the above method embodiment when executed by a processor.

[0103] That is, those skilled in the art will understand that all or part of the steps in the above-described method embodiments can be implemented by instructing related hardware through a program. The program is stored in a storage medium and includes a number of instructions for causing a device (such as a microcontroller or chip) or a processor to execute all or part of the steps in the method embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0104] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0105] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0106] Those skilled in the art will appreciate that the above-mentioned embodiments are specific examples for implementing the present invention, and that in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present invention.

Claims

1. A microphone signal beamforming processing method, characterized in that: include: Performing time-frequency transformation on first output signals corresponding to a plurality of microphones (number not less than three) to obtain frequency domain signals, and performing multiple groups of different beamforming preprocessing on the frequency domain signals of the plurality of microphones to obtain multiple different beam signals; performing multiple groups of different cross-mode analyses on the multiple beam signals to obtain multiple positive weighting coefficients, wherein the positive weighting coefficients are used to evaluate similarities between the multiple beam signals; multiplying the multiple positive weighting coefficients to obtain a combination coefficient, and multiplying the combination coefficient with the frequency domain signal of any one of the microphones to obtain a weighted spectral component; Performing inverse time-frequency transformation on the weighted spectral components to obtain second output signals corresponding to the multiple microphones.

2. The microphone signal beamforming processing method according to claim 1, characterized in that: The distance between every two microphones in the plurality of microphones is less than half of the highest frequency wavelength in the target application scenario.

3. The microphone signal beamforming processing method according to claim 1, wherein: The step of performing time-frequency transformation on first output signals corresponding to a plurality of microphones (not less than three) to obtain frequency domain signals includes: A plurality of time-frequency conversion modules corresponding one-to-one to the plurality of microphones are used to perform time-frequency conversion on the first output signal of the corresponding microphone to obtain the frequency domain signal of the microphone.

4. The microphone signal beamforming processing method according to claim 3, characterized in that: Performing multiple groups of different beamforming preprocessing on the frequency domain signals of the multiple microphones to obtain multiple different beam signals includes: A plurality of beamformers, not less than two in number, are used to perform multiple groups of different beamforming preprocessing on the frequency domain signals of the plurality of microphones; wherein each beamformer performs a group of beamforming preprocessing on all frequency domain signals output by the given at least three time-frequency transformation modules to obtain one beam signal.

5. The microphone signal beamforming processing method according to claim 4, characterized in that: The beamformer is a steerable beamformer, and beam signals formed by each of the at least two beamformers have different widths and / or directions.

6. The microphone signal beamforming processing method according to any one of claims 1 to 5, characterized in that: Performing multiple groups of different cross-mode analyses on the multiple beam signals to obtain multiple positive weighting coefficients, including: Two cross analysis modules are used to respectively perform measurement calculations on the correlation and / or coherence between the multiple beam signals to obtain the positive weighting coefficients.

7. The microphone signal beamforming processing method according to claim 1, characterized in that: Before multiplying the combination coefficient with the frequency domain signal of any microphone, the method includes: The combination coefficients are gain normalized based on a gain normalization factor and a bottom value, so as to selectively attenuate inputs in directions with cross-mode similarity below a predetermined threshold, so as to obtain a desired gain in the main lobe direction of the beam generated after the multiplication.

8. The microphone signal beamforming processing method according to any one of claims 1 to 7, characterized in that: Dividing each of the frequency domain signals into multiple parts input into multiple frequency windows, wherein the frequency windows are determined according to a sampling frequency and a size of a time-frequency transform module; Performing multiple groups of different beamforming preprocessing on the frequency domain signals of the multiple microphones to obtain multiple different beam signals, including: performing multiple groups of different beamforming preprocessing on the frequency domain signals of the multiple microphones belonging to the current frequency window in each frequency window to obtain multiple different beam signals belonging to the current frequency window; Performing multiple groups of different cross-mode analyses on the multiple beam signals to obtain multiple positive weighting coefficients, including: in each frequency window, performing multiple groups of different cross-mode analyses on multiple different beam signals belonging to the current frequency window to obtain multiple positive weighting coefficients belonging to the current frequency window; Multiplying the multiple positive weighting coefficients to obtain a combination coefficient, and multiplying the combination coefficient with the frequency domain signal of any of the microphones to obtain a weighted spectral component, including: in each frequency window, multiplying the multiple positive weighting coefficients belonging to the current frequency window to obtain a combination coefficient, and multiplying the combination coefficient with the frequency domain signal of any of the microphones belonging to the current frequency window to obtain a weighted spectral component belonging to the current frequency window; The performing an inverse time-frequency transform on the weighted spectral components to obtain the second output signals corresponding to the multiple microphones includes: combining the weighted spectral components belonging to each frequency window and performing an inverse time-frequency transform to obtain the second output signals corresponding to the multiple microphones.

9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the microphone signal beamforming processing method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the microphone signal beamforming processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Microphone array, signal processing method and device thereof, equipment and medium

    CN116017230A

  • Device, method and program for acquiring sound

    JP2011119898A

  • Microphone array, method and apparatus for forming constant directivity beams using the same, and method and apparatus for estimating acoustic source direction using the same

    US20040175006A1

  • Method for spatial filtering of at least one sound signal, computer readable storage medium and spatial filtering system based on cross-pattern coherence

    US20150304766A1