Directional sound pickup method, system and device based on small-aperture dual-microphone array

Through the frequency band processing of the small-aperture dual-microphone array, combined with constrained white noise gain and differential beamformer, the voice signal processing is optimized, the problem of low white noise gain in the low-frequency band of the differential beamformer is solved, and better directional pickup effect and voice quality are achieved.

CN118921596BActive Publication Date: 2025-09-19YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410960074.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2025-09-19
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

The existing differential beamformer has low white noise gain in the low-frequency band, which causes the microphone's own noise to be amplified, affecting voice quality. In addition, the multi-channel method has poor directional pickup effect.

Method used

A small-aperture dual-microphone array is used for frequency band processing. A beamformer with constrained white noise gain is used in the low-frequency band, and a differential beamformer is used in the high-frequency band. The differential beamformer with different preset zero points selects the one with the smallest amplitude as the optimal beamformer. Combined with short-time Fourier transform and inverse Fourier transform, speech signal processing is optimized.

Benefits of technology

It effectively suppresses interference signals from other directions, improves directional sound pickup, reduces the impact of the microphone's own noise, and improves voice quality and communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118921596B_ABST
    Figure CN118921596B_ABST
Patent Text Reader

Abstract

The present application discloses a directional sound pickup method, system, and device based on a small-aperture dual-microphone array. The method includes: collecting time domain signals from the dual microphones and converting the time domain signals into frequency domain signals; the frequency domain signals include low-frequency band signals and high-frequency band signals; for the high-frequency band signals, a plurality of differential beamformers with different zero points are preset, and multiple amplitudes of the high-frequency band frequency domain signals are calculated according to the plurality of differential beamformers, and the differential beamformer with the smallest amplitude is selected as the optimal differential beamformer; for the low-frequency band signals, the low-frequency band frequency domain signals are calculated by a beamformer with constrained white noise gain; the frequency domain signals calculated by the beamformer with constrained white noise gain and the optimal differential beamformer are transformed to obtain directionally picked-up speech. The present application can extract speech in the target direction, suppress sound signals in other directions, and improve the directional sound pickup effect of the dual-microphone array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of sound signal processing technology, and in particular to a directional sound pickup method, system and device based on a small-aperture dual-microphone array. Background Art

[0002] When using a microphone to capture a speaker's speech, interference from noise, other speakers, and reverberation often occurs. These interferences can reduce speech quality and intelligibility, and can easily cause auditory fatigue. To reduce these interferences, speech enhancement algorithms need to be designed to effectively improve speech communication efficiency. Existing speech enhancement algorithms are primarily categorized into single-channel and multi-channel approaches. Single-channel approaches typically perform signal processing using a single microphone, but their performance is limited by their inability to obtain spatial information about the sound source. Multi-channel approaches utilize multiple microphones to capture sound signals and extract speech signals in the target direction through methods such as spatial filtering. Within multi-channel approaches, beamforming is an important spatial filtering technique. Differential beamforming, as a type of beamforming technique, uses differentials to approximate the differential of the sound pressure field, thereby obtaining spatial directivity. This approach offers advantages such as good frequency consistency of the beam pattern and suitability for processing wideband speech signals. However, differential beamformers suffer from low white noise gain in the low-frequency band, which amplifies the microphone's own noise and affects speech quality. Therefore, improving the speech quality of directional sound pickup is an urgent issue. Summary of the Invention

[0003] The main purpose of this application is to overcome the shortcomings and deficiencies of the existing technology and provide a directional sound pickup method, system and equipment based on a small-aperture dual-microphone array. By using different beamformers to process microphone signals in different frequency bands, the speech in the target direction is extracted, the sound signals in other directions are suppressed, and the directional sound pickup effect of the dual-microphone array is improved.

[0004] In order to achieve the above objectives, this application adopts the following technical solutions:

[0005] In a first aspect, the present application provides a directional sound pickup method based on a small-aperture dual-microphone array, comprising the following steps:

[0006] Collecting time domain signals from dual microphones and converting the time domain signals into frequency domain signals; the frequency domain signals include low frequency band signals and high frequency band signals;

[0007] For the high-frequency band signal, a plurality of differential beamformers with different zero points are preset, multiple amplitudes of the high-frequency band frequency domain signal are calculated according to the plurality of differential beamformers, and the differential beamformer with the smallest amplitude is selected as the optimal differential beamformer;

[0008] For the low-frequency band signal, calculate the low-frequency band frequency domain signal by using a beamformer with constrained white noise gain;

[0009] The frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer are transformed to obtain directionally picked-up speech.

[0010] As a preferred technical solution, the collecting of time domain signals from dual microphones and converting the time domain signals into frequency domain signals includes:

[0011] The time domain signals collected by the dual microphones are subjected to short-time Fourier transform to obtain frequency domain signals; wherein the short-time Fourier transform includes adopting a window function.

[0012] As a preferred technical solution, the beamformer constraining the white noise gain is specifically:

[0013] min w H (f)Γ(f)w(f)

[0014] stw H (f)d(f,0)=1

[0015]

[0016] Where Γ(f) is the spatial coherence moment of the diffuse noise model, which is calculated based on the microphone spacing, sound speed, and sampling frequency information; w(f) is the beamformer that constrains the white noise gain; w H (f) represents the conjugate transpose of w(f); WNG min (f) is the minimum white noise gain.

[0017] As a preferred technical solution, the preset several different zero points are set in the interval [90°, 180°].

[0018] As a preferred technical solution, for the high frequency band signal, a plurality of differential beamformers with different zero points are preset, specifically:

[0019]

[0020] Among them, w p (f) is the pth zero point α p Differential beamformer; d H (f, α p ) is d(f,α p ), d(f, α p ) indicates that the incident angle is α p The guiding vector.

[0021] As a preferred technical solution, the frequency domain signal calculated by the constrained white noise gain beamformer and the optimal differential beamformer is transformed to obtain a directionally picked-up speech output, including:

[0022] The frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer are subjected to inverse Fourier transform to obtain time domain signals, and overlap-addition is performed on the time domain signals to obtain a directionally picked-up speech output.

[0023] In a second aspect, the present application provides a directional sound pickup system based on a small-aperture dual-microphone array, which is applied to the directional sound pickup method based on a small-aperture dual-microphone array.

[0024] It includes a signal acquisition module, a high-frequency band frequency domain signal calculation module, a low-frequency band frequency domain signal calculation module, and a directional voice pickup module;

[0025] The signal acquisition module is used to acquire time domain signals from the dual microphones and convert the time domain signals into frequency domain signals; the frequency domain signals include low frequency band signals and high frequency band signals;

[0026] The high-frequency band frequency domain signal calculation module is configured to preset a plurality of differential beamformers with different zero points for the high-frequency band signal, calculate multiple amplitudes of the high-frequency band frequency domain signal based on the plurality of differential beamformers, and select the differential beamformer with the smallest amplitude as the optimal differential beamformer;

[0027] The module for calculating the low-frequency band frequency domain signal is configured to calculate the low-frequency band frequency domain signal for the low-frequency band signal by using a beamformer with constrained white noise gain;

[0028] The directional voice pickup module is used to transform the frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer to obtain directionally picked-up voice.

[0029] As a preferred technical solution, the signal acquisition module is specifically used to:

[0030] The time domain signals collected by the dual microphones are subjected to short-time Fourier transform to obtain frequency domain signals; wherein the short-time Fourier transform includes adopting a window function.

[0031] In a third aspect, the present application provides an electronic device, comprising:

[0032] at least one processor; and,

[0033] a memory communicatively connected to the at least one processor; wherein,

[0034] The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to perform the directional sound pickup method based on the small-aperture dual-microphone array.

[0035] In a fourth aspect, the present application provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the directional sound pickup method based on a small-aperture dual-microphone array.

[0036] In summary, compared with the prior art, the effective effects brought about by the technical solution provided by this application include at least:

[0037] The present application proposes a directional sound pickup method based on a small-aperture dual-microphone array, which collects time domain signals from the dual microphones and converts the time domain signals into frequency domain signals; the frequency domain signals include low-frequency band signals and high-frequency band signals; for the high-frequency band signals, several differential beamformers with different zero points are preset, and multiple amplitudes of the high-frequency band frequency domain signals are calculated according to the several differential beamformers, and the differential beamformer with the smallest amplitude is selected as the optimal differential beamformer; for the low-frequency band signals, the low-frequency band frequency domain signals are calculated by a beamformer with constrained white noise gain; the frequency domain signals calculated by the beamformer with constrained white noise gain and the optimal differential beamformer are transformed to obtain directionally picked up speech. In order to avoid the disadvantage of low white noise gain in the low-frequency band of the differential beamformer, the present application performs frequency band processing, adopting a constrained white noise gain beamformer in the low-frequency band and a differential beamformer in the high-frequency band; in addition, multiple differential beamformers with different zero points are preset, and the output of the beamformer with the smallest beam result amplitude is selected as the final output, thereby improving the directional sound pickup effect; therefore, the present application improves the degree of suppression of interference signals in other directions and effectively improves the directional sound pickup effect of the dual-microphone array. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 A flowchart of a directional sound pickup method based on a small-aperture dual-microphone array provided in one embodiment of the present application;

[0040] Figure 2 A schematic diagram of a desired direction provided for one embodiment of the present application;

[0041] Figure 3 A schematic diagram showing a spectrum comparison between the method used in the present application and a method using a super-directional beamformer, provided as an embodiment of the present application;

[0042] Figure 4 A block diagram of a directional sound pickup system based on a small-aperture dual-microphone array provided in accordance with one embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0044] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0045] Example:

[0046] See also Figure 1 In one embodiment of the present application, a directional sound pickup method based on a small-aperture dual-microphone array is provided, comprising the following steps:

[0047] S1. Collect time domain signals from two microphones and convert the time domain signals into frequency domain signals; the frequency domain signals include low frequency band signals and high frequency band signals.

[0048] Furthermore, the time domain signals collected by the dual microphones are subjected to short-time Fourier transform to obtain frequency domain signals; the time domain signals input in each frame are processed by frequency band division, so that different beamforming methods can be used in different frequency bands; in addition, in order to achieve better frequency consistency of the beam pattern, this embodiment uses a small-aperture dual-microphone array. The small aperture can be understood as a relatively small distance between the two microphones, for example, the distance between the two microphones can be 0.8 cm.

[0049] Among them, for short-time Fourier transform, a window function can be selected, the frame length can be set to F, and the frame shift can be set to 50% of the frame length. For the convenience of representation, it can be written in the form of a vector, that is, the signal of the fth frequency point in the tth frame can be expressed as:

[0050]

[0051] where x1(t, f)∈C and x2(t, f)∈C are the signals of microphone 1 and microphone 2 at the time-frequency point {t, f}, respectively, and f∈{0, 1, …, F / 2} is the frequency index. In the far-field environment, equation (1) can be expressed as:

[0052] x(t,f)=d(f,θ)s1(t,f)+v(t,f) (2)

[0053] In the above formula (2), is the steering vector of the sound source s with an incident azimuth angle of θ, where T represents the transpose, f s represents the sampling rate, s1(t, f) is the signal of the sound source at microphone 1, and v(t, f) is the additive noise signal.

[0054] For a dual-microphone array, the pickup direction of the differential beamformer can only be the end-fire direction (0° or 180°). The 0° desired direction is used as an example. Figure 2 As shown, the 0 degree direction is the direction from microphone 1 to microphone 2, or in other words, it is directed to microphone 1. For the target sound source in the 0 degree direction, its steering vector can be expressed as τ0 is the time delay of the sound propagating from microphone 1 to microphone 2. Therefore, a beamformer w(f)∈C can be designed 2 , C 2 is a complex vector identifier that suppresses noise signals (outputs frequency domain signals) while ensuring no distortion of the target sound source s:

[0055] y(t,f)=w H (f)x(t,f) (3)

[0056] In the above formula (3), H represents conjugate transposition.

[0057] S2. For the high-frequency band signal, preset several differential beamformers with different zero points, calculate multiple amplitudes of the high-frequency band frequency domain signal according to the several differential beamformers, and select the differential beamformer with the smallest amplitude as the optimal differential beamformer.

[0058] Furthermore, in order to improve the degree of suppression of interference signals from other directions, for frequency index f>f TH , that is, the high-frequency signal, selects the one with the smallest amplitude from the results of the differential beamformers with P (P ≥ 2) different zero points as the output. In other words, the beamformer with the smallest output amplitude is selected.

[0059] In this embodiment, the zero point of the differential beamformer can be set within the interval [90°, 180°], where the zero point refers to the point where the beam has the lowest degree of suppression on a signal in a certain direction.

[0060] Calculate the pth zero point as α p The differential beamformer w p (f), namely:

[0061]

[0062] In the above formula (4), w p (f) is the pth zero point α p Differential beamformer; d H (f,α p ) is d(f,α p ), d(f,α p ) indicates that the incident angle is α p The steering vector of the pth beam is Afterwards, find the index of the beamformer that has the smallest beamformation magnitude:

[0063]

[0064] Finally, select the pth * The beam result is output at this frequency point, namely:

[0065]

[0066] In this embodiment, a plurality of differential beamformers with different zero points (i.e., differential beamformer 1, ..., differential beamformer P) are preset, and a beam result with the minimum amplitude is ultimately output, thereby achieving the effect of zero point adaptation, improving the degree of suppression of interference signals, and enhancing the noise reduction effect.

[0067] S3. For the low-frequency band signal, calculate the low-frequency band frequency domain signal by using a beamformer with constrained white noise gain.

[0068] Furthermore, in order to avoid the disadvantage of the differential beamformer having a low white noise gain in the low frequency band, which will amplify the microphone self-noise, for the frequency index f≤f TH , that is, low-frequency signals, using a beamformer with constrained white noise gain; where f TH It can be set according to actual needs.

[0069] Specifically, a convex optimization toolbox such as CVX can be used to solve the following optimization problem to obtain a beamformer:

[0070]

[0071] In the above equation (7), w(f) is the beamformer with constrained white noise gain; w H (f) represents the conjugate transpose of w(f); WNG min (f) is the minimum white noise gain (this value can be set according to specific needs); Γ(f)∈R 2x2 is the spatial coherence matrix of the diffuse noise model, which can be calculated based on the microphone spacing, sound speed and sampling frequency f s The specific calculation formula of the {i, j}th element of the Γ(f) matrix is ​​calculated as follows:

[0072]

[0073] In the above formula, sin[·] represents the sine trigonometric function.

[0074] The white noise gain in this embodiment refers to the degree to which the beamformer suppresses the microphone self-noise (such as the thermal noise of the sensor). The larger the white noise gain, the better the suppression effect. When the white noise gain is negative, the microphone self-noise will be amplified.

[0075] In addition, for the low-frequency part in the frequency band processing, the use of a delay-sum beamformer or a robust super-directional beamformer (diagonally loaded super-directional beamformer) can also solve the problem of low white noise gain in the low-frequency band of the differential beam.

[0076] S4. Transforming the frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer to obtain directionally picked-up speech.

[0077] Furthermore, the frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer are subjected to an inverse Fourier transform to obtain a time domain signal, and the time domain signals are overlap-added to obtain a directionally picked-up speech output. The output of the beamformer (frequency domain signal) is a complex number, i.e., it includes amplitude and phase.

[0078] Specifically, to illustrate the technical advantages of the method of the present application, the spectrum of the method of the present application is compared with that of the super-directional beamformer method. Figure 3 ; Figure 3 The audio comparison 1 in the figure shows the spectrum of the reference channel, the audio comparison 2 shows the spectrum generated by the super-directional beamformer method, and the audio comparison 3 shows the spectrum generated by the method of the present application. In this example, a speaker is used to play white noise signals at different angles. In this experimental comparison, the pickup direction is 0°. For f>f THfrequency band (high frequency band signal), the zero points of the three differential beamformers are set to 90°, 150° and 180°, that is, P = 3, α1 = 90°, α2 = 150°, α3 = 180°; for f ≤ f TH Frequency band (low frequency signal), WNG min (f) Set according to actual needs. Figure 3 As can be seen from the 90°, 150° and 180° in the spectrum diagram, the method adopted in this application is compared with the method using the super-directional beamformer: (1) the degree of signal suppression in the undesired direction is higher, and (2) the degree of amplification of the microphone's own noise is lower; wherein, Figure 3 In the figure, the reference channel is the signal originally collected by the microphone. After passing through the super-directional beamformer, it can be seen that the amplitude of the low-frequency signal has increased a lot (this is manifested as the color of the low-frequency band on the spectrum graph becomes lighter), while this application has no obvious increase. In addition, Figure 3 The middle black vertical segment is a silent segment, which can be considered to contain only the microphone's own noise. This also shows that compared with the super-directional beam, the present application has a lower degree of amplification of the microphone's own noise. Overall, the method adopted in the present application has a better directional sound pickup effect.

[0079] To summarize, in order to avoid the disadvantage of low low-frequency white noise gain of the differential beamformer, this application uses a beamformer with constrained white noise gain in the low-frequency band and a differential beamformer in the high-frequency band. Several differential beamformers with different zero points are preset, and the beam result with the minimum amplitude is finally output, thereby achieving an effect similar to zero-point adaptation, improving the degree of suppression of interference signals, and improving the noise reduction effect, thereby improving the directional sound pickup effect.

[0080] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the described order of actions, because according to this application, certain steps can be performed in other orders or simultaneously.

[0081] Based on the same concept as the directional sound pickup method based on a small-aperture dual-microphone array in the above-mentioned embodiment, the present application also provides a directional sound pickup system based on a small-aperture dual-microphone array, which can be used to implement the above-mentioned directional sound pickup method based on a small-aperture dual-microphone array. For ease of explanation, the structural diagram of the embodiment of the directional sound pickup system based on a small-aperture dual-microphone array only shows the parts related to the embodiment of the present application. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the system, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0082] See also Figure 4In another embodiment of the present application, a directional sound pickup system based on a small-aperture dual-microphone array is provided, the system comprising a signal collection module 101, a high-frequency band frequency domain signal calculation module 102, a low-frequency band frequency domain signal calculation module 103, and a directional voice pickup module 104;

[0083] The signal acquisition module 101 is used to acquire time domain signals from dual microphones and convert the time domain signals into frequency domain signals; the frequency domain signals include low frequency band signals and high frequency band signals;

[0084] The high-frequency band frequency domain signal calculation module 102 is configured to preset a plurality of differential beamformers with different zero points for the high-frequency band signal, calculate multiple amplitudes of the high-frequency band frequency domain signal based on the plurality of differential beamformers, and select the differential beamformer with the smallest amplitude as the optimal differential beamformer;

[0085] The low-band frequency domain signal calculation module 103 is configured to calculate the low-band frequency domain signal for the low-band signal by using a beamformer with constrained white noise gain;

[0086] The directional speech pickup module 104 is configured to transform the frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer to obtain directionally picked-up speech.

[0087] It should be noted that the directional sound pickup system based on a small-aperture dual-microphone array of the present application corresponds one-to-one to the directional sound pickup method based on a small-aperture dual-microphone array of the present application. The technical features and beneficial effects described in the above-mentioned embodiment of the directional sound pickup method based on a small-aperture dual-microphone array are applicable to the embodiment of the directional sound pickup system based on a small-aperture dual-microphone array. For specific contents, please refer to the description in the embodiment of the method of the present application. No further details will be given here. This is hereby declared.

[0088] In addition, in the implementation of the directional sound pickup system based on a small-aperture dual-microphone array in the above-mentioned embodiment, the logical division of each program module is only an example. In actual applications, the above-mentioned functions can be assigned to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the directional sound pickup system based on a small-aperture dual-microphone array is divided into different program modules to complete all or part of the functions described above.

[0089] In another embodiment, an electronic device for implementing a directional sound pickup method based on a small-aperture dual-microphone array is provided, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the directional sound pickup method based on a small-aperture dual-microphone array of any embodiment of the present application is implemented.

[0090] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the device;

[0091] The device may be a computing device such as a desktop computer, a laptop, a PDA, a cloud server, etc. The device may include, but is not limited to, a processor, a memory;

[0092] The processor may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device, connecting various parts of the entire device using various interfaces and lines;

[0093] The memory can be used to store the computer programs and / or modules, and the processor realizes various functions of the device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; in addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0094] Accordingly, the present application also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the directional sound pickup method based on a small-aperture dual-microphone array as described in any of the above embodiments.

[0095] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0096] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] The above embodiments are preferred implementation modes of the present application, but the implementation modes of the present application are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present application should be considered as equivalent replacement methods and are included in the scope of protection of the present application.

Claims

1. A directional sound pickup method based on a small-aperture dual-microphone array, characterized in that: The steps include: Collecting time domain signals from dual microphones and converting the time domain signals into frequency domain signals; the frequency domain signals include low frequency band signals and high frequency band signals; For the high-frequency band signal, a plurality of differential beamformers with different zero points are preset, multiple amplitudes of the high-frequency band frequency domain signal are calculated according to the plurality of differential beamformers, and the differential beamformer with the smallest amplitude is selected as the optimal differential beamformer; For the low-frequency band signal, calculate the low-frequency band frequency domain signal by using a beamformer with constrained white noise gain; The frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer are transformed to obtain directionally picked-up speech.

2. The directional sound pickup method based on a small-aperture dual-microphone array according to claim 1, characterized in that: The collecting time domain signals of the dual microphones and converting the time domain signals into frequency domain signals includes: The time domain signals collected by the dual microphones are subjected to short-time Fourier transform to obtain frequency domain signals; wherein the short-time Fourier transform includes adopting a window function.

3. The directional sound pickup method based on a small-aperture dual-microphone array according to claim 1, characterized in that: The beamformer constrained by white noise gain is specifically: Where Γ(f) is the spatial coherence matrix of the diffuse noise model, which is calculated based on the microphone spacing, sound speed, and sampling frequency information; ω(f) is the beamformer that constrains the white noise gain; ω H (f) represents the conjugate transpose of ω(f); WNG min (f) is the minimum white noise gain.

4. The directional sound pickup method based on a small-aperture dual-microphone array according to claim 1, characterized in that: The preset multiple different zero points are set in the interval [90°, 180°].

5. The directional sound pickup method based on a small-aperture dual-microphone array according to claim 1, characterized in that: For the high frequency band signal, a plurality of differential beamformers with different zero points are preset, specifically: Among them, ω p (f) is the pth zero point α p Differential beamformer; d H (f, α p ) is d(f,α p ), d(f, α p ) indicates that the incident angle is α p The guiding vector.

6. The directional sound pickup method based on a small-aperture dual-microphone array according to claim 1, characterized in that: The step of transforming the frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer to obtain a directionally picked-up speech output includes: The frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer are subjected to inverse Fourier transform to obtain time domain signals, and overlap-addition is performed on the time domain signals to obtain a directionally picked-up speech output.

7. A directional sound pickup system based on a small-aperture dual-microphone array, characterized in that: A directional sound pickup method based on a small-aperture dual-microphone array applied to any one of claims 1-6, comprising a signal acquisition module, a high-frequency band frequency domain signal calculation module, a low-frequency band frequency domain signal calculation module, and a directional voice pickup module; The signal acquisition module is used to acquire time domain signals from the dual microphones and convert the time domain signals into frequency domain signals; the frequency domain signals include low frequency band signals and high frequency band signals; The high-frequency band frequency domain signal calculation module is configured to preset a plurality of differential beamformers with different zero points for the high-frequency band signal, calculate multiple amplitudes of the high-frequency band frequency domain signal based on the plurality of differential beamformers, and select the differential beamformer with the smallest amplitude as the optimal differential beamformer; The module for calculating the low-frequency band frequency domain signal is configured to calculate the low-frequency band frequency domain signal for the low-frequency band signal by using a beamformer with constrained white noise gain; The directional voice pickup module is used to transform the frequency domain signals calculated by the constrained white noise gain beamformer and the optimal differential beamformer to obtain directionally picked-up voice.

8. The directional sound pickup system based on a small-aperture dual-microphone array according to claim 7, characterized in that: The signal acquisition module is specifically used for: The time domain signals collected by the dual microphones are subjected to short-time Fourier transform to obtain frequency domain signals; wherein the short-time Fourier transform includes adopting a window function.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the directional sound pickup method based on a small-aperture dual-microphone array as described in any one of claims 1 to 6.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the directional sound pickup method based on a small-aperture dual-microphone array according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Double layer ring microphone array voice enhancement method

    CN108447499A

  • Pickup method and device and electronic equipment

    CN113393856A