Directional optional sound pickup method based on dual microphones, electronic device and storage medium
Through the directional optional sound pickup method of dual microphones, the target beam angle and phase adjustment are used, combined with zero-desist technology, the problem of poor interference suppression effect in complex environments is solved, and effective suppression of non-stationary noise and vocal interference and directional selection of sound pickup areas is achieved.
Patent Information
- Application Number
- CN202210417778.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-04-20
AI Technical Summary
The traditional single microphone noise reduction algorithm has poor noise reduction capabilities in non-stationary noise environments. The single microphone algorithm based on deep learning has poor effect on suppressing vocal interference and is highly complex. The traditional dual microphone solution has poor interference suppression effect in complex environments and is difficult to adjust the sound pickup direction in real time.
The directional optional sound pickup method based on the dual microphone is adopted. By obtaining the frequency domain signal of the dual microphone, using the target beam angle for phase adjustment, forming a directionally enhanced audio signal, combining the preset direction zero trap and control factor to achieve suppression of non-target direction noise.
It realizes effective suppression of non-stationary noise and vocal interference in complex environments, supports directional selection and real-time adjustment of sound picking areas, and improves the sound picking effect of voice signals.
Smart Images

Figure CN114708881B_ABST
Abstract
Description
Technical field
[0001] The embodiments of the present application relate to the field of smart terminal technology, and in particular to a directional optional sound pickup method based on dual microphones, an electronic device, and a storage medium. [Background Technology]
[0002] With the development of smartphones, wearable devices, and smart speakers, voice terminal devices equipped with at least two microphones have been widely used. Traditional single-microphone noise reduction algorithms have poor noise reduction capabilities for non-stationary noise. Single-microphone noise reduction algorithms based on deep learning can improve noise reduction performance in non-stationary noise scenarios, but due to the characteristics of the algorithm itself, the algorithm has poor performance in suppressing human voice interference and is more complex than traditional noise reduction algorithms. Therefore, providing a directional optional pickup method based on dual microphones that can adjust the pickup area in real time and suppress interference from non-target directions is an urgent problem to be solved by those skilled in the art. [Summary of the invention]
[0003] The embodiments of the present application provide a directional and optional sound pickup method based on dual microphones, an electronic device, and a storage medium, so that when directional sound pickup is required, the phase of the microphone signal can be adjusted through a phase control module to output a voice signal that supports a selectable sound pickup area and has good suppression capabilities against non-stationary noise and human voice interference.
[0004] In a first aspect, an embodiment of the present application provides a directional optional sound pickup method based on dual microphones, which is applied to an electronic device, wherein the electronic device includes a first microphone and a second microphone, and the method includes: obtaining a first frequency domain signal corresponding to the first microphone, and a second frequency domain signal corresponding to the second microphone; obtaining a target beam angle corresponding to the audio signal to be output; determining a first phase adjustment value corresponding to the first frequency domain signal and a second phase adjustment value corresponding to the second frequency domain signal based on the target beam angle; adjusting the phase of the first frequency domain signal according to the first phase adjustment value to obtain a third frequency domain signal, and adjusting the phase of the second frequency domain signal according to the second phase adjustment value to obtain a fourth frequency domain signal; and forming the audio signal to be output based on the third frequency domain signal and the fourth frequency domain signal.
[0005] In the above-mentioned directional optional sound pickup method based on dual microphones, after obtaining the first frequency domain signal and the second frequency domain signal corresponding to the dual microphones, the first frequency domain signal and the second frequency domain signal are phase-adjusted according to the target beam angle of the audio signal to be output to obtain the phase-adjusted third frequency domain signal and the fourth frequency domain signal, and then the audio signal to be output is formed based on the third frequency domain signal and the fourth frequency domain signal. Compared with the prior art, the present application determines the sound pickup area that needs to be directionally enhanced by determining the target beam angle, so as to support the optional directionality of the sound pickup area and output the directionally enhanced audio signal.
[0006] In one embodiment, obtaining the target beam angle corresponding to the to-be-output audio signal includes: responding to a sound pickup area selection instruction triggered by a user, obtaining the target beam angle corresponding to the sound pickup area selection instruction.
[0007] In one embodiment, the first phase adjustment value and the second phase adjustment value are obtained based on the following formula:
[0008]
[0009] Where θ1(k) is the first phase adjustment value, θ2(k) is the second phase adjustment value, α is the target beam angle, d is the distance between the two microphones, c is the speed of sound in air, and N is the length of the time window corresponding to the frequency domain signal.
[0010] In one implementation manner, the third frequency domain signal is obtained based on the following formula:
[0011] S p1 [k]=S1[k]*θ1(k)
[0012] Among them, S p1 [k] is the third frequency domain signal, S1[k] is the first frequency domain signal, and θ1(k) is the first phase adjustment value;
[0013] The fourth frequency domain signal is obtained based on the following formula:
[0014] S p2 [k]=S2[k]*θ2(k)
[0015] Among them, S p2 [k] is the fourth frequency domain signal, S2[k] is the second frequency domain signal, and θ2(k) is the second phase adjustment value.
[0016] In one embodiment, the corresponding audio signal is formed based on the third frequency domain signal and the fourth frequency domain signal, including: forming a frequency domain signal after zero-notching in a corresponding direction according to a preset direction based on the third frequency domain signal and the fourth frequency domain signal; obtaining the amplitude value of each frequency point in the target frequency domain signal based on the frequency domain signal after zero-notching in the corresponding direction to obtain the target frequency domain signal; and performing time domain conversion on the target frequency domain signal to form the audio signal to be output.
[0017] In one embodiment, the frequency domain signal after the corresponding direction nulling is obtained based on the following formula:
[0018] S w1 =[S p1 S p2 ]H1, S w2 =[Sp1 S p2 ]H2,S w3 =[S p1 S p2 ]H3
[0019] in, H1, H2, and H3 are weight vectors corresponding to 0-degree null sink, 90-degree null sink, and 180-degree null sink, respectively. w1 、S w2 、S w3 They are the frequency domain signals after zero sinking in the corresponding directions of 0 degrees, 90 degrees, and 180 degrees respectively.
[0020] In one embodiment, the target frequency domain signal is obtained based on the following formula:
[0021]
[0022] Among them, S a1 、S a2 、S a3 S w1 、S w2 、S w3 The corresponding amplitude spectrum, S ang is the phase spectrum of the first frequency domain signal, H L is the compensation filter factor, γ is the factor that controls the beam width of the sound pickup area, and β is the factor that controls the suppression strength of the non-sound pickup area.
[0023] In a second aspect, the present application provides an electronic device, which includes a first microphone and a second microphone, and the electronic device includes: a frequency domain signal acquisition module, used to obtain a first frequency domain signal corresponding to the first microphone, and a second frequency domain signal corresponding to the second microphone; a target beam angle acquisition module, used to obtain a target beam angle corresponding to the audio signal to be output; a phase adjustment value determination module, used to determine a first phase adjustment value corresponding to the first frequency domain signal and a second phase adjustment value corresponding to the second frequency domain signal based on the target beam angle; a frequency domain signal adjustment module, used to adjust the phase of the first frequency domain signal according to the first phase adjustment value to obtain a third frequency domain signal, and adjust the phase of the second frequency domain signal according to the second phase adjustment value to obtain a fourth frequency domain signal; an audio signal output module, used to form the audio signal to be output based on the third frequency domain signal and the fourth frequency domain signal.
[0024] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the above-mentioned directional optional sound pickup method based on dual microphones.
[0025] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the above-mentioned directional optional sound pickup method based on dual microphones.
[0026] It should be understood that the second to fourth aspects of the embodiments of the present application are consistent with the technical solutions of the first aspect of the embodiments of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated.
Brief Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 A schematic diagram of a flow chart of a directional and selective sound pickup method based on dual microphones provided in one embodiment of the present application;
[0029] Figure 2 A schematic diagram of the structure of a dual-microphone array provided in one embodiment of the present application, wherein: Figure 2 A is a schematic diagram of the structure of a dual-microphone array of a smartphone. Figure 2 B is a schematic diagram of the structure of the dual-microphone array of the watch;
[0030] Figure 3 A schematic diagram of a target beam angle provided for one embodiment of the present application;
[0031] Figure 4 The beam patterns under different beam width factors provided by one embodiment of the present application, wherein: Figure 4 A is the beam pattern when the beam width factor is 1, Figure 4 B is the beam pattern when the beam width factor is 5;
[0032] Figure 5 The beam patterns under different suppression strength factors provided by an embodiment of the present application, wherein: Figure 5 A is the beam pattern when the suppression strength factor is 0.01, Figure 5 B is the beam pattern when the suppression strength factor is 0.0001;
[0033] Figure 6 This is a spectrum diagram after suppressing the interference signal under different β parameters provided by one embodiment of the present application;
[0034] Figure 7A beam pattern of an audio signal to be output provided in one embodiment of the present application, wherein: Figure 7 A is the beam pattern of the audio signal to be output when the target beam angle is 0°, Figure 7 B is the beam pattern of the audio signal to be output when the target beam angle is 90°. Figure 7 C is the beam pattern of the audio signal to be output when the target beam angle is 180°;
[0035] Figure 8 A spectrum diagram when the target beam angle is 90° provided for one embodiment of the present application;
[0036] Figure 9 A schematic diagram of a flow chart of a directional and selective sound pickup method based on dual microphones provided in one embodiment of the present application;
[0037] Figure 10 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application;
[0038] Figure 11 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application;
[0039] Figure 12 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. [Specific implementation method]
[0040] In order to better understand the technical solutions of this specification, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0041] It should be clear that the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this specification.
[0042] The terms used in the examples of this application are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a," "an," "the," and "the" used in the examples of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0043] The dual-microphone directional selectable sound pickup method provided in the embodiments of the present application can be executed by an electronic device, which can be a terminal device such as a smartphone, a watch, a tablet computer, a PC, a laptop computer, etc. In an optional embodiment, a service program for executing the dual-microphone directional selectable sound pickup method can be installed on the electronic device.
[0044] In existing technology, dual microphones are widely used in various voice terminal products. Directional sound pickup often uses microphone array-based beamforming technology. Traditional fixed beamforming solutions have poor interference suppression when the number of microphones is small, and the beam pattern does not meet frequency invariance. First-order difference beamforming can achieve good frequency invariance through compensation filtering, but first-order difference can only null interference from a certain direction. In a complex environment, the interference direction is often unknown and the number is large, so the suppression effect is poor. As an alternative implementation, a dual-microphone solution using first-order difference combined with a generalized sidelobe canceller (GSC) outputs a noise estimation signal through a blocking matrix, while using fixed beamforming to enhance the speech signal. The noise estimate is used as a reference signal to adaptively cancel interference noise in the speech, and the final output is post-filtered. This solution can achieve better noise suppression than first-order difference, but the suppression strength of directional interference is uncontrollable, and the sound pickup direction is difficult to adjust in real time.
[0045] In response to the above technical issues, the following is a detailed explanation of this application:
[0046] Figure 1 A flow chart of a method for directional and selective sound pickup based on dual microphones provided in an embodiment of the present application is shown in the figure. The method for directional and selective sound pickup based on dual microphones may include the following steps:
[0047] Step S101 : obtaining a first frequency domain signal corresponding to a first microphone and a second frequency domain signal corresponding to a second microphone.
[0048] Optionally, the first microphone and the second microphone may be located on the same side of the electronic device, for example, Figure 2 As shown in FIG. 1 , the first microphone ( 101 ) and the second microphone ( 102 ) are both located at the bottom of the smartphone, or as shown in FIG. Figure 2 As shown in B, the first microphone (103) and the second microphone (104) are both located on the bottom side of the front of the watch.
[0049] Optionally, the signal first collected by the first microphone is a first time domain signal corresponding to the first frequency domain signal. The first time domain signal is framed and windowed, and then the time domain sub-signals of each frame are Fourier transformed respectively, and each time domain sub-signal of each frame is transformed into the frequency domain to obtain each frame frequency domain sub-signal corresponding to each frame time domain sub-signal, and each frame frequency domain sub-signal is integrated to form the first frequency domain signal. The integration process of each frame frequency domain sub-signal can be implemented based on the following formula:
[0050]
[0051] Where s1(n) is the first time domain signal, w(n) is the N-point window function, and the size of N corresponds to the time window length of the frequency domain analysis. Window functions include rectangular window (rectwin), sine window (Sine), Hanning window (Hanning), Hamming window (Hamming), Tukey window, etc.
[0052] Optionally, the signal first collected by the second microphone is a second time domain signal corresponding to the second frequency domain signal. The second time domain signal is framed and windowed, and then the time domain sub-signals of each frame are Fourier transformed respectively, and each time domain sub-signal of each frame is transformed into the frequency domain to obtain each frame frequency domain sub-signal corresponding to each frame time domain sub-signal. Then, the frequency domain sub-signals of each frame are integrated to form the second frequency domain signal. The integration process of the frequency domain sub-signals of each frame can be implemented based on the following formula:
[0053]
[0054] Where s2(n) is the second time domain signal, w(n) is the window function of N points, and the size of N corresponds to the time window length of the frequency domain analysis. Window functions include rectangular window (rectwin), sine window (Sine), Hanning window (Hanning), Hamming window (Hamming), Tukey window (Tukey), etc.
[0055] Step S102: Obtain a target beam angle corresponding to the audio signal to be output.
[0056] When two microphones receive audio signals from a sound source, due to the different distances between the sound source and the two microphones, there will be a time difference and intensity difference when the two microphones receive the audio signal. That is, the phase and amplitude values of the audio signals received by the two microphones are different. The electronic device can calculate the sound source position based on this difference after noise reduction and superposition processing. The sound source position can be expressed as the target beam angle. The schematic diagram of the target beam angle (α) is shown as follows Figure 3 shown.
[0057] Optionally, the target beam angle may be acquired in advance, or a pickup area selection request may be generated for the electronic device for the user to select the target pickup area, so that the electronic device may respond to the pickup area selection instruction triggered by the user and obtain the target beam angle corresponding to the pickup area selection instruction.
[0058] Optionally, the pickup areas included in the above pickup area selection request can be classified according to the collection scenario, for example, company group meeting, one-on-one meeting, face-to-face interview, etc., and can also be classified according to the direction of the sound source to be collected, for example, directly in front, diagonally left front, diagonally right front, directly left, directly right, etc.
[0059] Optionally, after the target pickup area is determined according to the pickup area selection instruction fed back by the user, the corresponding target beam angle can be determined according to the target pickup area. For example, when the pickup area is classified according to the collection scene, the target beam angles are: 0° or 180° for company group meetings, 90° for one-on-one meetings, and 90° for face-to-face interviews; when the pickup area is classified according to the direction of the sound source to be collected, the target beam angles are: 90° for the front, 135° for the diagonally left front, 45° for the diagonally right front, 180° for the left side, and 0° for the right side.
[0060] Step S103: determining a first phase adjustment value corresponding to the first frequency domain signal and a second phase adjustment value corresponding to the second frequency domain signal based on the target beam angle.
[0061] After the target beam angle is determined, the target beam angle is used as the final beam angle of the audio signal to be output, and the phases of the first frequency domain signal and the second frequency domain signal are adjusted according to the target beam angle, so that the audio signal output after the dual-channel frequency domain signals are denoised and superimposed is the audio signal to be output with the beam angle being the target beam angle.
[0062] Optionally, the first phase adjustment value and the second phase adjustment value may be obtained based on the following formula:
[0063]
[0064] Where θ1(k) is the first phase adjustment value, θ2(k) is the second phase adjustment value, α is the target beam angle, d is the distance between the two microphones, c is the speed of sound in air, which is generally 340 m / s, and N is the length of the time window corresponding to the frequency domain signal.
[0065] Optionally, after the value of θ1(k)-θ2(k) is calculated, the first phase adjustment value and the second phase adjustment value can be assigned according to the difference. For example, when the target beam angle α is 0°, that is, When , the first phase adjustment value is The second phase adjustment value is 0; when the target beam angle α is 90°, that is, θ1(k)-θ2(k)=0, the first frequency domain signal and the second frequency domain signal do not need to be phase adjusted, that is, the first phase adjustment value is 0, and the second phase adjustment value is 0; when the target beam angle α is 180°, that is When , the first phase adjustment value is The second phase adjustment value is 0.
[0066] In step S104 , the phase of the first frequency domain signal is adjusted according to the first phase adjustment value to obtain a third frequency domain signal, and the phase of the second frequency domain signal is adjusted according to the second phase adjustment value to obtain a fourth frequency domain signal.
[0067] Optionally, the third frequency domain signal may be obtained based on the following formula:
[0068] S p1 [k]=S1[k]*θ1(k)
[0069] Among them, S p1 [k] is the third frequency domain signal, S1[k] is the first frequency domain signal, and θ1(k) is the first phase adjustment value.
[0070] Optionally, the fourth frequency domain signal is obtained based on the following formula:
[0071] S p2 [k]=S2[k]*θ2(k)
[0072] Among them, S p2 [k] is the fourth frequency domain signal, S2[k] is the second frequency domain signal, and θ2(k) is the second phase adjustment value.
[0073] Step S105 : generating an audio signal to be output based on the third frequency domain signal and the fourth frequency domain signal.
[0074] Since the target beam angle is determined at this time, the corresponding time delay can be used to compensate the signals received by each microphone to align all signals. That is, one microphone is used as a reference point, and the signal received by another microphone is compensated in amplitude and phase. Finally, the signals are weighted and summed, thereby achieving the purpose of enhancing the audio signal in the target direction and suppressing noise interference in the non-target direction.
[0075] Optionally, the spatial characteristics of the signals collected by each microphone can be superimposed to output the audio signal to be output. In this case, the output target frequency domain signal is Y(m)=W H X(m), where X(m) is the far-field steering vector of each microphone, used to characterize the spatial characteristics of each microphone, and W H The weight vector W is a conjugate device. The value of the weight vector W is related to the far-field steering vector of each microphone. Beamforming uses the weight vector W to add appropriate delay compensation Y(m) to the signal, ensuring that the phase and amplitude of the signals received by all microphones are consistent. The signals received by all microphones are summed to form a beam in the direction of the incoming wave, thereby enhancing the audio signal. The weight vectors are used to make the direction of any interfering signals opposite to the output audio signal, thereby suppressing the interference source.
[0076] For example, the above step S105 may include the following steps:
[0077] Step S1051: generating a frequency domain signal after being nulled in a corresponding direction according to a preset direction based on the third frequency domain signal and the fourth frequency domain signal, wherein the preset direction is determined by a target beam angle.
[0078] The purpose of forming a zero notch is to eliminate interference or noise in a specified direction. After the target beam angle is determined in advance, interference noise in directions other than the target beam angle is removed. After the noise is removed, delay compensation is performed and then a summation operation is performed.
[0079] Optionally, the preset directions may be 0 degrees, target beam angle α, and 180 degrees, that is, the weight vectors H1, H2, and H3 corresponding to the three directions are:
[0080]
[0081] For example, when the target beam angle α is 90°, the weight vectors H1, H2, and H3 in the three directions are:
[0082]
[0083] Then the frequency domain signals after weight vector zero-stuck are:
[0084] S w1 =[S p1 S p2 ]H1, S w2 =[S p1 S p2 ]H2,S w3 =[S p1 S p2 ]H3
[0085] Step S1052 : obtaining the amplitude value of each frequency point in the target frequency domain signal based on the frequency domain signal after nulling in the corresponding direction, and obtaining the target frequency domain signal.
[0086] Since the suppression strength of interference noise in the prior art is uncontrollable, it will affect the ideal strength value of the output audio signal to a certain extent. Therefore, a factor for controlling the suppression strength needs to be introduced to control the suppression strength of directional interference.
[0087] Optionally, after calculating the frequency domain signals after the weight vectors in three directions are zero-stuck, calculate S respectively. w1 、S w2 、S w3 The amplitude spectrum of S a1 、S a2 、S a3 , that is, the output target frequency domain signal S out(k) is:
[0088]
[0089] Among them, S a1 、S a2 、S a3 S w1 、S w2 、S w3 The corresponding amplitude spectrum, S ang is the phase spectrum of the first frequency domain signal, H L is the compensation filter factor, γ is the factor that controls the beam width of the sound pickup area (γ≥1), and β is the factor that controls the suppression strength of the non-sound pickup area (0<β<<1).
[0090] After the introduction of the above-mentioned γ factor, those skilled in the art and users can adjust the beam width of the pickup area, for example, Figure 4 As shown in the figure, when the sampling rate of the sound source is 16kHz and the frequency range of the spectrum analysis is 0-8kHz, after adjusting the phase and amplitude of the dual-channel frequency domain signal for the output audio signal with a target beam angle of 90°, the following can be formed when γ=1. Figure 4 The beam pattern shown in A can be formed when γ=5. Figure 4 The beam pattern shown in B.
[0091] After the introduction of the above-mentioned β factor, those skilled in the art and users can adjust the suppression strength of the non-pickup area, for example, Figure 5 As shown in the figure, when the sampling rate of the sound source is 16kHz and the frequency range of the spectrum analysis is 0-8kHz, after adjusting the phase and amplitude of the dual-channel frequency domain signal for the output audio signal with a target beam angle of 90°, the following can be formed when β=0.01. Figure 5 The beam pattern shown in A can be formed when β = 0.0001. Figure 5 The beam pattern shown in B. In addition, the suppression strength of the non-pickup area can also be controlled accordingly. For example, when the target beam angle is selected as 90°, the audio signal with a beam angle of 0° is regarded as an interference signal. The spectrum diagram of the interference signal under different β parameters is as follows: Figure 6 As shown, at this time, β can be adjusted to output audio signals with different suppression strengths.
[0092] Step S1053 : performing time domain conversion on the target frequency domain signal to form an audio signal to be output.
[0093] Optionally, the target frequency domain signal may be converted into the audio signal to be output by using an inverse Fourier transform, as shown in the following formula:
[0094]
[0095] Optionally, the audio signal to be outputted at this time is represented by a beam corresponding to the target beam angle. For example, when the target beam angle is 0°, the beam corresponding to the audio signal to be outputted at this time is a beam with a beam angle of 0°. Figure 7 As shown in A; when the target beam angle is 90°, the beam corresponding to the audio signal to be output is the beam with a beam angle of 90°, as shown in Figure 7 As shown in B; when the target beam angle is 180°, the beam corresponding to the audio signal to be output is the beam with a beam angle of 180°, as shown in Figure 7 As shown in Figure B. Moreover, at this time, the interference signals other than the target beam angle are effectively suppressed. For example, when the target beam angle is 90°, the 0°, 45°, 135° and 180° angles other than 90° are all interference signals. The corresponding spectrum is shown in Figure 1. Figure 8 shown.
[0096] In the above-mentioned directional optional sound pickup method based on dual microphones, after obtaining the first frequency domain signal and the second frequency domain signal corresponding to the dual microphones, the first frequency domain signal and the second frequency domain signal are phase-adjusted according to the target beam angle of the audio signal to be output to obtain the phase-adjusted third frequency domain signal and the fourth frequency domain signal, and then the audio signal to be output is formed based on the third frequency domain signal and the fourth frequency domain signal. Compared with the prior art, the present application determines the sound pickup area that needs to be directionally enhanced by determining the target beam angle, so as to support the optional directionality of the sound pickup area and output the directionally enhanced audio signal.
[0097] Figure 9 A flow chart of a method for directional and selective sound pickup based on dual microphones provided in an embodiment of the present application is shown in the figure. The method for directional and selective sound pickup based on dual microphones may include the following steps:
[0098] Step S201: obtaining a first time domain signal (201) of a first microphone and a second time domain signal (202) of a second microphone, wherein the time domain signals are audio signals collected by the microphones.
[0099] Step S202: The first time domain signal (201) and the second time domain signal (202) are subjected to frame-based windowing and frequency domain conversion to obtain a first frequency domain signal (203) and a second frequency domain signal (204).
[0100] Step S203 , in response to the sound pickup area selection instruction triggered by the user, determining the sound pickup area corresponding to the sound pickup area selection instruction, obtaining the target beam angle corresponding to the sound pickup area, and determining that the target beam angle is 90°.
[0101] Step S204, calculating the phase shift to be added to the first frequency domain signal (203) and the second frequency domain signal (204) according to the target beam angle. When the target beam angle is 90°, the formula It is determined that the first frequency domain signal and the second frequency domain signal do not need to be phase shifted or need to be phase shifted the same, thereby obtaining a third frequency domain signal (205) and a fourth frequency domain signal (206).
[0102] Step S205, using the third frequency domain signal and the fourth frequency domain signal to perform null-notch beamforming, the three weight vectors correspond to the directions of 0°, 90° and 180° respectively, and the frequency domain signals after null-notch in the corresponding directions are obtained according to the weight vectors in the three directions, which are the frequency domain signal after the first null-notch (207), the frequency domain signal after the second null-notch (208) and the frequency domain signal after the third null-notch (209).
[0103] In step S206, the frequency domain signal after the first nulling (207), the frequency domain signal after the second nulling (208) and the frequency domain signal after the third nulling (209) are subjected to spectrum subtraction compensation filtering to obtain a target frequency domain signal (210).
[0104] Step S207 , performing time domain conversion and windowing synthesis on the target frequency domain signal to obtain an audio signal to be output ( 211 ).
[0105] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device includes a first microphone and a second microphone. As shown in the figure, the electronic device 30 may include:
[0106] A frequency domain signal acquisition module 301 is configured to obtain a first frequency domain signal corresponding to the first microphone and a second frequency domain signal corresponding to the second microphone;
[0107] The target beam angle acquisition module 302 is used to acquire the target beam angle corresponding to the audio signal to be output;
[0108] A phase adjustment value determining module 303 is configured to determine a first phase adjustment value corresponding to the first frequency domain signal and a second phase adjustment value corresponding to the second frequency domain signal based on a target beam angle;
[0109] a frequency domain signal adjustment module 304 configured to adjust the phase of the first frequency domain signal according to the first phase adjustment value to obtain a third frequency domain signal, and to adjust the phase of the second frequency domain signal according to the second phase adjustment value to obtain a fourth frequency domain signal;
[0110] The audio signal output module 305 is configured to generate an audio signal to be output based on the third frequency domain signal and the fourth frequency domain signal.
[0111] In one embodiment, the target beam angle acquisition module 302 includes:
[0112] The selection instruction response submodule is used to respond to the sound pickup area selection instruction triggered by the user and obtain the target beam angle corresponding to the sound pickup area selection instruction.
[0113] In one embodiment, the phase adjustment value determination module 303 may be operated according to the following formula:
[0114]
[0115] Where θ1(k) is the first phase adjustment value, θ2(k) is the second phase adjustment value, α is the target beam angle, d is the distance between the two microphones, c is the speed of sound in air, and N is the length of the time window corresponding to the frequency domain signal.
[0116] In one embodiment, the frequency domain signal adjustment module 304 may operate according to the following formula:
[0117] S p1 [k]=S1[k]*θ1(k)
[0118] Among them, S p1 [k] is the third frequency domain signal, S1[k] is the first frequency domain signal, and θ1(k) is the first phase adjustment value;
[0119] S p2 [k]=S2[k]*θ2(k)
[0120] Among them, S p2 [k] is the fourth frequency domain signal, S2[k] is the second frequency domain signal, and θ2(k) is the second phase adjustment value.
[0121] In one embodiment, the audio signal output module 305 may include:
[0122] A frequency domain signal forming submodule after nulling is used to form a frequency domain signal after nulling in a corresponding direction according to a preset direction based on the third frequency domain signal and the fourth frequency domain signal;
[0123] The target frequency domain signal acquisition submodule is used to obtain the amplitude value of each frequency point in the target frequency domain signal based on the frequency domain signal after nulling in the corresponding direction, thereby obtaining the target frequency domain signal;
[0124] The time domain conversion submodule is used to perform time domain conversion on the target frequency domain signal to form an audio signal to be output.
[0125] In one embodiment, the frequency domain signal forming submodule after nulling can be operated based on the following formula:
[0126] Sw1 =[S p1 S p2 ]H1, S w2 =[S p1 S p2 ]H2,S w3 =[S p1 S p2 ]H3
[0127] in, H1, H2, and H3 are weight vectors corresponding to 0-degree null sink, 90-degree null sink, and 180-degree null sink, respectively. w1 、S w2 、S w3 They are the frequency domain signals after zero sinking in the corresponding directions of 0 degrees, 90 degrees, and 180 degrees respectively.
[0128] In one embodiment, the target frequency domain signal acquisition submodule may be operated based on the following formula:
[0129]
[0130] Among them, S a1 、S a2 、S a3 S w1 、S w2 、S w3 The corresponding amplitude spectrum, S ang is the phase spectrum of the first frequency domain signal, H L is the compensation filter factor, γ is the factor that controls the beam width of the sound pickup area, and β is the factor that controls the suppression strength of the non-sound pickup area.
[0131] Figure 11 A structural diagram of an electronic device provided in an embodiment of the present application is shown. As shown in the figure, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C and an internal memory 121, etc.
[0132] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0133] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0134] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0135] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0136] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0137] The processor 110 executes various functional applications and data processing by running the programs stored in the internal memory 121, such as implementing the present invention. Figures 1 to 9 The illustrated embodiment provides a directional selectable sound pickup method based on two microphones.
[0138] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0139] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0140] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0141] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0142] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0143] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0144] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0145] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0146] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0147] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.
[0148] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.
[0149] like Figure 12 As shown, the embodiment of the present application also provides a structural diagram of an electronic device, which may include at least one processor; and at least one memory in communication with the above-mentioned processor, wherein: the memory stores program instructions that can be executed by the processor, and the above-mentioned processor calls the above-mentioned program instructions to execute the instructions of this specification. Figures 1 to 9 The illustrated embodiment provides a directional selectable sound pickup method based on two microphones.
[0150] The present application also provides a computer-readable storage medium having computer instructions stored thereon. When executed by a processor, the instructions implement the steps of the above-described dual-microphone-based directional selective sound pickup method. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0151] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0152] In the description of the embodiments of the present invention, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0153] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of this specification includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of this specification belong.
[0154] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0155] It should be noted that the terminals involved in the embodiments of the present application may include but are not limited to personal computers (PCs), personal digital assistants (PDAs), wireless handheld devices, tablet computers, mobile phones, MP3 players, MP4 players, etc.
[0156] In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms.
[0157] In addition, the functional units in the various embodiments of this specification may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional units.
[0158] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the methods of various embodiments of this specification. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0159] The above are only preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A directional optional sound pickup method based on dual microphones, characterized in that: Applied to an electronic device, the electronic device includes a first microphone and a second microphone, and the method includes: Obtaining a first frequency domain signal corresponding to the first microphone and a second frequency domain signal corresponding to the second microphone; Obtaining a target beam angle corresponding to the audio signal to be output; Determining a first phase adjustment value corresponding to the first frequency domain signal and a second phase adjustment value corresponding to the second frequency domain signal based on the target beam angle; adjusting the phase of the first frequency domain signal according to the first phase adjustment value to obtain a third frequency domain signal, and adjusting the phase of the second frequency domain signal according to the second phase adjustment value to obtain a fourth frequency domain signal; forming the to-be-output audio signal based on the third frequency domain signal and the fourth frequency domain signal; The forming of the to-be-output audio signal based on the third frequency domain signal and the fourth frequency domain signal includes: Generating a frequency domain signal after nulling in a corresponding direction according to a preset direction based on the third frequency domain signal and the fourth frequency domain signal; Obtaining the amplitude value of each frequency point in the target frequency domain signal based on the frequency domain signal after nulling in the corresponding direction, thereby obtaining the target frequency domain signal; The target frequency domain signal is converted into the time domain to form the audio signal to be output.
2. The method according to claim 1, wherein The obtaining of a target beam angle corresponding to the audio signal to be output includes: In response to a sound pickup area selection instruction triggered by a user, a target beam angle corresponding to the sound pickup area selection instruction is obtained.
3. The method according to claim 1, wherein The first phase adjustment value and the second phase adjustment value are obtained based on the following formula: Where θ1(k) is the first phase adjustment value, θ2(k) is the second phase adjustment value, α is the target beam angle, d is the distance between the two microphones, c is the speed of sound in air, and N is the length of the time window corresponding to the frequency domain signal.
4. The method according to claim 1, wherein The third frequency domain signal is obtained based on the following formula: S p1 [k]=S1[k]*θ1(k) Among them, S p1 [k] is the third frequency domain signal, S1[k] is the first frequency domain signal, and θ1(k) is the first phase adjustment value; The fourth frequency domain signal is obtained based on the following formula: S p2 [k]=S2[k]*θ2(k) Among them, S p2 [k] is the fourth frequency domain signal, S2[k] is the second frequency domain signal, and θ2(k) is the second phase adjustment value.
5. The method according to claim 1, wherein The frequency domain signal after the corresponding direction nulling is obtained based on the following formula: S w1 =[S p1 S p2 ]H1,S w2 =[S p1 S p2 ]H2,S w3 =[S p1 S p2 ]H3 in, H1, H2, and H3 are weight vectors corresponding to 0-degree null sink, 90-degree null sink, and 180-degree null sink, respectively. w1 、S w2 、S w3 They are the frequency domain signals after zero sinking in the corresponding directions of 0 degrees, 90 degrees, and 180 degrees respectively.
6. The method according to claim 5, wherein The target frequency domain signal is obtained based on the following formula: Among them, S a1 、S a2 、S a3 S w1 、S w2 、S w3 The corresponding amplitude spectrum, S ang is the phase spectrum of the first frequency domain signal, H L is the compensation filter factor, γ is the factor that controls the beam width of the sound pickup area, and β is the factor that controls the suppression strength of the non-sound pickup area.
7. An electronic device, characterized in that: The electronic device includes a first microphone and a second microphone, and the electronic device includes: a frequency domain signal obtaining module, configured to obtain a first frequency domain signal corresponding to the first microphone and a second frequency domain signal corresponding to the second microphone; A target beam angle acquisition module is used to obtain a target beam angle corresponding to the audio signal to be output; a phase adjustment value determining module, configured to determine a first phase adjustment value corresponding to the first frequency domain signal and a second phase adjustment value corresponding to the second frequency domain signal based on the target beam angle; a frequency domain signal adjustment module, configured to adjust the phase of the first frequency domain signal according to the first phase adjustment value to obtain a third frequency domain signal, and adjust the phase of the second frequency domain signal according to the second phase adjustment value to obtain a fourth frequency domain signal; an audio signal output module, configured to form the to-be-output audio signal based on the third frequency domain signal and the fourth frequency domain signal; The audio signal output module includes: a frequency domain signal forming submodule after nulling, configured to form a frequency domain signal after nulling in a corresponding direction according to a preset direction based on the third frequency domain signal and the fourth frequency domain signal; A target frequency domain signal acquisition submodule is used to obtain the amplitude value of each frequency point in the target frequency domain signal based on the frequency domain signal after nulling in the corresponding direction, so as to obtain the target frequency domain signal; The time domain conversion submodule is used to perform time domain conversion on the target frequency domain signal to form the audio signal to be output.
8. An electronic device, characterized in that: include: at least one processor; as well as at least one memory in communication with the processor, wherein: The memory stores program instructions that can be executed by the processor, and the processor can execute the method according to any one of claims 1 to 6 by calling the program instructions.
9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Directional pickup method and device under double microphones
CN111863018A
Audio signal processing method and device, and storage medium
CN113362847A