A directional sound pickup method, device and electronic equipment

By performing phase compensation and noise estimation on the microphone signal in the frequency domain, the problem of insufficient noise suppression under sampling rate limitations in traditional multi-microphone noise reduction algorithms is solved, achieving a more efficient environmental noise suppression effect.

CN115798503BActive Publication Date: 2026-02-03SHANGHAI FULLHAN MICROELECTRONICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211185768.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-02-03
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Traditional multi-microphone noise reduction algorithms struggle to accurately align microphone signals under sampling rate limitations, resulting in poor environmental noise suppression, especially in complex acoustic environments where they fail to meet specific requirements.

Method used

By performing linear compensation on the microphone phase in the frequency domain, the microphone signals in the target direction are made to be in phase. The real and imaginary parts with the smallest absolute values ​​of the aligned frequency domain signals are used as the output. Combined with the differential signal as an environmental noise estimate, noise is further suppressed.

Benefits of technology

The noise suppression effect of the delay and sum algorithm has been improved, theoretically able to suppress more than 6dB of environmental noise and significantly improve the signal-to-noise ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798503B_ABST
    Figure CN115798503B_ABST
Patent Text Reader

Abstract

The application discloses a directional sound pickup method and device and electronic equipment, and the method comprises the following steps: in step S1, a microphone frequency domain signal phase difference is calculated according to a target sound source azimuth angle, and the phase is compensated; in step S2, the compensated microphone signals are summed to obtain a first enhanced signal of a target signal. The application can break through the limitation of a sampling rate by linearly compensating the microphone phase in the frequency domain, so that the phases of the two microphone signals of the target direction are completely consistent, and the effect of a delay and sum algorithm is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, and in particular to a directional sound pickup method, device, and electronic device. Background Technology

[0002] Speech enhancement technology, generally speaking, refers to noise reduction processing of speech signals, and it has a wide range of applications in today's rapidly developing communication technology. Traditional speech noise reduction algorithms are divided into single-microphone noise reduction algorithms and multi-microphone noise reduction algorithms: single-microphone noise reduction algorithms are relatively mature and can easily suppress stationary noise, but non-stationary noise usually requires additional feature extraction for separate suppression; multi-microphone noise reduction algorithms utilize spatial information, and compared with single-microphone algorithms, they can suppress not only stationary noise but also non-stationary noise in specific directions. This invention mainly discusses improvements to multi-microphone noise reduction algorithms.

[0003] Multi-microphone arrays can utilize beamforming algorithms to amplify sound from a target direction while suppressing sound from other directions, offering a natural advantage for noise reduction tasks in complex acoustic environments. Beamforming algorithms leverage the phase information generated by the same sound source arriving at different microphones to amplify sound from a specific direction. For example... Figure 1 As shown, taking two microphones as an example, based on the target direction The time difference of arrival can be determined by the distance L between microphones A and B. .

[0004] (1)

[0005] Where c is the speed of sound. When the time difference...

[0006] After confirmation, a commonly used voice enhancement method is delay and sum, which involves aligning the microphone signals and then adding them together, such as... Figure 2 As shown. In the case of two microphones, this algorithm can theoretically suppress 3dB of ambient noise; the more microphones, the stronger the suppression of ambient noise, and there is no risk to the target signal. However, it still has the following drawbacks:

[0007] (1) At typical sampling rates (such as 8k, 16k, 32k, 48k), it is difficult to accurately align two microphone signals. For example, for an 8k sampling rate, the minimum time interval between sampling points is 1 / 8000 seconds. If the time difference calculated by equation (1) is not an integer multiple of this value, the two microphone signals cannot be perfectly aligned, which in turn affects the effectiveness of the algorithm.

[0008] (2) The algorithm can theoretically suppress 3dB of ambient noise in the case of two microphones. This suppression strength is not high and may not meet certain specific needs. Summary of the Invention

[0009] To overcome the shortcomings of the existing technology, the present invention provides a directional sound pickup method, device and electronic device. By performing linear compensation on the microphone phase in the frequency domain, the limitation of sampling rate can be broken, and the phases of the two microphone signals in the target direction can be made completely consistent, which is beneficial to improving the performance of delay and sum algorithm.

[0010] Another objective of this invention is to provide a directional sound pickup method and apparatus. Furthermore, by taking the real and imaginary parts with the smallest absolute values ​​of the real and imaginary parts of the aligned frequency domain signal as the output, the problem of insufficient environmental noise suppression in the delay and sum algorithm is compensated.

[0011] Another objective of this invention is to provide a directional sound pickup method and apparatus, which proposes a scheme to further suppress noise by using the differential signal of the aligned signal as an estimate of environmental noise, thereby achieving more thorough suppression of environmental noise.

[0012] To achieve the above objectives, the present invention provides a directional sound pickup method, comprising the following steps:

[0013] Step S1: Calculate the phase difference of the microphone frequency domain signal based on the azimuth angle of the target sound source, and compensate for the phase.

[0014] Step S2: Sum the compensated microphone signals to obtain the first enhanced signal of the target signal.

[0015] Optionally, the method further includes:

[0016] Step S3: Compare the real and imaginary parts of the compensated microphone frequency domain signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the second enhanced signal of the target signal.

[0017] Step S4: Compare the real and imaginary parts of the first and second enhanced signals of the target signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the third enhanced signal of the target signal.

[0018] Optionally, the method further includes:

[0019] Step S5: Differentiate the compensated microphone frequency domain signal to obtain the environmental noise estimation signal;

[0020] Step S6: Using the environmental noise estimation signal as a noise reference signal, filter the third enhanced signal to obtain the filtered signal;

[0021] Step S7: Perform time-frequency transformation on the filtered signal to obtain the final time-domain output signal.

[0022] Optionally, step S1 includes:

[0023] Step S101: For any two microphone signals, convert the microphone signals to the frequency domain and convert the frequency domain signals to polar coordinates.

[0024] Step S102: Calculate the phase difference between the two frequency domain signals based on the location of the target sound source, and compensate for the phase difference to ensure that the phases of the target signals in each microphone are completely consistent.

[0025] Optionally, in step S3, the real part of each frequency point of the first microphone spectrum signal is compared with the real part of the corresponding frequency point of the second microphone compensation signal, and the real part with the smaller absolute value is taken as the real part of the corresponding frequency point of the second enhancement signal; the imaginary part of each frequency point of the first microphone spectrum signal is compared with the imaginary part of the corresponding frequency point of the second microphone compensation signal, and the imaginary part with the smaller absolute value is taken as the imaginary part of the corresponding frequency point of the second enhancement signal.

[0026] Optionally, in step S5, the first microphone spectrum signal and the compensation signal of the second microphone are differentially divided to obtain the environmental noise estimation signal.

[0027] Optionally, in step S7, the filtered signal is subjected to time-frequency transformation and then windowed and synthesized to obtain the final time-domain output signal.

[0028] Optionally, the method further includes:

[0029] Step S8: When the number of microphones is greater than two, select one microphone as the reference, and use the above steps S1-S7 to obtain multiple time-domain output signals for it and other microphones respectively, and combine the multiple time-domain output signals to obtain the final time-domain output signal.

[0030] To achieve the above objectives, the present invention also provides a directional sound pickup device, comprising:

[0031] The time-frequency conversion and phase compensation unit is used to calculate the phase difference of the microphone frequency domain signal based on the azimuth angle of the target sound source and to compensate for the phase.

[0032] The summation unit is used to sum the compensated microphone signals to obtain the first enhanced signal of the target signal;

[0033] The second enhanced signal determination unit is used to compare the real part and imaginary part of the compensated microphone frequency domain signal, and take the real part and imaginary part with the smallest absolute value to obtain the second enhanced signal of the target signal;

[0034] The third enhanced signal determination unit is used to compare the real and imaginary parts of the first enhanced signal and the second enhanced signal of the target signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the third enhanced signal of the target signal.

[0035] The differential unit is used to perform differential analysis on the compensated microphone frequency domain signal to obtain the environmental noise estimation signal;

[0036] The filtering unit is used to use the environmental noise estimation signal as a noise reference signal to filter the third enhanced signal to obtain the filtered signal.

[0037] The time-frequency conversion unit is used to perform time-frequency conversion on the filtered signal to obtain the final time-domain output signal.

[0038] Compared with existing technologies, this invention provides a directional sound pickup method, device, and electronic device. By performing linear compensation of the microphone phase in the frequency domain, it overcomes the limitations of the sampling rate, ensuring that the phases of the two microphone signals in the target direction are completely aligned, which is beneficial for improving the performance of the delay and sum algorithm. Furthermore, this invention proposes to use the real and imaginary parts with the smallest absolute values ​​of the aligned frequency domain signal as the output, compensating for the insufficient environmental noise suppression of the delay and sum algorithm. This invention also proposes a scheme to use the differential signal of the aligned signal as an environmental noise estimate to further suppress noise, achieving a more thorough suppression of environmental noise.

[0039] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0040] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same parts or steps.

[0041] Figure 1 This is a schematic diagram of the phase difference of a sound source in existing technology;

[0042] Figure 2 This is an enhanced diagram of delay and sum;

[0043] Figure 3 This is a flowchart illustrating a directional sound pickup method provided in an exemplary embodiment of the present invention.

[0044] Figure 4 This is a system structure diagram of a directional sound pickup device provided in an exemplary embodiment of the present invention;

[0045] Figure 5 This is a flowchart of the directional sound pickup method provided in the embodiments of the present invention;

[0046] Figure 6 This is the structure of an electronic device provided in an exemplary embodiment of the present invention. Detailed Implementation

[0047] The following describes the embodiments of the present invention through specific examples and in conjunction with the accompanying drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific examples, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0048] Exemplary method

[0049] Figure 3 This is a schematic flowchart of a directional sound pickup method provided in an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices, such as... Figure 3 As shown, it includes the following steps:

[0050] Step S1: For the two microphone signals, calculate the phase difference of the microphone frequency domain signals based on the azimuth angle of the target sound source, and compensate for the phase.

[0051] Specifically, step S1 further includes:

[0052] Step S101: Select two microphone signals, convert the microphone signals to the frequency domain, and convert the frequency domain signals to polar coordinates.

[0053] In this invention, the number of directional microphones should be no less than two. Generally speaking, the more microphones there are, the better the suppression of ambient noise. For ease of explanation, this embodiment uses a dual-microphone setup as an example to illustrate the directional microphone pickup principle of this invention:

[0054] Time domain signals from two microphones , ( Corresponding to the first microphone, The frequency domain signal is obtained by performing frame segmentation and Discrete Fourier Transform (DFT) on the signal corresponding to the second microphone. , ,Signal and Each frequency point is divided into a real part and an imaginary part, which can be converted into polar coordinates:

[0055] (2)

[0056] (3)

[0057] (4)

[0058] (5)

[0059] in, , This represents the frequency domain amplitude of the two microphones. , This indicates the frequency domain angle between the two microphones.

[0060] Step S102: Calculate the phase difference between the two frequency domain signals based on the location of the target sound source, and compensate for the phase difference to ensure that the phases of the target signals in each microphone are completely consistent.

[0061] In this embodiment, if the two microphones are highly consistent and random noise from the microphones is excluded, the following relationship should hold:

[0062] (6)

[0063] (7)

[0064] in, Indicates frequency point index, It is a constant.

[0065] Equation (6) indicates that the amplitude of each frequency point of the two microphones is the same, and equation (7) indicates that the phase difference of each frequency point of the two microphones is linear. If b is 0, it means that the phase of the signals from the two microphones is completely consistent, that is, the sound source arrives at the two microphones at the same time. The larger the absolute value of b, the larger the phase difference between the signals from the two microphones, that is, the greater the delay of the sound source arriving at the two microphones. And the value of b corresponds one-to-one with the delay.

[0066] Therefore, when the azimuth angle of the target sound source Once determined, the value of b is uniquely determined.

[0067] When the delay is At that time, the phase difference is:

[0068] (8)

[0069] There is also a phase difference.

[0070] (9)

[0071] in, Where is the sampling rate, and N is the FFT length.

[0072] Combining equations (8) and (9), we can obtain

[0073] (10)

[0074] Where L is the distance between the two microphones.

[0075] In equation (10), b is the slope of the phase difference between the two microphone signals in the frequency domain after the position of the sound source is fixed. The accuracy of this slope is not limited in any way.

[0076] The slope is now converted into a phase difference, and the phase difference between the two microphone signals is compensated:

[0077] (11)

[0078] (12)

[0079] in This represents the phase-compensated spectrum of the second microphone, including both real and imaginary parts. Generally, for the spectrum in the target direction, Phase with the first microphone They are consistent, but generally inconsistent for non-target signals.

[0080] In other words, in this embodiment, the first microphone 1 It's fixed; you only need to adjust the frequency domain signal of the second microphone 2. The phase of 2 ensures that the phases of the two microphones are the same in the target direction.

[0081] Step S2: Sum the compensated microphone signals to obtain the first enhanced signal of the target signal.

[0082] In this embodiment, the two compensated microphone signals are summed to obtain the first enhanced signal of the target signal.

[0083] In this embodiment, the frequency domain output of the Delay and sum algorithm is...

[0084] (13)

[0085] It is evident that for signals in the target direction, the summed value should be the same as the original value, while for signals in non-target directions, the summed value should be less than 1 times the original value.

[0086] In this way, the first enhancement signal of the target signal can be... Performing the IDFT transform yields the delay and sum output. This method can theoretically improve the signal-to-noise ratio by 3dB.

[0087] Optionally, the directional sound pickup method of the present invention further includes:

[0088] Step S3: Compare the real and imaginary parts of the compensated microphone frequency domain signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the second enhanced signal of the target signal. .

[0089] In this embodiment, the spectrum signal of the first microphone 1 is compared. The real part of each frequency point and the compensation signal from the second microphone 2 The real part of the corresponding frequency point is selected, and the real part with the smaller absolute value is taken as the real part of the corresponding frequency point of the output signal; compare the spectrum signal of microphone 1. The imaginary part of each frequency point and the microphone 2 compensation signal The imaginary part of the corresponding frequency point is selected, and the imaginary part with the smaller absolute value is taken as the imaginary part of the corresponding frequency point of the output signal.

[0090] Specifically, select signal and compensation signal The values ​​with the smallest absolute values ​​of the real and imaginary parts are taken as the real and imaginary parts of the output signal. This is because for the target sound signals from the two microphones, the following relationship holds:

[0091] (14)

[0092] (15)

[0093] Among them, operators The operator represents taking the real part of a complex signal. Let represent the imaginary part of the complex signal. Equations (14) and (15) indicate that after phase compensation, the signals received by the two microphones in the target direction are exactly the same. In addition to the signal in the target direction, there is also random background noise generated by the microphones and signals in non-target directions, all of which need to be suppressed. Obviously, these signals do not satisfy equations (14) and (15), and the relationship between the real and imaginary parts is relatively random. Therefore, let the output be...

[0094]

[0095] (16)

[0096]

[0097] (17)

[0098] Where abs represents the modulo operation, and equations (16) and (17) indicate that the real part of the output is the real part with the smallest absolute value of the two real parts, and the imaginary part of the output is the imaginary part with the smallest absolute value of the two imaginary parts. For signals in the target direction, since equations (14) and (15) are satisfied, the real and imaginary parts of the output are the same as the original microphone signal. For signals in non-target directions, since the real and imaginary parts of the two microphone signals are different, the output signal will be significantly suppressed.

[0099] Step S4: Compare the first enhanced signal of the target signal. With the second enhanced signal The real and imaginary parts of the target signal are used to obtain the third enhanced signal by taking the real and imaginary parts with the smallest absolute values. .

[0100] Among them, the third enhanced signal The acquisition of [the substance] can be referred to in step S3, and the process is the same as in step S3, so it will not be described in detail here.

[0101] The third enhancement signal of the target signal can Performing the IDFT transform yields the delay and sum output. In this invention, by combining steps S3 and S4 with delay and sum, environmental noise exceeding 6dB can be suppressed.

[0102] Optionally, the directional sound pickup method of the present invention further includes:

[0103] Step S5: Differentiate the compensated microphone frequency domain signal to obtain the environmental noise estimation signal.

[0104] In this embodiment, the spectrum signal of the first microphone 1 is... Compensation signal with second microphone 2 By performing differentiation, the environmental noise estimation signal is obtained. .

[0105] Since the target sound source satisfies equations (14) and (15), the two signals can be differentially analyzed:

[0106] (18)

[0107] The signal necessarily does not contain the target sound signal, but only the environmental noise signal. Therefore, this embodiment can use... The signal is used as an estimate of environmental noise.

[0108] Step S6, estimate the environmental noise signal As a noise reference signal, the third enhancement signal Filtering is performed to obtain the filtered signal. .

[0109] In this embodiment, Wiener filtering or spectral subtraction can be used to enhance the third signal. Environmental noise suppression operations are performed. Since the noise printing method used here is conventional, it will not be elaborated upon further.

[0110] Step S7, filter the signal Time-frequency transformation is performed, and windowed synthesis is applied to obtain the final time-domain output signal.

[0111] It should be noted that this embodiment only describes the case of two microphones. When there are more than two microphones, it is only necessary to combine the microphones in pairs, use the above algorithm to obtain several time-domain output signals, and then combine the several time-domain output signals.

[0112] Therefore, optionally, the directional sound pickup method of this embodiment further includes:

[0113] Step S8: When the number of microphones is greater than two, select one microphone as the reference, and use the above steps S1-S7 to obtain multiple time-domain output signals for it and other microphones respectively, and combine the multiple time-domain output signals to obtain the final time-domain output signal.

[0114] In this embodiment, taking three microphones as an example, one microphone is selected as the reference, such as microphone 1; then the operations described in steps S1-S7 are performed on microphone 1 and microphone 2 to obtain the time-domain output signal Dout31, the phase of which is the same as that of microphone 1; then the operations described in the patent are performed on microphone 1 and microphone 3 to obtain the output signal Dout32, the phase of which is also the same as that of microphone 1; since Dout31 and Dout32 are both obtained with microphone 1 as the reference, these two outputs are naturally aligned, and finally Dout = (Dout31 + Dout32) / 2 is directly taken.

[0115] Exemplary apparatus

[0116] Figure 4 This is a system structure diagram of a directional sound pickup device provided in an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices, such as... Figure 4 As shown, it includes:

[0117] The time-frequency conversion and phase compensation unit 401 is used to calculate the phase difference of the microphone frequency domain signal based on the azimuth angle of the target sound source and to compensate for the phase.

[0118] Specifically, the time-frequency conversion and phase compensation unit 401 further includes:

[0119] The time-frequency conversion module is used to select any two microphone signals and convert them to the frequency domain.

[0120] The polar coordinate conversion module is used to convert frequency domain signals into polar coordinate form.

[0121] The phase compensation module calculates the phase difference between two frequency domain signals based on the location of the target sound source and compensates for this phase difference to ensure that the phase of the target signal in each microphone is completely consistent.

[0122] Summing unit 402 is used to sum the compensated microphone signals to obtain the first enhanced signal of the target signal. .

[0123] The second enhanced signal determination unit 403 is used to compare the real and imaginary parts of the compensated microphone frequency domain signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the second enhanced signal of the target signal. .

[0124] The third enhancement signal determination unit 404 is used to compare the first enhancement signal of the target signal. With the second enhanced signal The real and imaginary parts of the target signal are used to obtain the third enhanced signal by taking the real and imaginary parts with the smallest absolute values. .

[0125] Differential unit 405 is used to perform differential analysis on the compensated microphone frequency domain signal to obtain an environmental noise estimation signal.

[0126] Filtering unit 406 is used to filter the environmental noise estimation signal. As a noise reference signal, the third enhancement signal Filtering is performed to obtain the filtered signal. .

[0127] Time-frequency conversion unit 407 is used to convert the filtered signal Time-frequency transformation is performed, and windowed synthesis is applied to obtain the final time-domain output signal.

[0128] Embodiments

[0129] In this embodiment, as Figure 5 As shown, a directional sound pickup method has the following specific process:

[0130] Step A: Perform frame segmentation and time-frequency transformation on the microphone 1 signal to obtain the frequency domain signal of microphone 1. Frequency domain signal Convert to polar coordinates. Indicates the spectral amplitude. Indicates the spectral phase.

[0131] Step B, similar to step A, involves framing and time-frequency transforming the microphone 2 signal to obtain the frequency domain signal of microphone 2. Frequency domain signal Convert to polar coordinates. Indicates the spectral amplitude. Indicates the spectral phase.

[0132] Step C: Calculate the phase difference slope b between microphones 1 and 2 based on the microphone spacing and the target signal azimuth angle. Apply b to the linear phase difference formula to calculate the phase difference of the target signal at each frequency point in the two microphones. .

[0133] Step D: Compensate the phase of microphone 2 based on the calculated phase difference to obtain the phase-compensated signal of microphone 2. Regarding the signal in the direction of the target, The value and The values ​​are theoretically exactly the same.

[0134] Step E, the spectrum signal of microphone 1 Compensation signal with microphone 2 When summing, for signals in the target direction, the summed value should be twice the original value; for signals in non-target directions, the summed value should be less than twice the original value. This process is the "sum" process in delay and sum, which theoretically can improve the signal-to-noise ratio by 3dB. The output signal is designated as Enhanced Signal 1.

[0135] Step F: Compare the spectral signal of microphone 1. The real part of each frequency point and the microphone 2 compensation signal The real part of the corresponding frequency point is selected, and the real part with the smaller absolute value is taken as the real part of the corresponding frequency point of the output signal.

[0136] Step G, similar to F, compares the spectral signal of microphone 1. The imaginary part of each frequency point and the microphone 2 compensation signal The imaginary part of the corresponding frequency point is selected, and the imaginary part with the smaller absolute value is taken as the imaginary part of the corresponding frequency point of the output signal.

[0137] Step H, combined with steps F and G, yields enhanced signal 2.

[0138] Step I, similar to steps F, G, and H, involves performing corresponding operations on enhanced signal 1 and enhanced signal 2 to obtain enhanced signal 3.

[0139] Step J, the spectrum signal of microphone 1 Compensation signal with microphone 2 By performing differentiation, the environmental noise estimation signal is obtained. .

[0140] Step K, using a signal The enhanced signal 3 is filtered using a noise reference signal. Wiener filtering or spectral subtraction can be used to obtain the filtered signal. .

[0141] Step L, send the signal Time-frequency transformation is performed, and windowed synthesis is applied to obtain the final time-domain output signal.

[0142] Exemplary electronic device

[0143] Figure 6 This is the structure of an electronic device provided in an exemplary embodiment of the present invention. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them. Figure 6 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated. For example... Figure 6 As shown, the electronic device includes one or more processors 61 and memory 62.

[0144] The processor 61 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0145] The memory 62 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 61 may execute the program instructions to implement the directional sound pickup method and / or other desired functions of the software program of the various embodiments of this disclosure described above. In one example, the electronic device may also include an input device 63 and an output device 64, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0146] In addition, the input device 63 may also include, for example, a keyboard, a mouse, etc.

[0147] The output device 64 can output various information to the outside. The output device 64 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0148] Of course, for the sake of simplicity, Figure 6 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.

[0149] Exemplary computer program product and computer readable storage medium

[0150] In addition to the methods and devices described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the directional sound pickup methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0151] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0152] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the directional sound pickup methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0153] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0154] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0155] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0156] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0157] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.

[0158] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps are decomposable and / or recombinable. Such decomposition and / or recombination should be considered equivalent to the present disclosure. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0159] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A directional sound pickup method, comprising the following steps: Step S1: Calculate the phase difference of the microphone frequency domain signal based on the azimuth angle of the target sound source, and compensate for the phase difference. Step S2: Sum the compensated microphone frequency domain signals to obtain the first enhanced signal of the target signal; Step S3: Compare the real and imaginary parts of the compensated microphone frequency domain signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the second enhanced signal of the target signal; Step S4: Compare the real and imaginary parts of the first and second enhanced signals of the target signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the third enhanced signal of the target signal. Step S5: Differentiate the compensated microphone frequency domain signal to obtain the environmental noise estimation signal; Step S6: Using the environmental noise estimation signal as a noise reference signal, filter the third enhanced signal to obtain the filtered signal; Step S7: Perform time-frequency transformation on the filtered signal to obtain the final time-domain output signal.

2. The directional sound pickup method as described in claim 1, characterized in that, Step S1 includes: Step S101: For any two microphone signals, convert the microphone signals to the frequency domain and convert the frequency domain signals to polar coordinates. Step S102: Calculate the phase difference between the two frequency domain signals based on the location of the target sound source, and compensate for the phase difference to make the phase of the target signal in each microphone consistent.

3. The directional sound pickup method as described in claim 2, characterized in that, In step S3, the real part of each frequency point of the first microphone spectrum signal is compared with the real part of the corresponding frequency point of the second microphone compensation signal, and the real part with the smaller absolute value is taken as the real part of the corresponding frequency point of the second enhancement signal; the imaginary part of each frequency point of the first microphone spectrum signal is compared with the imaginary part of the corresponding frequency point of the second microphone compensation signal, and the imaginary part with the smaller absolute value is taken as the imaginary part of the corresponding frequency point of the second enhancement signal.

4. The directional sound pickup method as described in claim 3, characterized in that, In step S5, the first microphone spectrum signal and the compensation signal of the second microphone are differentially divided to obtain the environmental noise estimation signal.

5. The directional sound pickup method as described in claim 4, characterized in that, In step S7, the filtered signal is transformed by time and frequency, then windowed and synthesized to obtain the final time-domain output signal.

6. The directional sound pickup method as described in claim 1, characterized in that, The method further includes: Step S8: When the number of microphones is greater than two, select one microphone as the reference, and use the above steps S1-S7 to obtain multiple time-domain output signals for each of the other microphones, and combine the multiple time-domain output signals to obtain the final time-domain output signal.

7. A directional sound pickup device, comprising: The time-frequency conversion and phase compensation unit is used to calculate the phase difference of the microphone frequency domain signal based on the azimuth angle of the target sound source and to compensate for the phase difference. The summation unit is used to sum the compensated microphone frequency domain signal to obtain the first enhanced signal of the target signal; The second enhanced signal determination unit is used to compare the real part and imaginary part of the compensated microphone frequency domain signal, and take the real part and imaginary part with the smallest absolute value to obtain the second enhanced signal of the target signal; The third enhanced signal determination unit is used to compare the real and imaginary parts of the first enhanced signal and the second enhanced signal of the target signal, and take the real and imaginary parts with the smallest absolute values ​​to obtain the third enhanced signal of the target signal. The differential unit is used to perform differential analysis on the compensated microphone frequency domain signal to obtain the environmental noise estimation signal; The filtering unit is used to use the environmental noise estimation signal as a noise reference signal to filter the third enhanced signal to obtain the filtered signal. The time-frequency conversion unit is used to perform time-frequency conversion on the filtered signal to obtain the final time-domain output signal.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the directional sound pickup method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Pickup device

    CN106658296A

  • Wind noise processing method, device and system based on multiple microphones and storage medium

    CN110786022A

  • Voice processing method and device and device for voice processing

    CN113077808A

  • Directional selectable pickup method based on double microphones, electronic equipment and storage medium

    CN114708881A