Voice extraction method, device, storage medium and electronic device
By using microphone array and short-time Fourier transform technology, combined with beamforming and adaptive destroyers, the problem of severe distortion of speech extraction and inability to distinguish distances in the prior art is solved, and efficient and accurate speech extraction effect is achieved.
Patent Information
- Application Number
- CN202211250043.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-10-12
AI Technical Summary
In the prior art, distortion is serious during the speech extraction process, and distance distinction cannot be performed, and audio signals in the specified direction cannot be effectively picked up.
The microphone array is used for speech extraction, and the time domain signal is converted into a short-time time frequency domain signal through short-time Fourier transform. The relative delay and guidance vector of the microphone received signal are calculated as beamformer weights, and weighted processing is performed to enhance the target direction signal, and the signal is further optimized through the adaptive destroyer and post-processing module.
It effectively reduces distortion during speech extraction, can accurately distinguish audio signals in different directions, and improves the accuracy and cleanliness of speech extraction.
Smart Images

Figure CN115662394B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field related to speech extraction, and in particular to a speech extraction method, device, storage medium and electronic device. Background Art
[0002] With the development of intelligence, people have more and more demands for sound. In more and more cases, sound and audio data processing is needed to obtain semantically related recognition.
[0003] In the related technology, only audio from a specified direction can be picked up, and distance cannot be distinguished, resulting in severe distortion during voice extraction.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present invention provide a speech extraction method, device, storage medium and electronic device to at least solve the technical problem of high distortion in speech extraction in the prior art.
[0006] According to one aspect of an embodiment of the present invention, a speech extraction method is provided, comprising: obtaining an audio signal received by a microphone array, performing a short-time Fourier transform on the signal received by the microphone array, and converting a time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two subarrays, and each subarray microphone array includes at least two microphones; calculating a relative delay τ of the microphone received signal according to the microphone array topology and a target direction θ, and calculating a steering vector as a beamformer weight W; weighting the beamformer weight W to the short-time frequency domain signal, and adding the weighted signals to obtain a preliminarily enhanced target direction signal S; inputting the target direction signal S into an adaptive canceller, and outputting a signal Q; calculating the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller; determining an output gain G based on the energy and a set energy difference threshold; and applying the gain G to the output signal of the adaptive canceller. Performing short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal.
[0007] Optionally, the target direction signal S is input into an adaptive canceller, and the output signal is Q, including: caching M frames of short-time frequency domain signals and calculating the overall array DOA and the DOA of the two sub-arrays, which are respectively recorded as and Calculate the overall DOA estimate The absolute error Δα between the target direction α and the sub-array DOA is calculated. Let the adaptive canceller filter update flag be γ; if Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup range, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q.
[0008] Optionally, the determining the output gain G based on the energy and the set energy difference threshold includes: calculating the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller, which are respectively denoted as E 1 and E 2 , calculate the difference ΔE between the two; set the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1.
[0009] Optionally, the gain G is applied to the output signal of the adaptive canceller Performing a short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal includes: applying the gain G to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
[0010] According to another aspect of an embodiment of the present invention, a speech extraction device is also provided, comprising: a first acquisition unit, used to acquire an audio signal received by a microphone array, perform a short-time Fourier transform on the signal received by the microphone array, and convert the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two subarrays, and each subarray microphone array includes at least two microphones; a first calculation unit, used to calculate the relative delay τ of the microphone received signal according to the microphone array topology and the target direction θ, and calculate the steering vector as a beamformer weight W; a weighting unit, used to weight the beamformer weight W to the short-time frequency domain signal, and add the weighted signals to obtain a preliminarily enhanced target direction signal S; an adaptive unit, used to input the target direction signal S into an adaptive canceller, and the output signal is Q; a second calculation unit, used to calculate the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller; a determination unit, used to determine the output gain G based on the energy and the set energy difference threshold; an output unit, used to apply the gain G to the output signal of the adaptive canceller Performing short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal.
[0011] Optionally, the adaptive unit includes: a cache calculation module for caching M frames of short-time frequency domain signals and calculating the DOA of the entire array and the DOA of the two sub-arrays, respectively denoted as and The first calculation module is used to calculate the overall DOA estimate The absolute error Δα between the target direction α and the sub-array DOA is calculated. The adaptive canceller filter update flag is γ; the adaptive module is used if Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup range, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q.
[0012] Optionally, the determining unit includes: a second calculation module, used to calculate the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller, respectively denoted as E 1 and E 2 , calculate the difference ΔE between the two; determine the module for setting the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1.
[0013] Optionally, the output unit includes: an output module, which is used to apply the gain G to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
[0014] According to a first aspect of an embodiment of the present application, a computer-readable storage medium is provided, characterized in that a computer program is stored in the storage medium, wherein the computer program is configured to execute the above-mentioned speech extraction method when it is run.
[0015] According to a first aspect of an embodiment of the present application, there is provided an electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned speech extraction method.
[0016] In an embodiment of the present invention, an audio signal received by a microphone array is obtained, and a short-time Fourier transform is performed on the signal received by the microphone array to convert the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two subarrays, and each subarray microphone array includes at least two microphones; according to the microphone array topology structure and the target direction θ, a relative delay τ of the microphone received signal is calculated, and a steering vector is calculated as a beamformer weight W; the beamformer weight W is weighted to the short-time frequency domain signal, and the weighted signals are added to obtain a preliminarily enhanced target direction signal S; the target direction signal S is input into an adaptive canceller, and the output signal is Q; the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller are calculated; based on the energy and the set energy difference threshold, an output gain G is determined; and the gain G is applied to the output signal of the adaptive canceller The output signal of the canceller is subjected to short-time inverse Fourier transform to obtain the final output signal. In this embodiment, the signal distance is judged by the DOA of the two sub-arrays, and the audio of the specified area can be extracted in combination with the DOA information of the entire array; the accuracy of audio signal judgment is improved by calculating the DOA of continuous multi-frame signals; the noise interference outside the pickup area is further suppressed by the post-processing module, so that the extracted signal is cleaner, so as to at least solve the technical problem of high distortion in speech processing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a mobile terminal according to an optional voice extraction method of an embodiment of the present invention;
[0019] Figure 2 is a flow chart of an optional voice extraction method according to an embodiment of the present invention;
[0020] Figure 3 is a flow chart of an optional method for extracting audio from a specified area according to an embodiment of the present invention;
[0021] Figure 4 is a schematic diagram of an optional combination of two microphone arrays of arbitrary formations according to an embodiment of the present invention;
[0022] Figure 5 2 is a diagram of an optional speech extraction device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a sequence of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] The speech extraction method provided in the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a mobile terminal of a voice extraction method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal 10 may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0026] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the voice extraction method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0027] In this embodiment, a speech extraction method is also provided. Figure 2 is a flow chart of a speech extraction method according to an embodiment of the present invention. Figure 2 As shown, the speech extraction method process includes the following steps:
[0028] Step S202, obtaining the audio signal received by the microphone array, performing short-time Fourier transform on the signal received by the microphone array, and converting the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones.
[0029] Step S204 , calculating the relative delay τ of the microphone receiving signal according to the microphone array topology and the target direction θ, and calculating the steering vector as the beamformer weight W.
[0030] Step S206: weight the beamformer weight W to the short-time time-frequency domain signal, and add the weighted signals to obtain a preliminarily enhanced target direction signal S.
[0031] Step S208: input the target direction signal S into the adaptive canceller, and the output signal is Q.
[0032] Step S210, calculating the average energy of each channel of the buffered data and the energy of the output signal of the adaptive canceller.
[0033] Step S212, determining the output gain G based on the energy and the set energy difference threshold.
[0034] Step S214, applying the gain G to the output signal of the adaptive canceller The output signal of the canceller is subjected to inverse short-time Fourier transform to obtain the final output signal.
[0035] In this embodiment, the target query text and the target source text may be texts containing only Chinese, English, numbers or a mixture thereof.
[0036] The above-mentioned speech extraction method may include but is not limited to speech extraction in a long text, and the long text may include but is not limited to a text with a word count greater than or equal to a predetermined threshold, such as a long text with a word count of 300 words. The text may include but is not limited to a medical text, a doctor's prescription text, that is, the above-mentioned speech extraction method may include but is not limited to the processing of a medical text.
[0037] In this embodiment, the source text can be understood as the text to be processed, and the query text can be understood as part of the content in the source text. The specific method is to purify both the source text and the query text to obtain a standard comparison text, thereby improving the matching accuracy of the text.
[0038] Through the embodiments provided by the present application, an audio signal received by a microphone array is obtained, and a short-time Fourier transform is performed on the signal received by the microphone array to convert the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones; according to the microphone array topology structure and the target direction θ, the relative delay τ of the microphone received signal is calculated, and the steering vector is calculated as the beamformer weight W; the beamformer weight W is weighted to the short-time frequency domain signal, and the weighted signals are added to obtain a preliminarily enhanced target direction signal S; the target direction signal S is input into an adaptive canceller, and the output signal is Q; the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller are calculated; based on the energy and the set energy difference threshold, the output gain G is determined; the gain G is applied to the output signal of the adaptive canceller The output signal of the canceller is subjected to short-time inverse Fourier transform to obtain the final output signal. In this embodiment, the signal distance is judged by the DOA of the two sub-arrays, and the audio of the specified area can be extracted in combination with the DOA information of the entire array; the accuracy of audio signal judgment is improved by calculating the DOA of continuous multi-frame signals; the noise interference outside the pickup area is further suppressed by the post-processing module, so that the extracted signal is cleaner, so as to at least solve the technical problem of high distortion in speech processing in the prior art.
[0039] Optionally, the inputting the target direction signal S into the adaptive canceller and outputting the signal Q may include: buffering M frames of short-time frequency domain signals and calculating the overall array DOA and the DOA of the two sub-arrays, which are respectively recorded as and Calculate the overall DOA estimate The absolute error Δα between the target direction α and the sub-array DOA is calculated. Let the adaptive canceller filter update flag be γ; if Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup range, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q.
[0040] Optionally, determining the output gain G based on the energy and the set energy difference threshold may include: calculating the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller, which are respectively denoted as E 1 and E 2 , calculate the difference ΔE between the two; set the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1.
[0041] Optionally, the gain G is applied to the output signal of the adaptive canceller Performing an inverse short-time Fourier transform on the output signal of the canceller to obtain the final output signal may include: applying the gain G to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
[0042] As an optional embodiment, the present application also provides a method for extracting audio from a specified area. Figure 3 As shown, a flow chart of a method for extracting audio from a specified area includes the following contents.
[0043] Step S301, calculating the relative delay τ of the microphone receiving signal according to the microphone array topology and the target direction θ, and calculating the steering vector as the beamformer weight W;
[0044] Step S302, performing short-time Fourier transform on the microphone array received signal to convert the time domain signal into a short-time frequency domain signal;
[0045] Step S303, weighting the beamformer weight W to the short-time frequency domain signal, adding the weighted signals to obtain the initially enhanced target direction signal S, subtracting the two signals to obtain the interference signal N that suppresses the target direction signal, and sending the two signals to the adaptive canceller as the far-end signal and the near-end signal respectively;
[0046] Step S304: Buffer M frames of short-term time-frequency domain signals and calculate the overall array sound source localization DOA
[0047] (Direction of arrival) and the DOA of the two sub-arrays are recorded as and
[0048] Step S305, calculate the overall DOA estimate The absolute error Δα from the target direction α =
[0049] Calculate the absolute deviation between subarray DOAs Let the adaptive canceller filter update flag be γ, if Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup area, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q;
[0050] Step S306, calculating the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller, which are respectively denoted as E 1 and E 2 , calculate the difference between the two ΔE=|E 1 -E 2 |;
[0051] Step S307, setting the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1;
[0052] Step S308, applying the gain G to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
[0053] In this embodiment, two microphone arrays of arbitrary formations are combined to form a whole, such as Figure 4 A rectangular coordinate system is established for each of the two sub-arrays, and a rectangular coordinate system is established for the entire microphone array, as shown in Figure 4 As shown, the x-axes of the three coordinate systems coincide. The target audio direction is α, the target audio area angle range is δ, and the distance from the origin of the overall coordinate system is D 1 and D 2 The range between is the pickup distance range of the target area, such as Figure 4 As shown. Distance D 1 The angle relative to subarray 1 is denoted by θ 1 , the angle relative to subarray 2 is denoted as θ 2 , angle difference Δθ=|θ 1 -θ 2 |, distance D 2 The angle relative to subarray 1 is recorded as The angle relative to subarray 2 is recorded as Angle difference The DOA difference between subarray 1 and subarray 2 in the target pickup area is to between Δθ.
[0054] For the specified area, the DOA of the received signals of subarray 1 and subarray 2 are calculated. The difference between their DOA estimates can be used to determine whether the current signal is within the pickup distance area. The DOA estimation result of the entire array can be used to determine whether the current signal is within the pickup angle area. The combination of the two can complete the audio pickup in the specified area. In the pickup area, the adaptive canceller only performs data processing and does not update the filter. On the contrary, if it is outside the pickup area, the filter is updated to complete the audio extraction in the pickup area. When calculating DOA, multi-frame DOA of continuous signals is used to enhance the accuracy of audio signal judgment in the pickup area. At the same time, a post-processing module is added after the adaptive canceller to further suppress residual noise interference in the output result of the adaptive cancellation, making the output audio cleaner.
[0055] In this embodiment, the signal distance is judged by the DOA of two sub-arrays, and the audio of the specified area can be extracted by combining the DOA information of the entire array; the accuracy of audio signal judgment is improved by calculating the DOA of multiple consecutive frames of signals; and the noise interference outside the pickup area is further suppressed by the post-processing module, making the extracted signal cleaner.
[0056] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0057] In the present embodiment, a speech extraction device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0058] Figure 5 is a structural block diagram of a speech extraction device according to an embodiment of the present invention. Figure 5 As shown, the speech extraction device comprises:
[0059] The first acquisition unit 501 is used to acquire the audio signal received by the microphone array, perform short-time Fourier transform on the signal received by the microphone array, and convert the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones.
[0060] The first calculation unit 503 is used to calculate the relative delay τ of the microphone receiving signal according to the microphone array topology and the target direction θ, and calculate the steering vector as the beamformer weight W.
[0061] The weighting unit 505 is used to weight the beamformer weight W to the short-time time-frequency domain signal, and add the weighted signals to obtain a preliminarily enhanced target direction signal S.
[0062] The adaptive unit 507 is used to input the target direction signal S into the adaptive canceller, and the output signal is Q.
[0063] The second calculation unit 509 is used to calculate the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller.
[0064] The determination unit 511 is used to determine the output gain G based on the energy and a set energy difference threshold.
[0065] Output unit 513, used to apply the gain G to the output signal of the adaptive canceller Performing short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal.
[0066] Through the embodiments provided by the present application, the first acquisition unit 510 acquires the audio signal received by the microphone array, performs a short-time Fourier transform on the signal received by the microphone array, and converts the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones; the first calculation unit 503 calculates the relative delay τ of the microphone received signal according to the microphone array topology and the target direction θ, and calculates the steering vector as the beamformer weight W; the weighting unit 505 weights the beamformer weight W to the short-time frequency domain signal, and adds the weighted signals to obtain a preliminarily enhanced target direction signal S; the adaptive unit 507 inputs the target direction signal S into the adaptive canceller, and the output signal is Q; the second calculation unit 509 calculates the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller; the determination unit 511 determines the output gain G based on the energy and the set energy difference threshold; the output unit 513 acts on the output signal of the adaptive canceller The output signal of the canceller is subjected to a short-time inverse Fourier transform to obtain a final output signal. In this embodiment, the query text after noise reduction is matched with the source text to avoid noise interference, so as to at least solve the technical problem of low text matching accuracy in the prior art.
[0067] Optionally, the adaptive unit 507 may include: a cache calculation module for caching M frames of short-time frequency domain signals and calculating the DOA of the entire array and the DOA of the two sub-arrays, respectively denoted as and The first calculation module is used to calculate the overall DOA estimate The absolute error Δα between the target direction α and the sub-array DOA is calculated. The adaptive canceller filter update flag is γ; the adaptive module is used if Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup range, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q.
[0068] Optionally, the determining unit 511 may include: a second calculation module, configured to calculate the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller, respectively denoted as E 1 and E 2 , calculate the difference ΔE between the two; determine the module for setting the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1.
[0069] Optionally, the output unit 513 may include: an output module, configured to apply the gain G to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
[0070] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0071] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0072] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0073] S1, obtaining an audio signal received by a microphone array, performing a short-time Fourier transform on the signal received by the microphone array, and converting the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones.
[0074] S2, according to the microphone array topology and the target direction θ, calculate the relative delay τ of the microphone receiving signal and calculate the steering vector as the beamformer weight W.
[0075] S3, weighting the beamformer weight W to the short-time frequency domain signal, and adding the weighted signals to obtain a preliminary enhanced target direction signal S.
[0076] S4, input the target direction signal S into the adaptive canceller, and the output signal is Q.
[0077] S5, calculating the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller.
[0078] S6, determining the output gain G based on the energy and the set energy difference threshold.
[0079] S7, applying the gain G to the output signal of the adaptive canceller The output signal of the canceller is subjected to inverse short-time Fourier transform to obtain the final output signal.
[0080] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0081] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0082] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0083] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0084] S1, obtaining an audio signal received by a microphone array, performing a short-time Fourier transform on the signal received by the microphone array, and converting the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones.
[0085] S2, according to the microphone array topology and the target direction θ, calculate the relative delay τ of the microphone receiving signal and calculate the steering vector as the beamformer weight W.
[0086] S3, weighting the beamformer weight W to the short-time frequency domain signal, and adding the weighted signals to obtain a preliminary enhanced target direction signal S.
[0087] S4, input the target direction signal S into the adaptive canceller, and the output signal is Q.
[0088] S5, calculating the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller.
[0089] S6, determining the output gain G based on the energy and the set energy difference threshold.
[0090] S7, applying the gain G to the output signal of the adaptive canceller The output signal of the canceller is subjected to inverse short-time Fourier transform to obtain the final output signal.
[0091] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0092] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0093] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A speech extraction method, It is characterized in that include: Acquire an audio signal received by a microphone array, perform short-time Fourier transform on the signal received by the microphone array, and convert the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones; According to the microphone array topology and the target direction θ, the relative delay τ of the microphone receiving signal is calculated, and the steering vector is calculated as the beamformer weight W; Weighting the beamformer weight W to the short-time time-frequency domain signal, and adding the weighted signals to obtain a preliminary enhanced target direction signal S; The target direction signal S is input into the adaptive canceller, and the output signal is Q; Calculate the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller; Determine an output gain G based on the energy and a set energy difference threshold; The gain G is applied to the output signal of the adaptive canceller. Performing short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal.
2. The method according to claim 1, It is characterized in that The target direction signal S is input into an adaptive canceller, and the output signal is Q, including: Cache M frames of short-time frequency domain signals and calculate the DOA of the entire array and the DOA of the two sub-arrays, which are recorded as and Calculate the overall DOA estimate The absolute error Δα between the target direction α and the sub-array DOA is calculated. The adaptive canceller filter update flag is γ; If Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup range, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q.
3. The method according to claim 1, It is characterized in that The step of determining the output gain G based on the energy and a set energy difference threshold comprises: Calculate the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller, which are denoted as E 1 and E 2 , calculate the difference ΔE between the two; Set the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1.
4. The method according to claim 1, It is characterized in that The gain G is applied to the output signal of the adaptive canceller Performing a short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal, including: The gain G is applied to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
5. A speech extraction device, It is characterized in that include: A first acquisition unit is used to acquire an audio signal received by a microphone array, perform a short-time Fourier transform on the signal received by the microphone array, and convert the time domain signal into a short-time frequency domain signal, wherein the microphone array includes at least two sub-arrays, and each sub-array microphone array includes at least two microphones; A first calculation unit, configured to calculate a relative delay τ of a microphone receiving signal according to the microphone array topology and a target direction θ, and calculate a steering vector as a beamformer weight W; A weighting unit, configured to weight the beamformer weight W to the short-time time-frequency domain signal, and add the weighted signals to obtain a preliminarily enhanced target direction signal S; An adaptive unit, used for inputting the target direction signal S into an adaptive canceller, and outputting a signal Q; A second calculation unit, used to calculate the average energy of each channel of the cached data and the energy of the output signal of the adaptive canceller; A determination unit, configured to determine an output gain G based on the energy and a set energy difference threshold; An output unit, used to apply the gain G to the output signal of the adaptive canceller Performing short-time inverse Fourier transform on the output signal of the canceller to obtain a final output signal.
6. The device according to claim 5, It is characterized in that The adaptive unit comprises: The cache calculation module is used to cache M frames of short-time frequency domain signals and calculate the DOA of the entire array and the DOA of the two sub-arrays, which are respectively denoted as and The first calculation module is used to calculate the overall DOA estimate The absolute error Δα between the target direction α and the sub-array DOA is calculated. The adaptive canceller filter update flag is γ; Adaptive module, used if Δα≤δ and If the current signal is within the sound pickup area, set γ=0, the filter is not updated, and only data processing is performed; otherwise, if the current signal is outside the sound pickup range, set γ=1, update the filter, and record the output signal of the adaptive canceller as Q.
7. The device according to claim 5, It is characterized in that The determining unit comprises: The second calculation module is used to calculate the average energy of each channel of the cache data and the energy of the output signal of the adaptive canceller, which are respectively denoted as E 1 and E 2 , calculate the difference ΔE between the two; Determination module, used to set the energy difference threshold E th And the output gain G, if ΔE>E th , then G=0.1, otherwise G=1.
8. The device according to claim 5, It is characterized in that The output unit comprises: An output module is used to apply the gain G to the output signal of the adaptive canceller, that is, right Perform short-time inverse Fourier transform to obtain the final output signal.
9. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 4 when executed.
10. An electronic device comprising a memory and a processor, It is characterized in that A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Two-dimensional directional pickup method and device
CN113050035A
Directional pickup method and device, electronic equipment and storage medium
CN114023347A