A Speech Enhancement Method, Device, Terminal and Medium for Cross-Device Multi-Microphone Cooperative Processing

Through multi-microphones across devices, audio signals from the host microphone and external microphone are obtained, gain compensation and noise reduction are performed, which solves the problem of noise isolation in complex environments and improves voice clarity.

CN119905100BActive Publication Date: 2025-06-27ELEVOC TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510410030.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-27
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively isolate interfering noise in complex environments. The deep learning noise reduction method of a single microphone cannot remove interfering voices, and the noise suppression effect is limited, resulting in a decrease in speech clarity.

Method used

The method of cross-device multi-microphone collaborative processing is adopted to obtain the audio signals of the host microphone and external microphone, and the noise reduction processing of the external microphone audio signals is achieved through gain compensation, power spectrum calculation and noise estimation.

Benefits of technology

Effectively isolate peripheral noise signals, improve voice clarity, and enhance target speaking voices. It is suitable for voice input scenes in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119905100B_ABST
    Figure CN119905100B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice enhancement method, device, terminal and medium for cross-device multi-microphone collaborative processing. The method includes: obtaining a host microphone audio signal and an external microphone audio signal, and performing gain compensation on the external microphone audio signal; calculating power spectra based on the host microphone audio signal and the externally connected microphone audio signal after gain compensation, and determining the power spectrum difference ratio and the power spectrum difference spectrum of the host microphone audio signal and the externally connected microphone audio signal after gain compensation at each frequency; performing noise estimation based on the power spectrum difference ratio and the power spectrum difference spectrum at each frequency to obtain a noise power spectrum, and performing noise reduction processing on the externally connected microphone audio signal based on the noise power spectrum to obtain a noise-reduced externally connected microphone audio signal. The present invention can effectively isolate surrounding noise signals through the collaborative work of the host microphone and the external microphone, realize voice enhancement, and effectively improve voice clarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly to a voice enhancement method, device, terminal and medium for cross-device multi-microphone collaborative processing. Background Art

[0002] At present, external headphone devices usually come with built-in microphones. However, in some complex environments, a single microphone may not be able to effectively isolate interfering noises. For a single microphone, even using deep learning noise reduction methods, it is impossible to remove interfering human voices, and at the same time, the noise suppression effect is limited, resulting in a decrease in speech clarity. In addition, existing audio processing technologies often only rely on the microphones of external devices, while ignoring the potential value of the built-in microphones of host devices. Currently, in general external headphone or earphone devices on the market, after being connected to a host device (such as a laptop, mobile phone, tablet, etc.), they will automatically switch and enable the microphones inside the external headphone devices, while the built-in microphones of the host devices are disabled or not enabled. This design is effective in daily audio input scenarios, but in noisy environments, relying solely on the input of a single microphone is easily interfered by environmental noises, and it is impossible to clearly capture the voice of the target speaker, and the potential value of the built-in microphones of the host devices is not utilized.

[0003] Therefore, there are still defects in the prior art. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a voice enhancement method, device, terminal and medium for cross-device multi-microphone collaborative processing in view of the above-mentioned defects of the prior art. The technical solutions adopted by the present invention are as follows:

[0005] In a first aspect, the present invention provides a voice enhancement method for cross-device multi-microphone collaborative processing, wherein the method includes:

[0006] Obtain a host microphone audio signal and an external microphone audio signal, and perform gain compensation on the external microphone audio signal;

[0007] Based on the host microphone audio signal and the externally-connected microphone audio signal after gain compensation, perform power spectrum calculation, and determine the power spectrum difference ratio and the power spectrum difference spectrum of the host microphone audio signal and the externally-connected microphone audio signal at each frequency;

[0008] Based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum, perform noise estimation to obtain a noise power spectrum, and perform noise reduction processing on the external microphone audio signal based on the noise power spectrum to obtain a noise-reduced external microphone audio signal.

[0009] In one implementation, the method is applied to a microphone array formed by a host microphone and an external microphone, and speech enhancement is achieved by simultaneously acquiring the host microphone audio signal and the external microphone audio signal.

[0010] In one implementation, the gain compensation for the external microphone audio signal includes:

[0011] Perform time-domain synchronization on the host microphone audio signal and the external microphone audio signal;

[0012] Perform background noise tracking detection on the time-domain synchronized host microphone audio signal and external microphone audio signal to obtain the noise levels corresponding to the host microphone audio signal and the external microphone audio signal;

[0013] Based on the noise levels corresponding to the host microphone audio signal and the external microphone audio signal, determine the gain compensation coefficient corresponding to the external microphone audio signal;

[0014] Based on the gain compensation coefficient, perform gain compensation on the external microphone audio signal.

[0015] In one implementation, the performing gain compensation on the external microphone audio signal based on the gain compensation coefficient includes:

[0016] Multiply the gain compensation coefficient by the external microphone audio signal to obtain the gain-compensated external microphone audio signal.

[0017] In one implementation, the calculating the power spectrum based on the host microphone audio signal and the gain-compensated external microphone audio signal includes:

[0018] Convert the host microphone audio signal and the gain-compensated external microphone audio signal into frequency-domain signals;

[0019] Based on the frequency-domain signals, calculate the power spectra corresponding to the host microphone audio signal and the gain-compensated external microphone audio signal, and the power spectra are used to reflect the power distributions of the host microphone audio signal and the gain-compensated external microphone audio signal at different frequencies.

[0020] In one implementation, the performing noise estimation based on the power spectrum difference ratio and the power spectrum difference spectrum at each frequency to obtain the noise power spectrum includes:

[0021] Obtain a preset noise estimation model;

[0022] Obtain the preset first threshold, second threshold, and third threshold in the noise estimation model, where the first threshold is less than the second threshold;

[0023] Compare the power spectrum difference ratio and the power spectrum difference spectrum at each frequency with the first threshold, the second threshold, and the third threshold respectively;

[0024] If the power spectrum difference ratio at a certain frequency is between the first threshold and the second threshold, and the corresponding power spectrum difference spectrum is less than the third threshold, determine that there is noise at this frequency and obtain the noise power spectrum.

[0025] In one implementation, the noise reduction processing of the external microphone audio signal based on the noise power spectrum to obtain the noise-reduced external microphone audio signal includes:

[0026] Calculate the filtering coefficient based on the noise power spectrum and the power spectrum corresponding to the original external microphone audio signal;

[0027] Filter the frequency-domain representation of the original external microphone audio signal based on the filtering coefficient to obtain the frequency-domain representation of the noise-reduced signal;

[0028] Convert the frequency-domain representation of the noise-reduced signal into a time-domain representation to obtain the noise-reduced external microphone audio signal.

[0029] In a second aspect, an embodiment of the present invention further includes a voice enhancement device for cross-device multi-microphone collaboration, where the device is used to implement the steps of the voice enhancement method for cross-device multi-microphone collaboration described in the above solution, and the device includes:

[0030] A gain compensation module, configured to obtain the host microphone audio signal and the external microphone audio signal, and perform gain compensation on the external microphone audio signal;

[0031] A power spectrum analysis module, configured to perform power spectrum calculation based on the host microphone audio signal and the gain-compensated external microphone audio signal, and determine the power spectrum difference ratio and the power spectrum difference spectrum at each frequency between the host microphone audio signal and the gain-compensated external microphone audio signal;

[0032] A noise reduction processing module, configured to perform noise estimation based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum to obtain the noise power spectrum, and perform noise reduction processing on the external microphone audio signal based on the noise power spectrum to obtain the noise-reduced external microphone audio signal.

[0033] In a third aspect, an embodiment of the present invention further provides a terminal. The terminal includes a memory, a processor, and a voice enhancement program for cross-device multi-microphone collaborative processing stored in the memory and executable on the processor. When the processor executes the voice enhancement program for cross-device multi-microphone collaborative processing, the steps of the voice enhancement method for cross-device multi-microphone collaborative processing in any one of the above solutions are implemented.

[0034] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. A voice enhancement program for cross-device multi-microphone collaborative processing is stored on the computer-readable storage medium. When the voice enhancement program for cross-device multi-microphone collaborative processing is executed by a processor, the steps of the voice enhancement method for cross-device multi-microphone collaborative processing in any one of the above solutions are implemented.

[0035] Beneficial effects: Compared with the prior art, the present invention provides a voice enhancement method for cross-device multi-microphone collaborative processing. The method is applied to a microphone array formed by a host microphone and an external microphone, and realizes voice enhancement by simultaneously acquiring the host microphone audio signal and the external microphone audio signal. The present invention first acquires the host microphone audio signal and the external microphone audio signal, and performs gain compensation on the external microphone audio signal. Then, based on the host microphone audio signal and the externally microphone audio signal after gain compensation, power spectrum calculation is performed, and the power spectrum difference ratio and the power spectrum difference spectrum of the host microphone audio signal and the externally microphone audio signal after gain compensation at each frequency are determined. Finally, based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum, noise estimation is performed to obtain a noise power spectrum, and the external microphone audio signal is denoised based on the noise power spectrum to obtain a denoised external microphone audio signal. The present invention utilizes the potential value of the host microphone, and through the collaborative work of the host microphone and the external microphone, effectively isolates the surrounding noise signals, realizes voice enhancement, and effectively improves voice clarity. Description of the Drawings

[0036] Figure 1 It is a flowchart of a preferred embodiment of the voice enhancement method for cross-device multi-microphone collaborative processing provided by an embodiment of the present invention.

[0037] Figure 2 It is a schematic diagram of a preferred embodiment of the present invention applied to an external headphone device connected to a laptop computer.

[0038] Figure 3 It is a schematic diagram of a preferred embodiment of the present invention applied to an external headphone device connected to a mobile phone.

[0039] Figure 4 It is a flowchart of performing gain compensation on the external microphone audio signal in the present invention.

[0040] Figure 5 It is the spectrogram of the external microphone audio signal obtained.

[0041] Figure 6 It is the spectrogram of the host microphone audio signal obtained.

[0042] Figure 7 It is the spectrogram of the speech signal after speech enhancement processing using the method of the present invention.

[0043] Figure 8 It is a schematic diagram of the architecture of the speech enhancement device for cross-device multi-microphone collaborative processing provided by the embodiment of the present invention.

[0044] Figure 9 It is a schematic block diagram of the terminal provided by the embodiment of the present invention. Detailed implementation manners

[0045] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations or steps, nor do they necessarily need to be executed in the described order. For example, some operations or steps can also be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0047] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless otherwise clearly specified in the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0048] It should be understood that, in order to facilitate a clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first control information and the second control information are only used to distinguish different control information, and do not limit their order.

[0049] Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit to be different.

[0050] It should also be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0051] Based on the deficiencies of the prior art, this embodiment provides a voice enhancement method for cross-device multi-microphone collaborative processing. This method can be applied to terminals, and the terminals can be intelligent terminal products such as computers, smart TVs, mobile phones, etc. As Figure 1 shown in

[0052] Step S100: Obtain the host microphone audio signal and the external microphone audio signal, and perform gain compensation on the external microphone audio signal.

[0053] The voice enhancement method for cross-device multi-microphone collaborative processing in this embodiment is applied to a microphone array formed by a host microphone and an external microphone. By simultaneously obtaining the host microphone audio signal and the external microphone audio signal to achieve voice enhancement, it is a cross-device implementation method. Specifically, the host microphone in this embodiment is the microphone built into the host device, such as the microphone built into the edge of a laptop or the microphone built into a mobile phone. The external microphone is the microphone in an external wired earphone or a Bluetooth earphone. The core of the present invention lies in the collaborative work of the external microphone and the host microphone. By analyzing the stable feature of the amplitude difference of the sound signals collected by the two microphones, effective isolation of interference noise is achieved, and the target speaker's voice is enhanced. Since the two microphones are different in spatial position, for sounds from different directions, the amplitudes of the sound signals received by them will be different. The target speaker's voice is usually closer to the external microphone (such as a Bluetooth earphone) and has a larger amplitude; while the interference noise is relatively farther away, and the amplitudes when reaching the two microphones are relatively closer. Therefore, this embodiment can utilize this characteristic to process the signals collected by the two microphones through a specific algorithm, thereby achieving the effects of noise reduction and voice enhancement.

[0054] Specifically, as Figure 2 shown in Figure 2Schematic diagram of a preferred embodiment of the present invention applied to connecting an external headphone device to a laptop computer. In the prior art, when a user connects an external wired headphone with a microphone, the system automatically switches the data source of the system audio to the external wired headphone, ignoring the built-in microphone device of the laptop computer. To solve this problem, in this embodiment, while the system automatically switches the audio signal source to the external wired headphone, a built-in microphone audio signal acquisition module of the laptop computer is constructed; when the user makes a call using the external wired headphone, the audio signals of both the external wired headphone and the laptop computer's microphone can be acquired simultaneously, thereby constructing a microphone array to enhance the microphone audio signals collaboratively, and further ensuring effective isolation of surrounding interference signals and effectively improving the call quality. In another embodiment, as Figure 3 shown in Figure 3 Schematic diagram of a preferred embodiment of the present invention applied to connecting an external headphone device to a mobile phone. Similarly, in this embodiment, the mobile phone, as the host device, has a built-in microphone. When a user connects a Bluetooth headphone device with a microphone, the audio signals of both the host device and the external microphone will be acquired simultaneously, enabling collaborative processing of the microphone data of the two devices.

[0055] Based on this, the voice enhancement method for cross-device multi-microphone collaborative processing in this embodiment is applied to a microphone array. The microphone array includes at least two microphones, one of which is the host microphone (such as a mobile phone), and the other is the external microphone (such as a Bluetooth headphone). When the user makes a call using the external microphone, this embodiment can acquire the host microphone audio signal and the external microphone audio signal simultaneously. Then, since there are many types of external devices with microphones in the market, and the enhancement method in this embodiment mainly uses the amplitude difference of the sound signals collected by two microphones (the external device and the host device) as a feature, and the microphones of different external devices have different microphone sensitivities. In order to effectively adapt to various types of microphone external devices in the market, it is necessary to perform adaptive gain compensation on the external microphone audio signal.

[0056] In one implementation, as Figure 4 shown, the steps for performing gain compensation on the external microphone audio signal in this embodiment are as follows:

[0057] Step S101: Perform time-domain synchronization on the host microphone audio signal and the external microphone audio signal;

[0058] Step S102: Perform background noise tracking detection on the time-domain synchronized host microphone audio signal and external microphone audio signal to obtain the noise levels corresponding to the host microphone audio signal and the external microphone audio signal;

[0059] Step S103: Determine the gain compensation coefficient corresponding to the external microphone audio signal based on the noise levels corresponding to the host microphone audio signal and the external microphone audio signal;

[0060] Step S104: Perform gain compensation on the external microphone audio signal based on the gain compensation coefficient.

[0061] Specifically, in this embodiment, the cross-correlation function of the dual-channel signals (i.e., the host microphone audio signal and the external microphone audio signal) is first calculated. Based on this cross-correlation function, the peak position is determined, and then time-domain synchronization is achieved through peak alignment. In specific applications, this embodiment can determine the offset of these two signals based on the determined peak position. Based on the determined offset, the external microphone audio signal is translated in the time domain so that the peak of the external microphone audio signal is aligned with the peak of the host microphone audio signal, thereby ensuring that the external microphone audio signal and the host microphone audio signal are strictly synchronized on the time axis.

[0062] Next, in the case of no speech or mute state, the noise floor tracking detection is respectively performed on the time-domain synchronized host microphone audio signal and the external microphone audio signal to obtain the noise levels corresponding to the host microphone audio signal and the external microphone audio signal, which can be specifically implemented by means of Voice Activity Detection (VAD). Then this embodiment determines the gain compensation coefficient corresponding to the external microphone audio signal based on the noise levels corresponding to the host microphone audio signal and the external microphone audio signal. Specifically, the gain compensation coefficient corresponding to the external microphone audio signal is the square root of the ratio between the noise levels of the two signals. After calculating the gain compensation coefficient, this embodiment can perform gain compensation on the external microphone audio signal based on the gain compensation coefficient. Specifically, the gain compensation coefficient can be multiplied by the external microphone audio signal to obtain the gain-compensated external microphone audio signal, realizing the amplitude adjustment of the external microphone audio signal to make its amplitude consistent with the host microphone audio signal (i.e., the reference signal), and completing the automatic adjustment of the microphone gains of different devices.

[0063] Step S200: Calculate the power spectrum based on the host microphone audio signal and the gain-compensated external microphone audio signal, and determine the power spectrum difference ratio and the power spectrum difference spectrum at each frequency between the host microphone audio signal and the gain-compensated external microphone audio signal.

[0064] After performing gain compensation on the external microphone audio signal, this embodiment can calculate and analyze the power spectra of the signals of the two channels after gain compensation. The power spectrum reflects the distribution of the signals of each channel at different frequencies. Therefore, based on the calculated power spectra, the power spectrum difference ratio and the power spectrum difference spectrum of the host microphone audio signal and the external microphone audio signal after gain compensation at each frequency can be determined respectively. By analyzing the power spectrum difference ratio and the power spectrum difference spectrum, this embodiment can effectively distinguish noise signals.

[0065] Specifically, this embodiment can adopt the Discrete Time Fourier Transform (DTFT) algorithm to convert the host microphone audio signal and the external microphone audio signal after gain compensation into frequency-domain signals, obtain the power distribution of each channel signal at different frequencies, and then calculate the power spectra corresponding to the host microphone audio signal and the external microphone audio signal after gain compensation based on the frequency-domain signals. For example, assume that the external microphone audio signal after gain compensation is x1(t), and the host microphone audio signal is x2(t), where t represents the time corresponding to the audio signal. After the discrete-time Fourier transform, their frequency-domain signals are X1(f) and X2(f) respectively. Further, the power spectrum of the external microphone audio signal after gain compensation is calculated as P1(f) = |X1(f)|², and the power spectrum of the host microphone audio signal is P2(f) = |X2(f)|², where f represents frequency.

[0066] Furthermore, in order to more accurately measure the difference degree of the powers of the two microphone audio signals at different frequencies, this embodiment introduces the power spectrum difference ratio P(f), and its calculation method is: P(f) = P1(f) / P2(f), where P1(f) is the power spectrum of the external microphone audio signal, and P2(f) is the power spectrum of the host microphone audio signal. P(f) reflects the relative magnitude relationship of the signal powers collected by the two microphones at frequency f. When the target speaker's voice is dominant, since the target speaker is closer to the external microphone, P1(f) will be relatively larger, making the value of P(f) larger.

[0067] In addition, after obtaining the power spectrum difference ratio P(f), the power spectrum difference spectrum D(f) is introduced in this embodiment. The calculation method of the power spectrum difference spectrum is D(f) = P1(f) - P2(f), and the power spectrum difference spectrum in this embodiment is a specific value. It can be seen that this embodiment can calculate the power spectrum difference ratio and the power spectrum difference spectrum at each frequency. The power spectrum difference ratio and the power spectrum difference spectrum complement each other and jointly reflect the power difference of the signals in two channels at different frequencies. Since the propagation characteristics of the target speaker's voice and the interfering noise at the two microphones are different, they both have obvious characteristics in the power spectrum difference ratio and the power spectrum difference spectrum, providing richer information for subsequent noise identification and processing.

[0068] Step S300: Based on the power spectrum difference ratio and the power spectrum difference spectrum at each frequency, perform noise estimation to obtain a noise power spectrum, and based on the noise power spectrum, perform noise reduction processing on the external microphone audio signal to obtain a noise-reduced external microphone audio signal.

[0069] After calculating the power spectrum difference ratio and the power spectrum difference spectrum at each frequency, this embodiment can combine a noise estimation model to perform noise estimation, identify the noise distribution, and obtain a noise power spectrum. Specifically, the noise estimation model in this embodiment can preset a first threshold T1, a second threshold T2, and a third threshold T3, where the first threshold T1 is less than the second threshold T2. Then, the power spectrum difference ratio and the power spectrum difference spectrum corresponding to each frequency are respectively compared with the first threshold, the second threshold, and the third threshold. If the power spectrum difference ratio at a certain frequency is between the first threshold and the second threshold (i.e., T1 < P(f) < T2), and the corresponding power spectrum difference spectrum is less than the third threshold (i.e., D(f) < T3), it is determined that there is noise in the signal at this frequency, realizing the identification of noise. Based on the same method, it is possible to analyze whether there is noise at all other frequencies, and then obtain the noise power spectrum according to the power spectrum characteristics of this part of the signal with noise. The noise power spectrum reflects the power distribution of the noise signal at different frequencies.

[0070] The above first threshold, second threshold, and third threshold in this embodiment are set based on a large amount of experimental data and actual application scenarios. In addition, this embodiment also collects a large amount of speech data in different noise environments, analyzes the distribution characteristics of the target speaker's voice and noise in the power spectrum difference ratio and the power spectrum difference spectrum, so as to set each threshold that can effectively distinguish noise and the target speaker's voice. By continuously tracking and updating the power spectrum characteristics of the noise, the noise signal can be more accurately identified.

[0071] Considering the time-varying characteristics of noise in the actual environment, the noise estimation model of this embodiment has the ability of adaptive adjustment. As time goes by and signals are continuously input, the noise estimation model will update the estimation of the noise power spectrum in real time according to new observation data. For example, when the environmental noise suddenly increases or changes in type, the model can quickly capture the changes in the power spectrum difference ratio and the power spectrum difference spectrum, and re-estimate the new noise power spectrum to ensure the stability and effectiveness of the noise reduction effect.

[0072] Furthermore, after obtaining the noise power spectrum, this embodiment can adopt algorithms such as Wiener filtering and deep learning to calculate the filtering coefficients according to the noise power spectrum and the power spectrum corresponding to the original external microphone audio signal, and filter the original external microphone audio signal to remove the noise components and retain the target speaker's voice. Specifically, for each frequency f, the filtering coefficient H(f) is calculated according to the Wiener filtering formula, and then the frequency-domain representation of the original external microphone audio signal is filtered based on the filtering coefficient to obtain the frequency-domain representation of the noise-reduced signal Y(f)=H(f)*X2(f). Then, the frequency-domain representation of the noise-reduced signal is converted into the time-domain representation through the Inverse Fast Fourier Transform (IFFT) to obtain the noise-reduced external microphone audio signal.

[0073] After verification, using the speech enhancement method of the present invention can greatly suppress the interfering human voice and background noise. Specifically, as Figure 5 shown, Figure 5 is the spectrogram of the external microphone audio signal obtained, which includes the voice of the external microphone wearer, the surrounding interfering human voices, and environmental noise. Figure 6 is the spectrogram of the host microphone audio signal obtained, which includes the target human voice, the surrounding interfering human voices, and environmental noise. After being processed by the method of the present invention, as Figure 7 shown, Figure 7 is the spectrogram of the speech signal after being enhanced by the method of the present invention. It can be seen from Figure 7 that the interfering human voices and background noise are greatly suppressed, and the voice of the external microphone wearer is effectively enhanced.

[0074] In summary, in this embodiment, the host microphone audio signal and the external microphone audio signal are first obtained, and the external microphone audio signal is subjected to gain compensation. Then, based on the host microphone audio signal and the gain-compensated external microphone audio signal, power spectrum calculation is performed, and the power spectrum difference ratio and the power spectrum difference spectrum of the host microphone audio signal and the gain-compensated external microphone audio signal at each frequency are determined. Finally, based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum, noise estimation is performed to obtain the noise power spectrum, and based on the noise power spectrum, noise reduction processing is performed on the external microphone audio signal to obtain the noise-reduced external microphone audio signal. This embodiment utilizes the potential value of the host microphone, and through the collaborative work of the host microphone and the external microphone, effectively isolates the surrounding noise signals, realizes speech enhancement, and effectively improves speech clarity.

[0075] The present invention can be widely applied to various devices that require high-quality voice input, such as teleconferences, speech recognition, live broadcasts, and online communication scenarios, and is particularly suitable for scenarios with high requirements for speech clarity in noisy environments, such as outdoor scenarios, open office environments, public places, etc.

[0076] Based on the above embodiment, the present invention provides a voice enhancement device for cross-device multi-microphone collaborative processing. The device can be used to implement the steps of the above-mentioned voice enhancement method for cross-device multi-microphone collaborative processing, as Figure 8 shown, the device includes: a gain compensation module 10, a power spectrum analysis module 20, and a noise reduction processing module 30. Specifically, the gain compensation module 10 is used to obtain the host microphone audio signal and the external microphone audio signal, and perform gain compensation on the external microphone audio signal. The power spectrum analysis module 20 is used to perform power spectrum calculation based on the host microphone audio signal and the gain-compensated external microphone audio signal, and determine the power spectrum difference ratio and the power spectrum difference spectrum of the host microphone audio signal and the gain-compensated external microphone audio signal at each frequency. The noise reduction processing module 30 is used to perform noise estimation based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum to obtain the noise power spectrum, and perform noise reduction processing on the external microphone audio signal based on the noise power spectrum to obtain the noise-reduced external microphone audio signal.

[0077] The working principle of each module in the voice enhancement system for cross-device multi-microphone collaborative processing in this embodiment is the same as that of each step in the above method embodiment, and will not be elaborated here.

[0078] Each module in the above-mentioned voice enhancement system for cross-device multi-microphone collaborative processing can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor in the terminal in hardware form or independent of the processor, or stored in the memory in the terminal in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0079] Based on the above embodiments, the present invention further provides a terminal, and the principle block diagram of the terminal can be as Figure 9 shown. The terminal may include one or more processors 100 ( Figure 9 only one is shown in the figure), a memory 101, and a computer program 102 stored in the memory 101 and executable on one or more processors 100. For example, a voice enhancement program for cross-device multi-microphone collaborative processing. When one or more processors 100 execute the computer program 102, each step in the embodiment of the voice enhancement method for cross-device multi-microphone collaborative processing can be implemented. Alternatively, when one or more processors 100 execute the computer program 102, the functions of each module / unit in the embodiment of the voice enhancement system for cross-device multi-microphone collaborative processing can be implemented, which is not limited herein.

[0080] In one embodiment, the so-called processor 100 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0081] In one embodiment, the memory 101 may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 101 may also include both the internal storage unit and the external storage device of the electronic device. The memory 101 is used to store the computer program and other programs and data required by the terminal. The memory 101 may also be used to temporarily store the data that has been output or will be output.

[0082] Those skilled in the art can understand that Figure 9 the block diagram of the principle shown in Figure 9 is only the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0083] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, operational database or other medium used in the various embodiments provided by the present invention may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech enhancement method for cross-device multi-microphone collaborative processing, characterized in that: The method comprises: Acquire a host microphone audio signal and an external microphone audio signal, and perform gain compensation on the external microphone audio signal; Performing power spectrum calculation based on the host microphone audio signal and the gain-compensated external microphone audio signal, and determining a power spectrum difference ratio and a power spectrum difference spectrum between the host microphone audio signal and the gain-compensated external microphone audio signal at each frequency; Performing noise estimation based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum to obtain a noise power spectrum, and performing noise reduction processing on the external microphone audio signal based on the noise power spectrum to obtain a noise-reduced external microphone audio signal; The power spectrum calculation based on the host microphone audio signal and the gain-compensated external microphone audio signal includes: Converting the host microphone audio signal and the gain-compensated external microphone audio signal into frequency domain signals; Calculating a power spectrum corresponding to the host microphone audio signal and the gain-compensated external microphone audio signal based on the frequency domain signal, wherein the power spectrum is used to reflect the power distribution of the host microphone audio signal and the gain-compensated external microphone audio signal at different frequencies; The audio signal of the external microphone after gain compensation is x1(t), and the audio signal of the host microphone is x2(t), where t represents the time corresponding to the audio signal. After short-time Fourier transform, the corresponding frequency domain signals are X1(f) and X2(f), respectively. The power spectrum of the external microphone audio signal after gain compensation is calculated as P1(f)=|X1(f)|², and the power spectrum of the host microphone audio signal is calculated as P2(f)=|X2(f)|², where f represents the frequency. The power spectrum difference ratio is calculated as: P(f) = P1(f) / P2(f), where P(f) reflects the relative magnitude of the signal power collected by the host microphone and the external microphone at frequency f; The power spectrum difference spectrum is calculated as: D(f)=P1(f)-P2(f). The power spectrum difference spectrum is a specific value. The power spectrum difference ratio and the power spectrum difference spectrum complement each other and jointly reflect the power difference of the signals of the host microphone and the external microphone at different frequencies. The performing noise estimation based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum to obtain the noise power spectrum includes: Obtaining a preset noise estimation model; Acquire a first threshold, a second threshold, and a third threshold preset in the noise estimation model, wherein the first threshold is less than the second threshold; Compare the power spectrum difference ratio and the power spectrum difference spectrum at each frequency with a first threshold, a second threshold, and a third threshold, respectively; If the power spectrum difference ratio at a certain frequency is between the first threshold and the second threshold, and the corresponding power spectrum difference spectrum is less than the third threshold, it is determined that noise exists at the frequency, and the noise power spectrum is obtained; The noise estimation model has an adaptive adjustment capability. When the ambient noise suddenly increases or changes in type, the noise estimation model can quickly capture the changes in the power spectrum difference ratio and the power spectrum difference spectrum, and re-estimate a new noise power spectrum. The performing noise reduction processing on the external microphone audio signal based on the noise power spectrum to obtain the external microphone audio signal after noise reduction includes: Calculating a filter coefficient based on the noise power spectrum and the original power spectrum corresponding to the external microphone audio signal; Filtering the original frequency domain representation of the external microphone audio signal based on the filter coefficient to obtain a frequency domain representation of the signal after noise reduction; The frequency domain representation of the noise-reduced signal is converted into a time domain representation to obtain a noise-reduced external microphone audio signal.

2. The method for speech enhancement by cross-device multi-microphone collaborative processing according to claim 1, characterized in that: The method is applied to a microphone array formed by a host microphone and an external microphone, and voice enhancement is achieved by simultaneously acquiring the host microphone audio signal and the external microphone audio signal.

3. The method for speech enhancement by cross-device multi-microphone collaborative processing according to claim 1, characterized in that: The step of performing gain compensation on the external microphone audio signal comprises: Performing time domain synchronization on the host microphone audio signal and the external microphone audio signal; Performing background noise tracking detection on the host microphone audio signal and the external microphone audio signal after time domain synchronization to obtain noise levels corresponding to the host microphone audio signal and the external microphone audio signal; Determining a gain compensation coefficient corresponding to the external microphone audio signal based on noise levels corresponding to the host microphone audio signal and the external microphone audio signal; Based on the gain compensation coefficient, gain compensation is performed on the external microphone audio signal.

4. The method for speech enhancement by cross-device multi-microphone collaborative processing according to claim 3, characterized in that: The step of performing gain compensation on the external microphone audio signal based on the gain compensation coefficient includes: The gain compensation coefficient is multiplied by the external microphone audio signal to obtain a gain-compensated external microphone audio signal.

5. A speech enhancement device for cross-device multi-microphone collaborative processing, characterized in that: The device is used to implement the steps of the method for speech enhancement by cross-device multi-microphone collaborative processing according to any one of claims 1 to 4, and the device comprises: A gain compensation module, used to obtain a host microphone audio signal and an external microphone audio signal, and perform gain compensation on the external microphone audio signal; A power spectrum analysis module, configured to calculate a power spectrum based on the host microphone audio signal and the gain-compensated external microphone audio signal, and determine a power spectrum difference ratio and a power spectrum difference spectrum between the host microphone audio signal and the gain-compensated external microphone audio signal at each frequency; The noise reduction processing module is used to perform noise estimation based on the power spectrum difference ratio at each frequency and the power spectrum difference spectrum to obtain a noise power spectrum, and perform noise reduction processing on the external microphone audio signal based on the noise power spectrum to obtain a noise-reduced external microphone audio signal.

6. A terminal, characterized in that: The terminal includes a memory, a processor, and a speech enhancement program for cross-device multi-microphone collaborative processing stored in the memory and executable on the processor. When the processor executes the speech enhancement program for cross-device multi-microphone collaborative processing, the steps of the speech enhancement method for cross-device multi-microphone collaborative processing as described in any one of claims 1-4 are implemented.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a speech enhancement program for cross-device multi-microphone collaborative processing. When the speech enhancement program for cross-device multi-microphone collaborative processing is executed by the processor, the steps of the speech enhancement method for cross-device multi-microphone collaborative processing as described in any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Wired and wireless microphone arrays

    US20140355775A1