A noise estimation method, device, chip, medium and module equipment
By dividing molecular bands based on the coherence coefficient and adjusting the coherence coefficient, the problem of insufficient noise estimation performance of the existing dual-mix noise reduction algorithm is solved, and more efficient noise estimation and lower voice loss are achieved.
Patent Information
- Application Number
- CN202211256812.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-10-14
AI Technical Summary
The existing dual-mile noise reduction algorithm still has shortcomings in noise estimation performance, which is difficult to meet the high requirements of users for their usage experience.
In the double-mill noise estimation method, the frequency band is divided into subbands based on the coherence coefficients of the first sound signal and the second sound signal, and the coherence coefficient is adjusted according to the mean value of the subband coherence coefficients, and the double-mill noise estimation result is optimized.
It improves noise estimation performance, reduces voice loss, and enhances noise estimation robustness in different scenarios.
Smart Images

Figure CN115631763B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing, and in particular to a method, device, chip, medium and module equipment for dual-microphone noise estimation. Background Art
[0002] Noise estimation plays an important role in speech noise reduction, and the accuracy of noise estimation determines the quality of noise reduction. As users have higher and higher requirements for user experience, the noise reduction performance of single-microphone noise reduction algorithms is difficult to meet user requirements. Currently, mobile phones, tablets and other devices are usually equipped with more than one microphone, so the industry has proposed a dual-microphone noise reduction algorithm. The dual-microphone noise reduction algorithm can optimize noise estimation using information such as the coherence of dual-microphone signals. However, in practice, it is found that the noise reduction performance of the existing dual-microphone noise reduction algorithm is still insufficient. Therefore, improving the dual-microphone noise reduction algorithm and improving the noise estimation performance is an urgent problem to be solved. Summary of the invention
[0003] The present application provides a robust dual-microphone noise estimation method, which is beneficial to improving the noise estimation performance.
[0004] In a first aspect, the present application provides a noise estimation method, the method comprising: collecting a first sound signal through a first microphone, and collecting a second sound signal through a second microphone; performing noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in a first frequency band, the first frequency band being the frequency band where the first sound signal and the second sound signal are located; determining a coherence coefficient corresponding to each frequency point in the first frequency band between the first sound signal and the second sound signal; dividing the first frequency band into M sub-bands, and based on the coherence coefficient corresponding to each frequency point in the first frequency band, determining a mean coherence coefficient corresponding to each sub-band in the M sub-bands, where M is greater than 1. integer; if the mean value of the coherence coefficient corresponding to the sub-band is greater than the first threshold value corresponding to the sub-band, the coherence coefficient of each frequency point in the sub-band is increased; if the mean value of the coherence coefficient of the sub-band is less than the first threshold value corresponding to the sub-band, the coherence coefficient of each frequency point in the sub-band is kept unchanged; determine the dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, the dual-microphone noise estimation result corresponding to the frequency point is the weighted sum of the single-microphone noise estimation result corresponding to the frequency point and the power spectrum of the first sound signal corresponding to the frequency point, the weight value of the single-microphone noise estimation result corresponding to the frequency point is positively correlated with the coherence coefficient corresponding to the frequency point, and the weight value of the power spectrum of the first sound signal corresponding to the frequency point is negatively correlated with the coherence coefficient corresponding to the frequency point.
[0005] Based on the method described in the first aspect, considering the difference in the coherence of noise in different frequency bands, the coherence between sound signals can be optimized by sub-band. The sub-band coherence coefficient mean can be compared with the corresponding sub-band threshold, and the coherence coefficient value in the corresponding sub-band greater than the sub-band threshold can be enhanced. In this way, when the dual-microphone noise estimation result is subsequently determined based on the single-microphone noise estimation result and the power spectrum corresponding to the first sound signal, the weight of the single-microphone noise estimation result can be increased, and the weight of the power spectrum corresponding to the first sound signal can be reduced, thereby reducing speech loss and improving noise estimation performance.
[0006] In a possible implementation, the M subbands include a first subband and a second subband, the frequency of the first subband is lower than the frequency of the second subband, and the first threshold corresponding to the first subband is lower than the first threshold corresponding to the second subband.
[0007] In one possible implementation, the method also includes: determining a mean value of the coherence coefficient corresponding to the first frequency band based on the coherence coefficient corresponding to each frequency point in the first frequency band; determining whether the mean value of the coherence coefficient corresponding to the first frequency band is greater than or equal to a second threshold; if the mean value of the coherence coefficient corresponding to the first frequency band is greater than or equal to the second threshold, denoising the sound signal based on a single-microphone noise estimation result corresponding to each frequency point in the first frequency band; if the mean value of the coherence coefficient corresponding to the first frequency band is less than the second threshold, executing the step of dividing the first frequency band into M sub-bands.
[0008] In a possible implementation, the method further includes: performing denoising processing on the sound signal based on a dual-microphone noise estimation result corresponding to each frequency point in the first frequency band.
[0009] In a possible implementation, increasing the coherence coefficient of each frequency point in the sub-band includes: updating the coherence coefficient of the frequency point to a smaller value between a multiple of the coherence coefficient of the frequency point and 1; or updating the coherence coefficient of the frequency point to 1.
[0010] Based on these possible implementations, over-estimation of noise can be reduced, speech loss can be avoided, and noise estimation can be optimized.
[0011] In a possible implementation, when the electronic device is in the handheld mode, energy of the first sound signal is greater than energy of the second sound signal.
[0012] Based on this possible implementation, the robustness of noise estimation in a handheld scenario can be improved.
[0013] In one possible implementation, the method also includes: performing denoising processing on the sound signal based on a target noise estimation result corresponding to each frequency point in the first frequency band, the target noise estimation result corresponding to the frequency point being the maximum value of a single-microphone noise estimation result corresponding to the frequency point and a dual-microphone noise estimation result corresponding to the frequency point.
[0014] Based on this possible implementation method, underestimation of the noise signal in the dual-microphone noise estimation process can be avoided, thereby improving the overall noise estimation level.
[0015] In a second aspect, the present application provides a noise estimation device, which includes a collection unit, an estimation unit, a determination unit and an update unit, and is used to execute the method of the first aspect above.
[0016] In a third aspect, the present application provides a noise estimation device, which includes a microphone, a memory and a processor. The microphone is used to collect sound signals, the memory stores program instructions, and the processor is configured to call the program instructions to execute the method of the first aspect above.
[0017] In a fourth aspect, the present application provides a module device, which includes a microphone, a power module, a storage module, a communication module and a chip, wherein: the microphone is used to collect sound signals; the power module is used to provide power to the module device; the storage module is used to store data and instructions; the communication module is used for internal communication within the module device, or for the module device to communicate with external devices; the chip is used to execute the method of the first aspect above.
[0018] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the processor is configured to enable the chip to execute the method of the first aspect.
[0019] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed on an electronic device, the electronic device executes the method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 It is a structural schematic diagram of a mobile phone equipped with dual microphones provided in an embodiment of the present application;
[0022] Figure 2 is a flowchart of a noise estimation method provided in an embodiment of the present application;
[0023] Figure 3 is a flowchart of a noise estimation method provided in an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of the structure of a noise estimation device provided in an embodiment of the present application;
[0025] Figure 5 is a schematic diagram of the structure of a noise estimation device provided in an embodiment of the present application;
[0026] Figure 6 It is a structural schematic diagram of a module device provided in an embodiment of the present application;
[0027] Figure 7 is a schematic diagram of the structure of the chip provided in the embodiment of the present application;
[0028] Figure 8 It is a schematic diagram of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more listed items.
[0031] It should be noted that the terms "first", "second", "third", etc. in the specification and claims of the present application and in the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0032] The embodiment of the present application proposes a noise estimation method, device, chip, medium and module device. The method can be performed by an electronic device equipped with at least two microphones. Optionally, the electronic device may include a terminal with a voice signal acquisition function, such as a mobile phone, a tablet computer, a computer with a wireless transceiver function, a virtual reality (VR) user device, an augmented reality (AR) user device, a wireless terminal in industrial control, a vehicle-mounted user device, a wireless terminal in unmanned driving, a wireless terminal in telemedicine, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a wearable user device, etc. The embodiment of the present application does not limit the application scenario. Electronic devices may sometimes also be referred to as terminals, user devices, access user devices, vehicle-mounted terminals, industrial control terminals, UE units, UE stations, mobile stations, mobile stations, remote stations, remote user devices, mobile devices, UE user devices, user devices, wireless communication devices, UE agents or UE devices, etc. The electronic device may also be any fixed or mobile device with a voice signal acquisition function. Alternatively, the method may also be executed by a chip in an electronic device. Alternatively, the method may also be implemented by a software module loaded in a terminal.
[0033] See also Figure 1 , Figure 1 This is a schematic diagram of the structure of a mobile phone 100 with dual microphones provided in an embodiment of the present application. Figure 1 As shown, the mobile phone 100 equipped with dual microphones can collect the first sound signal and the second sound signal through the two microphones respectively. It can be understood that Figure 1 The mobile phone with dual microphones shown is merely an example of an electronic device that can implement the method of the present application, and does not limit the electronic device of the embodiments of the present application.
[0034] The method provided in the embodiment of the present application is further described below:
[0035] See also Figure 2 , Figure 2 A flowchart of a noise estimation method provided in an embodiment of the present application. Figure 2 The method described takes an electronic device as an example of an execution subject.
[0036] 201. An electronic device collects a first sound signal through a first microphone and collects a second sound signal through a second microphone.
[0037] 202. The electronic device performs noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in a first frequency band, where the first frequency band is a frequency band where the first sound signal and the second sound signal are located.
[0038] Optionally, the electronic device may use a traditional single-microphone noise estimation method to perform noise estimation on the first sound, including but not limited to minimum tracking, minimum control recursive averaging and other methods. Alternatively, the electronic device may use other methods to perform noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in the first frequency band. The single-microphone noise estimation result corresponding to the frequency point is expressed as δ mono (ω,k), where ω is the frequency index of the sound signal and k is the frame index of the sound signal.
[0039] Optionally, if the electronic device is in hands-free mode, the first microphone may be any microphone in the electronic device. If the electronic device is in handheld mode, the first microphone may be a microphone closest to the speaker, that is, the energy of the first sound signal is greater than the energy of the second sound signal.
[0040] 203. The electronic device determines a coherence coefficient between the first sound signal and the second sound signal corresponding to each frequency point in the first frequency band.
[0041] Optionally, the electronic device may perform coherence calculation on the first sound signal and the second sound signal to determine the coherence coefficient corresponding to each frequency point of the first sound signal and the second sound signal in the first frequency band. The coherence coefficient corresponding to the first sound signal and the second sound signal at the frequency point is used to characterize the correlation between the first sound signal and the second sound signal at the frequency point. For example, the method for coherence calculation may use magnitude squared coherence (MSC). The MSC calculation formula is as follows:
[0042]
[0043] in, represents the self-power of the first sound signal, represents the self-power of the second sound signal, represents the mutual power of two sound signals, ω is the frequency index of the sound signal, and k is the frame index of the sound signal.
[0044] Alternatively, other methods may be used to calculate the coherence coefficients of the first sound signal and the second sound signal at the frequency points, which is not limited in the embodiment of the present application.
[0045] 204. The electronic device divides the first frequency band into M sub-bands, and determines a mean value of the coherence coefficient corresponding to each sub-band in the M sub-bands based on the coherence coefficient corresponding to each frequency point in the first frequency band, where M is an integer greater than 1.
[0046] In one possible implementation, the first frequency band may be evenly divided into M sub-bands. In another possible implementation, the first frequency band may not be evenly divided into M sub-bands, for example, the frequency band may be divided based on arithmetic differences or geometric ratios of a preset first sub-bandwidth and subsequent sub-bandwidths and their respective previous sub-bandwidths, or the frequency band may be divided based on other suitable frequency band division methods.
[0047] 205. The electronic device determines whether a mean value of a coherence coefficient corresponding to the sub-band is greater than a first threshold corresponding to the sub-band.
[0048] Optionally, the first threshold may be preset or dynamically adjusted based on the use environment or operation mode of the electronic device.
[0049] 206. If the mean value of the coherence coefficient corresponding to the sub-band is greater than the first threshold corresponding to the sub-band, the electronic device increases the coherence coefficient of each frequency point in the sub-band.
[0050] Specifically, if the sub-band coherence coefficient mean is greater than the first threshold corresponding to the sub-band, it indicates that there may be speech in the sub-band. In this case, the coherence coefficient of each frequency point in the sub-band is increased, and when the dual-microphone noise estimation result is subsequently determined based on the single-microphone noise estimation result and the power spectrum corresponding to the first sound signal, the weight of the single-microphone noise estimation result is increased and the weight of the power spectrum corresponding to the first sound signal is reduced, thereby reducing speech loss.
[0051] In a possible implementation, increasing the coherence coefficient of each frequency point in the sub-band includes: updating the coherence coefficient of the frequency point to a smaller value between a multiple of the coherence coefficient of the frequency point and 1; or updating the coherence coefficient of the frequency point to 1.
[0052] For example, let Or, To further protect the voice and audio points.
[0053] 207. If the mean value of the coherence coefficient of the sub-band is less than the first threshold value corresponding to the sub-band, the electronic device keeps the coherence coefficient of each frequency point in the sub-band unchanged.
[0054] In a possible implementation, the noise has strong coherence at low frequencies and weak coherence at high frequencies. In order to further reduce the speech loss at low frequencies, in noise estimation, the first thresholds corresponding to the various subbands are not necessarily the same. Generally, the threshold corresponding to the low-frequency subband is smaller, and the threshold corresponding to the high-frequency subband is larger. In other words, the divided M subbands include at least the first subband and the second subband, the frequency of the first subband is smaller than the frequency of the second subband, and the first threshold corresponding to the first subband is smaller than the first threshold corresponding to the second subband.
[0055] In another possible implementation, the first thresholds corresponding to the various sub-bands may also be the same.
[0056] 208. The electronic device determines a dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, the dual-microphone noise estimation result corresponding to the frequency point is a weighted sum of a single-microphone noise estimation result corresponding to the frequency point and a power spectrum of the first sound signal corresponding to the frequency point, the weight value of the single-microphone noise estimation result corresponding to the frequency point is positively correlated with the coherence coefficient corresponding to the frequency point, and the weight value of the power spectrum of the first sound signal corresponding to the frequency point is negatively correlated with the coherence coefficient corresponding to the frequency point.
[0057] Specifically, the electronic device calculates the single-microphone noise estimation result δ according to the coherence coefficient and the power spectrum of the current frame of the first sound signal. mono (ω, k) is optimized to determine the dual-microphone noise estimation result corresponding to each frequency point in the first frequency band. Determine the weighting coefficient. The weight value of the single-microphone noise estimation result corresponding to the frequency point is positively correlated with the coherence coefficient corresponding to the frequency point. The weight value of the power spectrum of the first sound signal corresponding to the frequency point is negatively correlated with the coherence coefficient corresponding to the frequency point. When the coherence between the first sound signal and the second sound signal is weak, it can be considered that the current frequency point noise is dominant. Therefore, a larger weight is given to the power spectrum of the current frame signal during the dual-microphone noise estimation process. When the coherence is strong, it can be considered that the current frequency point speech is dominant. Therefore, the dual-microphone noise estimation gives the single-microphone estimated noise δ mono (ω,k) is weighted more heavily, thereby reducing speech loss. For example, the update formula for the dual-microphone noise estimation result can be as follows:
[0058]
[0059] In a possible implementation, the electronic device can also perform denoising on the sound signal based on the target noise estimation result corresponding to each frequency point in the first frequency band, where the target noise estimation result corresponding to the frequency point is the maximum value of the single-microphone noise estimation result corresponding to the frequency point and the dual-microphone noise estimation result corresponding to the frequency point.
[0060] In other words, in order to avoid underestimation of the noise signal in the dual-microphone noise estimation process, the maximum value of the dual-microphone estimation result and the single-microphone estimation result can be taken as the final dual-microphone noise estimation result, that is:
[0061] δ dual (ω,k)=max(δ dual (ω,k),δ mono (ω,k)
[0062] In another possible implementation, the electronic device may also directly perform denoising on the sound signal based on the dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, that is, the electronic device directly uses the dual-microphone noise estimation result as the final dual-microphone noise estimation result.
[0063] By implementing Figure 2 The described method optimizes the coherence between sound signals by sub-band, taking into account the difference in the coherence of noise in different frequency bands. The sub-band coherence coefficient mean is compared with the corresponding sub-band threshold, and the coherence coefficient value in the corresponding sub-band greater than the sub-band threshold is enhanced. In this way, when the dual-microphone noise estimation result is subsequently determined based on the single-microphone noise estimation result and the power spectrum corresponding to the first sound signal, the weight of the single-microphone noise estimation result can be increased, and the weight of the power spectrum corresponding to the first sound signal can be reduced, thereby reducing speech loss.
[0064] Please refer to Figure 3 , Figure 3 It is a flowchart of the noise estimation method provided in an embodiment of the present application. Figure 3 The method described takes an electronic device as an example of an execution subject.
[0065] 301. An electronic device collects a first sound signal through a first microphone and collects a second sound signal through a second microphone.
[0066] 302. The electronic device performs noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in a first frequency band, where the first frequency band is a frequency band where the first sound signal and the second sound signal are located.
[0067] 303. The electronic device determines a coherence coefficient corresponding to each frequency point in the first frequency band between the first sound signal and the second sound signal.
[0068] in, Figure 3 Steps 301-303 of the method shown are similar to Figure 2 Steps 201-203 in the method shown are the same and will not be repeated here.
[0069] 304. Determine a mean value of the coherence coefficient corresponding to the first frequency band based on the coherence coefficient corresponding to each frequency point in the first frequency band.
[0070] Optionally, the mean value of the coherence coefficient corresponding to the first frequency band may be an arithmetic mean value, a geometric mean value, a sliding mean value or any other suitable mean value of the coherence coefficient corresponding to the first frequency band.
[0071] 305. Determine whether the mean value of the coherence coefficient corresponding to the first frequency band is greater than a second threshold.
[0072] Optionally, the second threshold may be preset or dynamically adjusted based on the use environment or operation mode of the electronic device.
[0073] 306. If the mean value of the coherence coefficient corresponding to the first frequency band is greater than or equal to the second threshold, denoising the sound signal based on the single-microphone noise estimation result corresponding to each frequency point in the first frequency band.
[0074] It can be understood that if the mean value of the coherence coefficient of the first frequency band is greater than or equal to the second threshold, it indicates that the coherent component in the frame is dominant, that is, the speech is dominant, and the coherence is not used to further optimize the noise estimation, and the sound signal is denoised using the single-microphone noise estimation result.
[0075] 307. If the mean coherence coefficient corresponding to the first frequency band is less than the second threshold, the first frequency band is divided into M sub-bands, and based on the coherence coefficient corresponding to each frequency point in the first frequency band, the mean coherence coefficient corresponding to each sub-band in the M sub-bands is determined, where M is an integer greater than 1.
[0076] It can be understood that if the mean value of the coherence coefficient is less than the second threshold, it indicates that there are more irrelevant components, ie, noise, in the frame, and the coherence is used to estimate the noise for further optimization.
[0077] 308. The electronic device determines whether a mean value of a coherence coefficient corresponding to the sub-band is greater than a first threshold corresponding to the sub-band.
[0078] 309. If the mean value of the coherence coefficient corresponding to the sub-band is greater than the first threshold corresponding to the sub-band, the electronic device increases the coherence coefficient of each frequency point in the sub-band.
[0079] 310. If the mean value of the coherence coefficient of the sub-band is less than a first threshold value corresponding to the sub-band, the electronic device keeps the coherence coefficient of each frequency point in the sub-band unchanged.
[0080] 311. The electronic device determines a dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, the dual-microphone noise estimation result corresponding to the frequency point is a weighted sum of a single-microphone noise estimation result corresponding to the frequency point and a power spectrum of the first sound signal corresponding to the frequency point, the weight value of the single-microphone noise estimation result corresponding to the frequency point is positively correlated with the coherence coefficient corresponding to the frequency point, and the weight value of the power spectrum of the first sound signal corresponding to the frequency point is negatively correlated with the coherence coefficient corresponding to the frequency point.
[0081] Figure 3 The subsequent steps 308-311 in the method shown are similar to Figure 2 Steps 205-208 in the method shown are the same and will not be repeated here.
[0082] exist Figure 3In the described embodiment, a second threshold is set to determine whether the sound signal is dominated by the speech signal, and further noise estimation optimization is performed only on frames where the speech signal is not dominant, that is, where there are more noise signals. This can avoid the problem of over-estimation of noise in speech frames, save resource overhead of the noise estimation method, and reduce the complexity of noise estimation processing.
[0083] See also Figure 4 , Figure 4 4 is a schematic diagram of the structure of a noise estimation device 400 provided in an embodiment of the present application. The noise estimation device 400 can be used to perform part or all of the functions of the electronic device in the above method embodiment. Figure 4 As shown, the noise estimation device 400 includes a collection unit 401, an estimation unit 402, a determination unit 403, and an update unit 404. Among them:
[0084] The collecting unit 401 is used to collect a first sound signal through a first microphone and collect a second sound signal through a second microphone;
[0085] The estimation unit 402 is used to perform noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in a first frequency band, where the first sound signal and the second sound signal are located;
[0086] The determination unit 403 is used to determine the coherence coefficient corresponding to each frequency point of the first sound signal and the second sound signal in the first frequency band;
[0087] The determining unit 403 is further configured to divide the first frequency band into M sub-bands, and determine a mean value of the coherence coefficient corresponding to each sub-band in the M sub-bands based on the coherence coefficient corresponding to each frequency point in the first frequency band, where M is an integer greater than 1;
[0088] The updating unit 404 is used to increase the coherence coefficient of each frequency point in the sub-band if the mean value of the coherence coefficient corresponding to the sub-band is greater than the first threshold value corresponding to the sub-band; if the mean value of the coherence coefficient of the sub-band is less than the coherence coefficient threshold value corresponding to the sub-band, keep the coherence coefficient of each frequency point in the sub-band unchanged;
[0089] The determination unit 403 is also used to determine the dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, and the dual-microphone noise estimation result corresponding to the frequency point is the weighted sum of the single-microphone noise estimation result corresponding to the frequency point and the power of the first sound signal corresponding to the frequency point, wherein the weight value of the single-microphone noise estimation result corresponding to the frequency point is positively correlated with the coherence coefficient corresponding to the frequency point, and the weight value of the power spectrum of the first sound signal corresponding to the frequency point is negatively correlated with the coherence coefficient corresponding to the frequency point.
[0090] In a possible implementation, the M subbands include a first subband and a second subband, the frequency of the first subband is lower than the frequency of the second subband, and the first threshold corresponding to the first subband is lower than the first threshold corresponding to the second subband.
[0091] In a possible implementation, the noise estimation device 400 further includes a denoising unit; wherein:
[0092] The determination unit 403 is further configured to determine a mean value of the coherence coefficient corresponding to the first frequency band based on the coherence coefficient corresponding to each frequency point in the first frequency band; and determine whether the mean value of the coherence coefficient corresponding to the first frequency band is greater than or equal to a second threshold;
[0093] The denoising unit is used to perform denoising on the sound signal based on the single-microphone noise estimation result corresponding to each frequency point in the first frequency band if the mean value of the coherence coefficient corresponding to the first frequency band is greater than or equal to the second threshold; if the mean value of the coherence coefficient corresponding to the first frequency band is less than the second threshold, trigger the determination unit 403 to execute the step of dividing the first frequency band into M sub-bands.
[0094] In a possible implementation, the noise estimation device 400 further includes a denoising unit; the denoising unit is configured to perform denoising processing on the sound signal based on the dual-microphone noise estimation result corresponding to each frequency point in the first frequency band.
[0095] In a possible implementation, the updating unit 404 increases the coherence coefficient of each frequency point in the subband by: updating the coherence coefficient of the frequency point to a smaller value between a multiple of the coherence coefficient of the frequency point and 1; or updating the coherence coefficient of the frequency point to 1.
[0096] In a possible implementation, when the electronic device including the noise estimation apparatus 400 is in a handheld mode, the energy of the first sound signal is greater than the energy of the second sound signal.
[0097] In one possible implementation, the noise estimation device 400 also includes a denoising unit; the denoising unit is used to perform denoising processing on the sound signal based on the target noise estimation result corresponding to each frequency point in the first frequency band, and the target noise estimation result corresponding to the frequency point is the maximum value of the single-microphone noise estimation result corresponding to the frequency point and the dual-microphone noise estimation result corresponding to the frequency point.
[0098] Figure 5 is a schematic diagram of the structure of the noise estimation device 500 provided in an embodiment of the present application. Figure 5 As shown, the apparatus 500 may include microphones 501 - 1 . . . 501 - n , a memory 502 , and a processor 503 . The microphones 501 - 1 . . . 501 - n , the memory 502 , and the processor 503 are connected via one or more buses 504 .
[0099] Microphones 501 - 1 . . . 501 - n are used to collect sound signals.
[0100] The memory 502 is used to store program instructions. The memory 502 may include a read-only memory and a random access memory, and provides instructions and data to the processor 503. A portion of the memory 502 may also include a non-volatile random access memory.
[0101] The processor 503 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or alternatively, the processor 503 may be any conventional processor, etc. The processor 503 calls the program instructions stored in the memory 502 to enable the device 500 to execute Figure 2 or Figure 3 The noise estimation method described in the corresponding method embodiment.
[0102] Figure 6 600 is a schematic diagram of the structure of the module device 600 provided in the embodiment of the present application. Figure 6 As shown, the module device 600 can execute the relevant steps of the noise estimation method in the above method embodiment. The module device 600 includes: microphones 601 - 1 . . . 601 - n, a power module 602 , a storage module 603 , a communication module 604 and a chip 605 .
[0103] Among them, microphones 601-1...601-n are used to collect sound signals; power module 602 is used to provide power to the module device; storage module 603 is used to store data and instructions; communication module 604 is used for internal communication of the module device; chip 605 is used to execute Figure 2 or Figure 3 The noise estimation method described in the corresponding method embodiment.
[0104] Figure 7 Schematic diagram of the structure of the chip 700 provided in the embodiment of the present application. Figure 7 As shown, the chip 700 includes a processor 701 and a communication interface 702. The processor 701 is configured to enable the chip 700 to execute Figure 2 or Figure 3The noise estimation method described in the corresponding method embodiment.
[0105] Figure 8 800 is a schematic diagram of a computer-readable storage medium 800 provided in an embodiment of the present application. Figure 8 As shown, a computer-readable storage medium 800 stores a computer-readable instruction 801. When the instruction is executed on a processor, the method flow of the above method embodiment is implemented. The computer-readable storage medium 800 includes, but is not limited to, for example, a volatile memory and / or a non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc.
[0106] The embodiment of the present application also provides a computer program product. When the computer program product runs on a processor, the method flow of the above method embodiment is implemented.
[0107] Regarding the various modules / units included in the various devices and products described in the above embodiments, they can be software modules / units, or hardware modules / units, or they can be partially software modules / units and partially hardware modules / units. For example, for various devices and products applied to or integrated in a chip, the various modules / units included therein can all be implemented in the form of hardware such as circuits, or at least some of the modules / units can be implemented in the form of software programs, which run on the integrated processor inside the chip, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits; for various devices and products applied to or integrated in a chip module, the various modules / units included therein can all be implemented in the form of hardware such as circuits, and different modules / units can be located in the same part of the chip module (such as a chip, circuit module, etc.) or in different components, or at least some of the modules / units can be implemented in the form of software programs. The software programs run on the integrated processor inside the chip, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits. It can be implemented in the form of a software program, which runs on a processor integrated inside the chip module, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits; for various devices and products applied to or integrated in the terminal, the modules / units contained therein can all be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (for example, chip, circuit module, etc.) or in different components in the terminal, or, at least some modules / units can be implemented in the form of a software program, which runs on a processor integrated inside the terminal, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits.
[0108] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain operations can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0109] The descriptions of the various embodiments provided in this application can refer to each other, and the descriptions of the various embodiments have different focuses. For parts that are not described in detail in a certain embodiment, refer to the relevant descriptions of other embodiments. For the convenience and simplicity of description, for example, the functions of the various devices and equipment provided in the embodiments of this application and the operations performed can refer to the relevant descriptions of the method embodiments of this application, and the various method embodiments and the various device embodiments can also refer to, combine or quote each other.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A noise estimation method, characterized in that: The method comprises: Collecting a first sound signal through a first microphone and collecting a second sound signal through a second microphone; Performing noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in a first frequency band, where the first frequency band is a frequency band where the first sound signal and the second sound signal are located; Determine a coherence coefficient between the first sound signal and the second sound signal corresponding to each frequency point in the first frequency band; Divide the first frequency band into M sub-bands, and determine a mean value of the coherence coefficient corresponding to each sub-band in the M sub-bands based on the coherence coefficient corresponding to each frequency point in the first frequency band, where M is an integer greater than 1; If the mean value of the coherence coefficient corresponding to the sub-band is greater than the first threshold value corresponding to the sub-band, then the coherence coefficient of each frequency point in the sub-band is increased; otherwise, the coherence coefficient of each frequency point in the sub-band is kept unchanged; Determine a dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, the dual-microphone noise estimation result corresponding to the frequency point is a weighted sum of a single-microphone noise estimation result corresponding to the frequency point and a power spectrum corresponding to the first sound signal at the frequency point, the weight value of the single-microphone noise estimation result corresponding to the frequency point is positively correlated with the coherence coefficient corresponding to the frequency point, and the weight value of the power spectrum corresponding to the first sound signal at the frequency point is negatively correlated with the coherence coefficient corresponding to the frequency point.
2. The method according to claim 1, characterized in that: The M sub-bands include a first sub-band and a second sub-band, a frequency of the first sub-band is smaller than a frequency of the second sub-band, and a first threshold corresponding to the first sub-band is smaller than a first threshold corresponding to the second sub-band.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Determine a mean value of the coherence coefficient corresponding to the first frequency band based on the coherence coefficient corresponding to each frequency point in the first frequency band; Determining whether a mean value of coherence coefficients corresponding to the first frequency band is greater than or equal to a second threshold; If the mean value of the coherence coefficient corresponding to the first frequency band is greater than or equal to the second threshold, denoising the sound signal based on the single-microphone noise estimation result corresponding to each frequency point in the first frequency band; If the mean value of the coherence coefficient corresponding to the first frequency band is less than the second threshold, the step of dividing the first frequency band into M sub-bands is performed.
4. The method according to claim 1 or 2, characterized in that: The method further comprises: The sound signal is denoised based on the dual-microphone noise estimation result corresponding to each frequency point in the first frequency band.
5. The method according to claim 1 or 2, characterized in that: The increasing the coherence coefficient of each frequency point in the sub-band includes: updating the coherence coefficient of the frequency point to a smaller value between a multiple of the coherence coefficient of the frequency point and 1; or, Update the coherence coefficient of the frequency point to 1.
6. The method according to claim 1 or 2, characterized in that: When the electronic device is in the handheld mode, the energy of the first sound signal is greater than the energy of the second sound signal.
7. The method according to claim 1 or 2, characterized in that: The method further comprises: The sound signal is denoised based on a target noise estimation result corresponding to each frequency point in the first frequency band, wherein the target noise estimation result corresponding to the frequency point is a maximum value between a single-microphone noise estimation result corresponding to the frequency point and a dual-microphone noise estimation result corresponding to the frequency point.
8. A noise estimation device, characterized in that: The noise estimation device comprises: A collection unit, configured to collect a first sound signal through a first microphone and a second sound signal through a second microphone; an estimating unit, configured to perform noise estimation on the first sound signal to obtain a single-microphone noise estimation result corresponding to each frequency point in a first frequency band, where the first frequency band is a frequency band where the first sound signal and the second sound signal are located; a determining unit, configured to determine a coherence coefficient corresponding to each frequency point in the first frequency band between the first sound signal and the second sound signal; The determining unit is further configured to divide the first frequency band into M sub-bands, and determine a mean value of the coherence coefficient corresponding to each sub-band in the M sub-bands based on the coherence coefficient corresponding to each frequency point in the first frequency band, where M is an integer greater than 1; an updating unit, configured to increase the coherence coefficient of each frequency point in the sub-band if the mean value of the coherence coefficient corresponding to the sub-band is greater than the first threshold value corresponding to the sub-band; otherwise, keep the coherence coefficient of each frequency point in the sub-band unchanged; The determination unit is further used to determine a dual-microphone noise estimation result corresponding to each frequency point in the first frequency band, the dual-microphone noise estimation result corresponding to the frequency point being a weighted sum of a single-microphone noise estimation result corresponding to the frequency point and the power of the first sound signal corresponding to the frequency point, the weight value of the single-microphone noise estimation result corresponding to the frequency point being positively correlated with the coherence coefficient corresponding to the frequency point, and the weight value of the power spectrum of the first sound signal corresponding to the frequency point being negatively correlated with the coherence coefficient corresponding to the frequency point.
9. A noise estimation device, comprising a microphone, a memory and a processor, wherein the microphone is used to collect sound signals, the memory stores program instructions, and the processor is configured to call the program instructions so that the device executes the method according to any one of claims 1 to 7.
10. A module device, comprising a microphone, a power module, a storage module, a communication module and a chip, wherein: The microphone is used to collect sound signals; The power module is used to provide electrical energy to the module device; The storage module is used to store data and instructions; The communication module is used for internal communication of the module device, or for communication between the module device and an external device; The chip is used to execute the method according to any one of claims 1 to 7.
11. A chip, comprising a processor and a communication interface, wherein the processor is configured to enable the chip to execute the method according to any one of claims 1 to 7.
12. A computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by an electronic device, the electronic device executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Noise suppression method and electronic equipment
CN113160846A
Voice processing method and electronic equipment
CN113823314A