Adjustment of noise reduction
By adjusting the noise reduction intensity based on the signal-to-noise ratio (SNR) estimate, the problem of signal separation difficulties in deep neural networks under low SNR environments is solved, the noise reduction performance of hearing devices is optimized, and speech comprehension and sound quality are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- COCHLEAR LIMITED
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-29
AI Technical Summary
Existing noise reduction technologies based on deep neural networks struggle to effectively separate target signals in low signal-to-noise ratio environments, resulting in severe distortion of the input signal and affecting speech comprehension and sound quality for users of hearing aids.
The noise reduction intensity is automatically adjusted based on the signal-to-noise ratio (SNR) estimate. The noise reduction intensity is determined by the ratio of the output signal of the deep neural network (DNN) to the original input signal. The mixing ratio uses a higher DNR intensity under high SNR and a lower DNR intensity under low SNR to optimize the noise reduction effect.
The noise reduction intensity is automatically adjusted under different signal-to-noise ratio environments, which improves the speech comprehension and sound quality of hearing device users, reduces signal distortion, and enhances the applicability and effectiveness of noise reduction technology.
Smart Images

Figure CN122122920A_ABST
Abstract
Description
background Technical Field
[0001] This invention generally relates to techniques for adjusting noise reduction. Background Technology
[0002] In recent decades, medical devices have provided a wide range of therapeutic benefits to recipients. Medical devices can include internal or implantable components / devices, external or wearable components / devices, or combinations thereof (e.g., devices having an external component that communicates with the implantable component). Medical devices, such as conventional hearing aids, partially or fully implantable hearing prostheses (e.g., bone conduction devices, mechanical stimulators, cochlear implants), etc. wait Pacemakers, defibrillators, functional electrical stimulation devices, and other medical devices have been successful for many years in performing life-saving and / or lifestyle improvement functions and / or recipient monitoring.
[0003] Over the years, the types of medical devices and the range of functions they perform have increased. For example, many medical devices, sometimes referred to as “implantable medical devices,” now typically include one or more instruments, devices, sensors, processors, controllers, or other functional mechanical or electrical components that are permanently or temporarily implanted into a recipient’s body. These functional devices are typically used to diagnose, prevent, monitor, treat, or manage diseases / injuries or their symptoms, or to study, replace, or modify anatomical structures or physiological processes. Many of these functional devices utilize power and / or data received from an external device that is part of or operates in conjunction with the implantable component. Summary of the Invention
[0004] In one aspect, a method is provided. The method includes: receiving at least one input signal at one or more input devices; processing the at least one input signal using a noise reduction process to generate a spectral gain for application to the at least one input signal; and determining the signal-to-noise ratio (SNR) of the at least one input signal based on the spectral gain.
[0005] In another aspect, one or more non-transitory computer-readable storage media are provided. The one or more non-transitory computer-readable storage media include instructions that, when executed by one or more processors of a hearing device, cause the one or more processors to: calculate a noise reduction gain for application to one or more sound signals received at an audio input of the hearing device; determine a signal-to-noise ratio (SNR) of the one or more sound signals based on the noise reduction gain; apply the noise reduction gain to the one or more sound signals to generate one or more noise-reduced signals; and mix the one or more noise-reduced signals with the one or more sound signals according to a mixing ratio, wherein the mixing ratio is a function of the SNR.
[0006] In another aspect, one or more non-transitory computer-readable storage media are provided. The one or more non-transitory computer-readable storage media include instructions that, when executed by one or more processors of a hearing device, cause the one or more processors to: calculate a noise reduction gain for application to one or more sound signals received at an audio input of the hearing device; determine a signal-to-noise ratio (SNR) of the one or more sound signals based on the noise reduction gain; determine a noise reduction intensity parameter based on the SNR; and generate an output signal from the noise reduction gain and the noise reduction intensity parameter.
[0007] In another aspect, an apparatus is provided. The apparatus includes: one or more input devices configured to receive a noisy input signal, wherein the noisy input signal includes a target signal embedded in additive noise; a deep neural network (DNN)-based denoising module configured to generate a DNN output signal from the noisy input signal; a signal-to-noise ratio (SNR) estimation module configured to estimate the SNR of the noisy input signal based on the DNN output signal; a mapping module configured to generate a denoising intensity function that correlates the estimated SNR to the denoising intensity; and an application module configured to generate an output signal representing the target signal using the denoising intensity function. Attached Figure Description
[0008] Embodiments of the present invention are described herein in conjunction with the accompanying drawings, in which: Figure 1A This is a schematic diagram illustrating various aspects of a cochlear implant system that can be used to implement the techniques presented herein; Figure 1B It is wearing Figure 1A A side view of the recipient of the sound processing unit of the cochlear implant system; Figure 1C yes Figure 1A A schematic diagram of the components of a cochlear implant system; Figure 1D yes Figure 1A A block diagram of a cochlear implant system; Figure 2 This is a functional block diagram of a noise reduction system according to certain embodiments presented herein; Figure 3 This is a functional block diagram of another noise reduction system according to certain embodiments presented herein; Figure 4 This is a graph illustrating an example mapping function that correlates a priori SNR estimates to noise reduction intensity (mixing ratio, α) according to certain embodiments presented herein; Figure 5This is a graph showing the predicted SNR relative to the true SNR according to certain embodiments presented herein; and Figure 6 This is a flowchart of an example method based on some embodiments presented herein. Detailed Implementation
[0009] Noise reduction techniques, such as deep neural network (DNN)-based noise reduction (sometimes referred to as DNR in this paper), have shown promising improvements in speech understanding and sound quality for hearing-impaired users listening in noisy environments. Specifically referring to DNR, performance depends on several factors, such as network complexity and the scale / quality of its training, and generally varies with the input signal-to-noise ratio (SNR), where performance at high SNR is typically better than at low SNR. At low SNR, it is more challenging for the algorithm to correctly separate clean targets, and significant distortion of the input signal is undesirable.
[0010] To optimize the benefits of DNR, the original input signal can be mixed with the denoised signal (generated based on the output of a deep neural network (DNN)), where the mixing ratio is used to determine the overall strength of the denoising in the final output signal. This makes it useful to vary the denoising strength with the SNR, for example, using a higher DNR strength at high SNR and a lower DNR strength at low SNR. This paper presents a technique for automatically adjusting the strength of the DNR process based on an estimate of the SNR determined based on the DNN output (denoised signal). More specifically, according to the embodiments presented herein, the DNR strength is set based on the SNR, which is determined / estimated (e.g., from the DNN output) from the ratio of the input signal to the denoised signal.
[0011] For ease of illustration, the techniques presented herein will be described primarily with reference to deep neural network (DNN)-based denoising techniques / processes. However, it should be understood that the techniques presented herein can be used with other types of denoising techniques / processes. That is, the reference to deep neural network (DNN)-based denoising techniques is illustrative.
[0012] There are many different types of devices in which / utilizing them can implement embodiments of the present invention. For ease of description only, the techniques presented herein are described primarily with reference to specific devices in the form of cochlear implant systems. However, it should be understood that the techniques presented herein can also be implemented, in part or entirely, by any of many different types of devices, including consumer electronics (e.g., mobile phones), wearable devices (e.g., smartwatches), hearing devices, implantable medical devices, wearable devices, etc. As used herein, the term "hearing device" should be broadly interpreted as any device that delivers sound signals to a user in any form, including acoustic stimulation, mechanical stimulation, electrical stimulation, etc. Thus, hearing devices can be devices for use by people with hearing impairments (e.g., hearing aids, middle ear auditory prostheses, bone conduction devices, direct acoustic stimulators, electroacoustic hearing prostheses, auditory brainstem stimulators, bimodal hearing prostheses, bilateral hearing prostheses, dedicated tinnitus treatment devices, tinnitus treatment device systems, combinations or variations thereof, etc.) or devices for use by people with normal hearing (e.g., consumer devices providing audio streaming, consumer headphones, headphones, and other listening devices). In other examples, the techniques presented herein can be implemented by or in combination with a variety of implantable medical devices, such as vestibular devices (e.g., vestibular implants), visual devices (i.e., bionic eyes), sensors, pacemakers, drug delivery systems, defibrillators, functional electrical stimulation devices, catheters, seizure devices (e.g., devices for monitoring and / or treating epileptic events), sleep apnea devices, electroporation devices, etc.
[0013] Figure 1A-1D An exemplary cochlear implant system 102 is shown that can be used to implement aspects of the techniques presented herein. The cochlear implant system 102 includes an external component 104 configured to be directly or indirectly attached to a user's body, and an internal / implantable component 112 configured to be implanted in or worn on the user's head. Figure 1A-1D In the example, implantable component 112 is sometimes referred to as a "cochlear implant". Figure 1A The image shows a cochlear implant 112 implanted in the user's head 154, while... Figure 1B This is a schematic diagram of an external component 104 worn on the user's head 154. Figure 1C This is another schematic diagram of the cochlear implant system 102, and Figure 1D Further details of the cochlear implant system 102 are shown. For ease of description, they will generally be described together. Figure 1A-1D .
[0014] exist Figure 1A-1DIn the example, the external component 104 includes a sound processing unit 106, an external coil 108, and typically includes a magnet fixed relative to the external coil 108. The cochlear implant 112 includes an implantable coil 114, an implant body 134, and an elongated stimulation component 116 configured for implantation in a user's cochlea. In one example, the sound processing unit 106 is an off-ear (OTE) sound processing unit, sometimes referred to herein as an OTE component, configured to send data and power to the implantable component 112. Generally, an OTE sound processing unit is a component having a generally cylindrical housing 111 and configured to be magnetically coupled to a user's head 154 (e.g., including an integrated external magnet 150 configured to be magnetically coupled to an internal / implantable magnet 152 in the implantable component 112). The OTE sound processing unit 106 also includes an integrated external (head component) coil 108 (external coil 108) configured to be inductively coupled to the implantable coil 114.
[0015] It should be understood that the OTE sound processing unit 106 is merely an illustration of an external device that can operate in conjunction with the implantable component 112. For example, in an alternative example, the external component 104 may include a behind-the-ear (BTE) sound processing unit configured to attach to and be worn adjacent to the recipient's ear. Generally, the BTE sound processing unit includes a housing shaped to be worn on the user's outer ear and connected via a cable to a separate external coil assembly configured to be magnetically and inductively coupled to the implantable coil 114. It should also be understood that alternative external components may be located in the user's ear canal and worn on the body. wait .
[0016] Although the cochlear implant system 102 includes a sound processing unit 106 and a cochlear implant 112, as described below, the cochlear implant 112 can operate independently of the sound processing unit 106 for at least a certain period of time to stimulate the user. For example, the cochlear implant 112 can operate in a first general mode (sometimes referred to as "external hearing mode"), in which the sound processing unit 106 captures sound signals, which are then used as the basis for delivering stimulation signals to the user. The cochlear implant 112 can also operate in a second general mode (sometimes referred to as "invisible hearing" mode), in which the sound processing unit 106 cannot provide sound signals to the cochlear implant 112 (e.g., the sound processing unit 106 is absent, the sound processing unit 106 is powered off, the sound processing unit 106 malfunctions, etc.). Therefore, in invisible hearing mode, the cochlear implant 112 captures the sound signals themselves via an implantable sound sensor and then uses these sound signals as the basis for delivering stimulation signals to the user. Further details regarding the operation of the cochlear implant 112 in external hearing mode are provided below, followed by details regarding the operation of the cochlear implant 112 in invisible hearing mode. It should be understood that the references to external hearing mode and invisible hearing mode are merely illustrative, and the cochlear implant 112 can also be operated in alternating modes.
[0017] exist Figure 1A and 1C In the diagram, the cochlear implant system 102 is shown with an external device 110 configured to implement various aspects of the presented technology. The external device 110 (shown in more detail in Figure 1E) is a computing device, such as a personal computer (e.g., a laptop computer, desktop computer, tablet computer), a mobile phone (e.g., a smartphone), or a remote control unit. wait The external device 110 and the cochlear implant system 102 (e.g., sound processing unit 106 or cochlear implant 112) communicate wirelessly via a two-way communication link 126. The two-way communication link 126 may include, for example, short-range communication, such as a Bluetooth link, a Bluetooth Low Energy (BLE) link, a proprietary link, etc.
[0018] Return to Figure 1A-1D For example, the sound processing unit 106 of the external component 104 also includes one or more input devices configured to capture and / or receive input signals (e.g., audio signals or data signals) at the sound processing unit 106. The one or more input devices include, for example, one or more audio input devices 118 (e.g., one or more external microphones, audio input ports, pickup coils) each located in, on, or near the sound processing unit 106. waitOne or more auxiliary input devices 128 (e.g., audio ports, such as Direct Audio Input (DAI), data ports, such as Universal Serial Bus (USB) ports, cable ports). wait The input devices include a short-range wireless transmitter / receiver (wireless transceiver) 120 (e.g., for communicating with an external device 110). However, it should be understood that one or more input devices may include additional types of input devices and / or fewer input devices (e.g., the short-range wireless transceiver 120 and / or one or more auxiliary input devices 128 may be omitted).
[0019] The sound processing unit 106 also includes an external coil 108, a charging coil 130, a tightly coupled radio frequency transmitter / receiver (RF transceiver) 122, at least one rechargeable battery 132, and an external sound processing module 124. The external sound processing module 124 can be configured to perform a variety of operations. For ease of illustration, in Figure 1D Only operations relevant to the technology presented in this article are shown; other operations have been omitted. Specifically, Figure 1D The diagram shows a deep neural network-based noise reduction (DNR) module 131, an SNR estimation module 133, an SNR mapping module 135, and a mixing module 137. Again, modules 131, 133, 135, and 137 are intended only to represent a subset of functions / operations that can be performed by the external sound processing module 124.
[0020] Each of the DNN module 131, SNR estimation module 133, SNR mapping module 135, and mixing module 137 may be performed by one or more processors (e.g., one or more digital signal processors (DSPs), one or more uC cores) arranged to perform the operations described herein. wait firmware, software wait Formation. That is, modules 131, 133, 135, and 137 can each be implemented as firmware elements, partially or completely implemented using digital logic gates in one or more application-specific integrated circuits (ASICs), or partially or completely implemented in software. wait .although Figure 1D The DNN module 131, SNR estimation module 133, SNR mapping module 135, and mixing module 137 are shown as being implemented / executed at the external sound processing module 124; however, it should be understood that these elements (e.g., functional operations) may also or alternatively be implemented / executed as part of the implantable sound processing module 158, as part of the external device 110, etc. The optional presence of these modules in the implantable sound processing module 158 (which itself may optionally be present or omitted in different arrangements) is determined by… Figure 1D The dashed box in the image represents...
[0021] Return to Figure 1A-1D For example, the implantable component 112 includes an implant body (main module) 134, a lead area 136, and an intracochlear stimulation assembly 116, all configured to be implanted under the user's skin (tissue) 115. The implant body 134 generally includes an hermetically sealed housing 138, which in some examples includes at least one power source 125 (e.g., one or more batteries, one or more capacitors, etc.) 125, within which an RF interface circuitry 140 and a stimulator unit 142 are disposed. The implant body 134 also includes an internal / implantable coil 114, which is generally outside the housing 138 but connected via an hermetically sealed feedthrough (…). Figure 1D (Not shown) is connected to the RF interface circuit system 140.
[0022] As mentioned, the stimulation component 116 is configured to be at least partially implanted in the user's cochlea. The stimulation component 116 includes a plurality of longitudinally spaced intracochlear electrical stimulation contacts (electrodes) 144, which together form a contact array (electrode array) 146 for delivering electrical stimulation (current) to the recipient's cochlea. The stimulation component 116 extends through an opening in the recipient's cochlea (e.g., cochlear fenestration, round window). wait ), and has via lead area 136 and airtight feedthrough ( Figure 1D (Not shown) is connected to the proximal end of the stimulator unit 142. The lead region 136 includes a plurality of conductors (wires) that electrically couple the electrode 144 to the stimulator unit 142. The implantable component 112 also includes an electrode external to the cochlea, sometimes referred to as the external cochlear electrode (ECE) 139.
[0023] As described, the cochlear implant system 102 includes an external coil 108 and an implantable coil 114. An external magnet 150 is fixed relative to the external coil 108, and an internal / implantable magnet 152 is fixed relative to the implantable coil 114. The external magnet 150 and the internal / implantable magnet 152, fixed relative to the external coil 108 and the internal / implantable coil 114 respectively, facilitate operational alignment of the external coil 108 and the implantable coil 114. This operational alignment of the coils enables the external component 104 to transmit data and power to the implantable component 112 via a tightly coupled wireless link 148 formed between the external coil 108 and the implantable coil 114. In some examples, the tightly coupled wireless link 148 is a radio frequency (RF) link. However, various other types of power transfer (e.g., infrared (IR), electromagnetic, capacitive, and inductive transfer) can be used to transfer power and / or data from the external component to the implantable component, and therefore, Figure 1D Only one exemplary arrangement is shown.
[0024] As described above, the sound processing unit 106 includes an external sound processing module 124. The external sound processing module 124 is configured to process input audio signals received (received at one or more input devices, such as sound input device 118 and / or auxiliary input device 128), and to convert the received input audio signals into output control signals for stimulating the first ear of a receiver or user (i.e., the external sound processing module 124 is configured to perform sound processing on the input signals received at the sound processing unit 106). In other words, one or more processors (e.g., processing elements implementing firmware, software, etc.) in the external sound processing module 124 are configured to execute sound processing logic in memory to convert the received input audio signals into output control signals (stimulation signals) representing electrical stimulation to be delivered to a receiver.
[0025] As stated, Figure 1D An embodiment is shown in which an external sound processing module 124 in the sound processing unit 106 generates an output control signal. In an alternative embodiment, the sound processing unit 106 may send less processed information (e.g., audio data) to the implantable component 112, and sound processing operations (e.g., the conversion of input sound to output control signal 156) may be performed by a processor within the implantable component 112.
[0026] exist Figure 1D In this embodiment, according to an exemplary embodiment, an output control signal (stimulation signal) is provided to an RF transceiver 122, which transmits the output control signal (e.g., in an coded manner) transdermally to an implantable component 112 via an external coil 108 and an implantable coil 114. That is, the output control signal (stimulation signal) is received at the RF interface circuitry 140 via the implantable coil 114 and provided to the stimulator unit 142. The stimulator unit 142 is configured to generate an electrical stimulation signal (e.g., a current signal) using the output control signal to be delivered to the user's cochlea via one or more stimulation contacts (electrodes) 144. In this manner, the cochlear implant system 102 electrically stimulates the user's auditory nerve cells, thereby enabling the recipient to perceive one or more components of the input audio signal (received sound signal) in a way that bypasses the missing or defective hair cells that normally translate acoustic vibrations into neural activity.
[0027] As detailed above, in external hearing mode, the cochlear implant 112 receives processed sound signals from the sound processing unit 106. However, in invisible hearing mode, the cochlear implant 112 is configured to capture and process sound signals for electrical stimulation of the user's auditory nerve cells. Specifically, as Figure 1DAs shown, an exemplary embodiment of the cochlear implant 112 may include a plurality of implantable sound sensors 165(1), 165(2) and an implantable sound processing module 158, the plurality of implantable sound sensors collectively forming a sensor array 160. Similar to the external sound processing module 124, the implantable sound processing module 158 may include, for example, one or more processors and a memory device (memory) including sound processing logic. The memory device may include any or more of the following: non-volatile memory (NVM), ferroelectric random access memory (FRAM), read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical or other physical / tangible memory storage device. The one or more processors are, for example, microprocessors or microcontrollers that execute instructions for sound processing logic stored in the memory device.
[0028] In invisible hearing mode, the implantable sound sensors 165(1), 165(2) of sensor array 160 are configured to detect / capture input sound signals 166 (e.g., acoustic sound signals, vibrations, etc.), which are provided to implantable sound processing module 158. Implantable sound processing module 158 is configured to convert the received input sound signals 166 (received at one or more implantable sound sensors 165(1), 165(2)) into output control signals 156 for stimulating the first ear of the recipient or user (i.e., implantable sound processing module 158 is configured to perform sound processing operations). In other words, one or more processors (e.g., multiple processing elements implementing firmware, software, etc.) in implantable sound processing module 158 are configured to execute sound processing logic in memory to convert the received input sound signals 166 into output control signals 156 provided to stimulator unit 142. The stimulator unit 142 is configured to generate an electrical stimulation signal (e.g., a current signal) for delivery to the user's cochlea using an output control signal 156, thereby bypassing missing or defective hair cells that typically translate acoustic vibrations into neural activity.
[0029] It should be understood that the above descriptions of the so-called external hearing mode and the so-called invisible hearing mode are merely illustrative, and the cochlear implant system 102 may operate in different ways in different embodiments. For example, in an alternative embodiment of the external hearing mode, the cochlear implant 112 may use signals captured by the sound input device 118 and the implantable sound sensors 165(1), 165(2) of the sensor array 160 to generate stimulation signals for delivery to the user.
[0030] As mentioned above, deep neural network (DNN)-based noise reduction (DNR) is a noise reduction technique that has shown good improvement in speech understanding and sound quality for hearing device users listening in noisy environments. However, the performance of DNR depends on several factors, such as network complexity and the scale / quality of its training, and can vary with the input signal-to-noise ratio (SNR) (e.g., DNR generally performs well for signals with relatively high SNR but relatively poorly for signals with relatively low SNR). Specifically, correctly separating the target signal (e.g., speech) is more challenging with signals having relatively low SNR, leading to undesirable and severe distortion of the input signal.
[0031] To optimize the benefits of DNR, the original input signal is mixed with the denoised signal (generated based on the DNN output), where the mixing ratio is used to determine the overall strength of the denoising in the final output signal. This makes it useful to vary the denoising strength based on the SNR (e.g., by changing or controlling the mixing ratio). For example, the mixing ratio can produce a higher DNR strength at a high SNR and a lower DNR strength at a low SNR. This paper presents a technique for automatically adjusting the denoising strength (DNR strength) based on an estimate of the SNR determined based on the DNN output. More specifically, according to the embodiments presented herein, the DNR strength is a function of the SNR, which is determined / estimated from the ratio of the input signal to the denoised signal.
[0032] Figure 2 This is a functional block diagram illustrating an example deep neural network (DNN) based noise reduction (DNR) system 262 according to certain embodiments presented herein. Specifically, as Figure 2 As shown, system 262 includes a deep neural network (DNN) denoising module 264 (sometimes referred to as a DNN module), a signal-to-noise ratio (SNR) estimation module 266, an SNR mapping module 268, and a hybrid module 270. The operation of these various functional blocks, along with other aspects of system 262, is described below.
[0033] like Figure 2 As shown, system 262 receives a noisy input x(t), which is defined as clean speech s(t) embedded in additive noise d(t), where t This is a time index. This situation is illustrated in Equation 1 below.
[0034] (Equation 1) At filter block 261, using overlapping window frames with frame n and frequency index k, a Fourier transform is applied to generate a frequency domain representation of the noisy input (x(t)). The frequency domain representation is given below in Equation 2.
[0035] (Equation 2) As shown in the figure, provide to DNN module 264 (Frequency domain representation of the input signal). DNN module 264 from Calculate the DNN spectral gain H n (k) (e.g., the DNN output signal or noise reduction gain), which is applied at block 263 This block then generates a clean / noise-reduced output signal. In other words, the noise reduction signal ( ) is achieved by inputting a noisy input signal ( ) Applying DNN spectral gain (H n (k) is generated. This case is given below in Equation 3.
[0036] (Equation 3) like Figure 2 As shown, the pure / denoised signal is mixed according to the mixing ratio or the alpha (α) parameter of the denoising intensity. ) and the original noisy input ( Generate the final output signal Z n (k), the mixing ratio or noise reduction intensity parameter can vary between values 0 and 1 with time and frequency. This mixing process is represented by Equation 4 below.
[0037] (Equation 4) Finally, in this specific example, via The inverse FFT is used to obtain the time-domain output signal Z. n (t). In other examples, such as in the context of cochlear implants, inverse FFT is not required, and the frequency domain output ( () is used for subsequent processing operations.
[0038] Based on the techniques presented in this article, the noise reduction intensity parameter or hybrid parameter α n (k) is defined as the prior SNR ( The prior SNR is a function of ), which needs to be estimated by the algorithm, as shown in Equation 5 below.
[0039] (Equation 5) In other words, the noise reduction intensity parameter / mixing ratio varies with the input signal's SNR, but the variation is non-linear (e.g., it changes as a function of SNR with time and frequency). To derive the instantaneous prior SNR ( The estimated value of ) is used as a starting point, as defined below in Equation 6.
[0040] (Equation 6) In equation 6, E{.} is the expectation operator, and the DNN outputs |Y| 2 Used as speech power |S| 2 The estimated value is assumed to be completely cleaned by the DNN and dominated by the target speech, and can be directly observed after applying the DNN gain. This is shown in Equation 7 below.
[0041] (Equation 7) Noise power |D| 2 The difference between the noisy input (X) and the clean speech estimate (Y) is estimated, assuming that the speech (S) and noise (D) sources are uncorrelated, and therefore the cross power spectrum SD is zero, as shown in Equation 8.
[0042] (Equation 8) The prior SNR estimate obtained by substituting the DNN-based input (X) and output (Y) is shown in Equation 9 below.
[0043] (Equation 9) In other words, in Figure 2 In the example, the SNR of the signal is determined based on the original noisy input signal (X) and the denoised signal (Y) output by the DNN module 264, and more specifically, based on the ratio of the original noisy input signal (X) to the denoised signal (Y). In other words, the SNR estimate is a function of the ratio of the original noisy input signal (X) to the denoised signal (Y). Figure 2 In the example, the final output (Z) is determined from the original noise input signal (X) and the denoised signal (Y) output by the DNN module 264, as well as the denoising intensity parameter / mixing ratio (α).
[0044] As mentioned above, in Figure 2 In the example, the denoised signal (Y) is generated at 263 and used to generate the estimated SNR. Figure 3 This is a functional block diagram illustrating an alternative arrangement in which no noise-reduced signal (Y) is generated, and the estimated SNR is estimated directly from the output spectral gain (H) (noise reduction gain) of the DNN module 264.
[0045] More specifically, Figure 3 The diagram shown is a functional block diagram illustrating an example deep neural network (DNN) based noise reduction (DNR) system 362 according to certain embodiments presented herein. Specifically, as Figure 3As shown, system 362 includes a deep neural network (DNN) denoising module 364 (sometimes referred to as the DNN module), a signal-to-noise ratio (SNR) estimation module 366, an SNR mapping module 368, and a hybrid module 370. The operation of these various functional blocks, along with other aspects of system 362, is described below.
[0046] like Figure 3 As shown, the noise reduction system 362 receives a noisy input x(t), which is defined as clean speech s(t) embedded in additive noise d(t), where t For the time index (Equation 1 above). At filter block 361, using overlapping window frames with frame n and frequency index k, a Fourier transform is applied to generate a frequency domain representation of the noisy input (x(t)). The frequency domain representation is given in Equation 2 above.
[0047] As shown in the figure, the DNN module 364 is provided with (Frequency domain representation of the input signal). DNN module 364 from Calculate the DNN spectral gain H n (k). As shown in the figure, by inputting the original noisy input ( ) Applying DNN spectral gain H n (k), and then the denoised signal obtained by mixing with the original noisy input according to the mixing ratio or alpha (α) parameter is used to generate the final output signal Z. n (k), the mixing ratio or noise reduction intensity parameter can vary between values 0 and 1 with time and frequency. By substituting Y=HX into Equation 4 above, Z can be expressed in terms of H, X, and α without the signal Y. This is shown in Equation 10 below.
[0048] (Equation 10) Finally, in this specific example, via The inverse FFT is used to obtain the time-domain output signal Z. n (t). In other examples, such as in the context of cochlear implants, inverse FFT is not required, and the frequency domain output ( () is used for subsequent processing operations.
[0049] according to Figure 3 For example, substituting Y = HX and H = Y / X into Equation 9, the prior SNR estimate can be obtained directly from the DNN gain (H) without the need for the denoised signal (Y). This is shown below in Equation 11.
[0050] (Equation 11) In other words, in Figure 3 In the example, the SNR of the signal is determined directly from the output gain (H) of the DNN module 364. However, the output gain (H) represents the ratio of the original noisy input signal (X) to the denoised signal (Y). In other words, in Figure 3 In this context, the SNR estimate remains a function of the ratio of the original noisy input signal (X) to the denoised signal (Y) (expressed as the output gain (H) of the DNN module). Figure 3 In the example, the final output (Z) is determined from the output gain (H) and the noise reduction intensity parameter / mixing ratio (α).
[0051] In short, Figure 2 and Figure 3 The paper presents the denoising technique in two mathematically equivalent configurations, where in Figure 2 In the middle, signal Y is generated, and Figure 3 Explicit calculation of signal Y is not used. Instead, in Figure 3 In this study, the SNR estimate is based solely on the output gain (H) of the DNN module.
[0052] As described above, the SNR will be estimated (regardless of whether it is referenced as described above). Figure 2 still Figure 3 (Defined) Used to control the noise reduction intensity parameter / mixing ratio α n (k) (e.g., α) n (k) is defined as the prior SNR ( (function of ). Figure 4 This is a graph illustrating an example mapping function 472 that correlates a priori SNR estimate to the noise reduction intensity (e.g., showing how the noise reduction intensity parameter / mixing ratio varies relative to SNR). The noise reduction intensity parameter / mixing ratio (alpha (α)) can vary between values 0 and 1 with time and frequency and is based on the SNR setting of the input signal. It should be understood that... Figure 4 The mapping function 472 shown is merely illustrative, and the mapping function may have different shapes for different receivers.
[0053] In some examples, the SNR estimate can be smoothed over an appropriate time period, corresponding to a time course suitable for causing changes in the mix ratio. An optional frequency integral can also be used, employing a frequency importance scale, such as the band importance function defined in ANSI S3.5–1997, “American National Standard Methods for the Calculation of the Speech Intelligibility Index,” New York: ANSI, 1997. Smoothing of the SNR estimate is described below, but smoothing can be similarly applied to other signals in the chain, such as noisy input (X), clean speech (Y), SNR, noise reduction intensity parameter / mix ratio, etc.
[0054] As mentioned above, for example, the SNR estimate can be smoothed over time by using a rectangular window averaging method, taking the mean of the last T time frames, as shown in Equation 12 below.
[0055] (Equation 12) Alternatively, first-order IIR exponential smoothing can be used, which is performed using the following difference equation with a forgetting factor β, as shown in Equation 13 below.
[0056] (Equation 13) The SNR estimate can be integrated over K frequency regions using the Speech Intelligibility Importance Function (SII). The Speech Intelligibility Importance Function provides a relative weight for each frequency region, and these relative weights are normalized so that the sum of all weights equals zero. As shown in Equation 14 below.
[0057] (Equation 14) Figure 5 An example of SNR estimation is shown, which uses smoothing configured for averaging rectangular windows over 5 seconds, and SII(k) = 1 / K (providing equal contributions from each frequency grid). Figure 5 The relationship between prior SNR (true SNR) and its estimated (predicted SNR) is shown on a wide database of speech samples in noisy environments (168 TIMIT sentences per SNR, using various noise types from YouTube, including noisy human voices, restaurant noise, cafe noise, city noise, etc.).
[0058] In some examples, the techniques presented herein leverage user interaction with the system for the purpose of designing and fine-tuning the noise reduction intensity parameter / mixing ratio (e.g., adjusting the mapping function that maps prior SNR estimates to noise reduction intensity) to suit individual requirements. In other words, the user can control, at least to some extent, the shape of the mapping function between the prior SNR estimate and the intensity parameter (α). For example, the system can receive one or more user input signals adjusting the relationship between the mixing ratio and SNR, and the SNR mapping module can determine the mixing ratio based on the SNR and one or more user input signals. Such a system would allow the user to control the mixing ratio value (α) through an appropriate user interface, such as a slider or dial pad interface on a smartphone application. When under manual oversight, this function itself is updated based on the then-current estimated SNR using the user-selected noise reduction intensity parameter / mixing ratio (α). Figure 2 and Figure 3 The dashed arrows 275 and 375 show this optional user control over the noise reduction intensity parameter / mixing ratio (e.g., receiving a user input signal to adjust the relationship between the mixing ratio and SNR).
[0059] In some embodiments, the noise reduction system according to the embodiments presented herein may use different mapping functions for different use cases. For example, different mapping functions may be used for different noise types, different sound environments, different implementations of DNR, or for different end users with unique needs and / or preferences. More generally, the techniques presented herein may use metrics other than SNR, which are calculated and used to at least partially control the strength of the noise reduction intensity parameter. Examples include things like root mean square (RMS) signal level, reverberation characteristics, and other signal statistics / properties.
[0060] Figure 6 This is a flowchart of an example method 680 according to certain embodiments presented herein. Method 680 begins at 682, where at least one input signal (e.g., one or more audio signals) is received at one or more input devices. At 684, the at least one input signal is processed using a deep neural network (DNN)-based denoising process to generate a spectral gain for application to the at least one input signal. At 686, the signal-to-noise ratio (SNR) of the at least one input signal is determined based on the spectral gain (the output of the DNN-based denoising process).
[0061] In some aspects, a spectral gain is applied to at least one input signal to generate at least one denoised signal. Additionally, a denoising intensity parameter is generated based on the SNR, and the at least one denoised signal is mixed with at least one input signal based on the denoising intensity parameter to generate at least one output signal. This at least one output signal can be used for subsequent processing operations. For example, in some embodiments, the at least one input signal is an audio signal, and the at least one output signal is an output signal used for subsequent processing and in delivering auditory perception to a user.
[0062] according to Figure 6 In some embodiments, a spectral gain is applied to at least one input signal to generate at least one denoised signal, and the SNR of the at least one input signal is determined based on the ratio of the at least one input signal to the at least one denoised signal. Figure 6 In some embodiments, the SNR of at least one input signal is determined directly from the spectral gain (e.g., without first generating at least one denoised signal). In such embodiments, the spectral gain represents / relates to a function of the ratio of at least one input signal to at least one denoised signal, so the SNR of at least one input signal is still a function of the ratio of at least one input signal to at least one denoised signal.
[0063] exist Figure 6 In some embodiments, the SNR is an instantaneous SNR based on spectral gain estimation. In other embodiments, the SNR is a smoothed SNR estimated over a period of time and / or estimated over a frequency range.
[0064] exist Figure 6 In some embodiments, the noise reduction intensity parameter is a function of SNR and one or more user input signals. That is, in some examples, one or more user input signals are received, which control the estimation of the relationship between SNR and noise reduction intensity. The noise reduction intensity parameter is then adjusted based on the one or more user input signals (i.e., relating SNR to a function of noise reduction intensity). As described above, this optional user control over the noise reduction intensity parameter / mixing ratio (in...) Figure 2 and Figure 3 (As shown by dashed arrows 275 and 375) the relationship between the mix ratio and SNR can be adjusted according to the receiver's preferences. For example, the receiver can adjust the mix ratio in different sound environments with different sound types, etc.
[0065] It should be understood that while the specific uses of this technology have been described and discussed above, the disclosed technology can be used with various apparatuses according to many examples of this technology. The foregoing discussion is not intended to suggest that the disclosed technology is only suitable for implementation in systems similar to those shown in the accompanying drawings. In general, additional configurations can be used to practice the processes and systems described herein, and / or some aspects can be excluded without departing from the processes and systems disclosed herein.
[0066] This disclosure describes some aspects of the invention with reference to the accompanying drawings, which illustrate only some possible aspects. However, other aspects may be embodied in many different forms and should not be construed as limited to those set forth herein. Rather, these aspects are provided to make this disclosure exhaustive and complete and to fully convey the scope of possible aspects to those skilled in the art.
[0067] It should be understood that this document is not intended to limit the system and process to the specific aspects described with respect to the various aspects (e.g., parts, components, etc.) in relation to the accompanying drawings. Therefore, the methods and systems described herein can be practiced using additional configurations, and / or some aspects described may be excluded without departing from the methods and systems disclosed herein.
[0068] According to some aspects, a system and a non-transitory computer-readable storage medium are provided. The system is configured with hardware configured to perform operations similar to those of the present disclosure. One or more non-transitory computer-readable storage media include instructions that, when executed by one or more processors, cause one or more processors to perform operations similar to those of the present disclosure.
[0069] Similarly, where the steps of a process are disclosed, these steps are described for illustrative purposes of the method and system and are not intended to limit this disclosure to a particular sequence of steps. For example, these steps may be performed in a different order, two or more steps may be performed simultaneously, additional steps may be performed, and the disclosed steps may be excluded without departing from this disclosure. Furthermore, the disclosed process may be repeated.
[0070] While specific aspects have been described herein, the scope of this technology is not limited to those specific aspects. Those skilled in the art will recognize other aspects or modifications within the scope of this technology. Therefore, specific structures, actions, or media are disclosed only as illustrative aspects. The scope of this technology is defined by the appended claims and any equivalents thereof.
[0071] It should also be understood that the embodiments presented herein are not mutually exclusive, and various embodiments can be combined with one embodiment in any of a variety of different ways.
Claims
1. A method comprising: Receive at least one input signal at one or more input devices; The at least one input signal is processed using a noise reduction process to generate a spectral gain for application to the at least one input signal; as well as The signal-to-noise ratio (SNR) of the at least one input signal is determined based on the spectral gain.
2. The method of claim 1, wherein determining the SNR of the at least one input signal based on the spectral gain comprises: Apply the spectral gain to the at least one input signal to generate at least one noise-reduced signal; as well as The SNR of the at least one input signal is determined based on the ratio of the at least one input signal to the at least one noise-reduced signal.
3. The method according to claim 2, further comprising: Determine the noise reduction intensity parameters based on the SNR; as well as The at least one noise-reduced signal is mixed with the at least one input signal according to the noise reduction intensity parameter to generate an output signal.
4. The method according to claim 3, further comprising: The noise reduction intensity parameter is determined based on the SNR and one or more user inputs.
5. The method of claim 1, wherein determining the SNR of the at least one input signal based on the spectral gain comprises: The SNR is determined directly from the spectral gain.
6. The method of claim 5, further comprising: Determine the noise reduction intensity parameters based on the SNR; as well as The output signal is generated from the spectral gain and the noise reduction intensity parameter.
7. The method according to claim 6, further comprising: The noise reduction intensity parameter is determined based on the SNR and one or more user inputs.
8. The method according to claim 1, 2, 3, 4, 5, 6 or 7, wherein determining the SNR of the at least one input signal based on the spectral gain comprises: Generate the instantaneous SNR based on the spectral gain estimate.
9. The method according to claim 1, 2, 3, 4, 5, 6 or 7, wherein determining the SNR of the at least one input signal based on the spectral gain comprises: Determine the smoothed SNR estimated over a period of time.
10. The method according to claim 1, 2, 3, 4, 5, 6 or 7, wherein determining the SNR of the at least one input signal based on the spectral gain comprises: Determine the smoothed SNR estimated over the frequency range.
11. The method according to claim 1, 2, 3, 4, 5, 6 or 7, wherein receiving the at least one input signal at one or more input devices comprises: Receive at least one audio signal at one or more audio inputs.
12. The method according to claim 1, 2, 3, 4, 5, 6 or 7, wherein the at least one input signal is a time-domain signal, and wherein the method comprises: Before processing the at least one input signal using a DNN-based denoising process, the at least one input signal is converted into a frequency domain representation using overlapping window frames.
13. The method according to claim 1, 2, 3, 4, 5, 6 or 7, wherein processing the at least one input signal using a noise reduction process comprises: The at least one input signal is processed using a noise reduction process based on a deep neural network (DNN).
14. One or more non-transitory computer-readable storage media, the one or more non-transitory computer-readable storage media comprising instructions that, when executed by one or more processors of a hearing device, cause the one or more processors to: Calculate the noise reduction gain for application to one or more sound signals received at the sound input of the hearing device; The signal-to-noise ratio (SNR) of the one or more audio signals is determined based on the noise reduction gain. Apply the noise reduction gain to the one or more audio signals to generate one or more noise-reduced signals; as well as The one or more noise-reduced signals are mixed with the one or more audio signals according to a mixing ratio, wherein the mixing ratio is a function of the SNR.
15. The one or more non-transitory computer-readable storage media of claim 14, wherein the instructions causing the one or more processors to calculate the noise reduction gain include instructions that, when executed, cause the one or more processors to perform the following operations: The one or more audio signals are processed using a noise reduction process based on a deep neural network (DNN) to generate the noise reduction gain.
16. One or more non-transitory computer-readable storage media according to claim 14 or 15, wherein the instruction causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain includes instructions that, when executed, cause the one or more processors to perform the following operations: Apply the noise reduction gain to the one or more audio signals to generate the one or more noise-reduced signals; and The SNR of the one or more audio signals is determined based on the ratio of the one or more audio signals to the one or more noise-reduced signals.
17. One or more non-transitory computer-readable storage media according to claim 14 or 15, wherein the instructions for causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain include instructions that, when executed, cause the one or more processors to perform the following operations: Generate an instantaneous SNR based on the noise reduction gain estimate.
18. One or more non-transitory computer-readable storage media according to claim 14 or 15, wherein the instructions for causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain include instructions that, when executed, cause the one or more processors to perform the following operations: Generate a smoothed SNR estimated over a period of time.
19. One or more non-transitory computer-readable storage media according to claim 14 or 15, wherein the instruction causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain includes instructions that, when executed, cause the one or more processors to perform the following operations: Generate a smoothed SNR estimated over the frequency range.
20. The one or more non-transitory computer-readable storage media of claim 14 or 15, further comprising instructions for causing the one or more processors to perform the following operations: Receive one or more user input signals that adjust the relationship between the mixing ratio and the SNR; and The mixing ratio is determined based on the SNR and based on one or more user input signals.
21. One or more non-transitory computer-readable storage media, the one or more non-transitory computer-readable storage media comprising instructions that, when executed by one or more processors of a hearing device, cause the one or more processors to: Calculate the noise reduction gain for application to one or more sound signals received at the sound input of the hearing device; The signal-to-noise ratio (SNR) of the one or more audio signals is determined based on the noise reduction gain. Determine the noise reduction intensity parameters based on the SNR; as well as An output signal is generated from the noise reduction gain and the noise reduction intensity parameter.
22. The one or more non-transitory computer-readable storage media of claim 21, wherein the instructions causing the one or more processors to calculate the noise reduction gain include instructions that, when executed, cause the one or more processors to perform the following operations: The one or more audio signals are processed using a noise reduction process based on a deep neural network (DNN) to generate the noise reduction gain.
23. The one or more non-transitory computer-readable storage media according to claim 21 or 22, wherein the instruction causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain includes instructions that, when executed, cause the one or more processors to perform the following operations: Generate an instantaneous SNR based on the noise reduction gain estimate.
24. The one or more non-transitory computer-readable storage media of claim 21, wherein the instructions for causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain include instructions that, when executed, cause the one or more processors to perform the following operations: Generate a smoothed SNR estimated over a period of time.
25. One or more non-transitory computer-readable storage media according to claim 21 or 22, wherein the instructions for causing the one or more processors to determine the SNR of the one or more audio signals based on the noise reduction gain include instructions that, when executed, cause the one or more processors to perform the following operations: Generate a smoothed SNR estimated over the frequency range.
26. The one or more non-transitory computer-readable storage media according to claim 21 or 22, further comprising instructions for causing the one or more processors to perform the following operations: Receive one or more user input signals that adjust the relationship between the noise reduction intensity parameter and the SNR; and The noise reduction intensity parameter is determined based on the SNR and based on one or more user input signals.
27. An apparatus comprising: One or more input devices configured to receive a noisy input signal, wherein the noisy input signal includes a target signal embedded in additive noise; A noise reduction module based on a deep neural network (DNN), the noise reduction module being configured to generate a DNN output signal from the noisy input signal; A signal-to-noise ratio (SNR) estimation module, configured to estimate the SNR of the noisy input signal based on the DNN output signal; A mapping module, configured to generate a noise reduction strength function that correlates the estimated SNR with the noise reduction strength; as well as An application module is configured to use the noise reduction intensity function to generate an output signal representing the target signal.
28. The apparatus of claim 27, wherein the DNN output signal is applied to the noisy input signal to generate a denoised signal, and wherein the SNR estimation module is configured to: The SNR of the noisy input signal is estimated based on the ratio of the noisy input signal to the denoised signal.
29. The apparatus of claim 27, wherein the SNR estimation module is configured to estimate the SNR directly from the DNN output signal.
30. The apparatus of claim 27, 28 or 29, wherein the SNR estimation module is configured to generate an instantaneous SNR estimated based on the DNN output signal.
31. The apparatus of claim 27, 28 or 29, wherein the SNR estimation module is configured to generate a smoothed SNR estimated over a period of time.
32. The apparatus of claim 27, 28 or 29, wherein the SNR estimation module is configured to generate a smoothed SNR estimated over a frequency range.
33. The apparatus of claim 27, 28, or 29, wherein the mapping module is configured to: Receive one or more user input signals that control the relationship between the estimated SNR and the noise reduction intensity; and The noise reduction intensity function is adjusted based on one or more user input signals.
34. The apparatus of claim 27, 28 or 29, wherein the one or more input devices include a voice input device.
35. The apparatus of claim 27, 28 or 29, wherein the DNN output signal is used to generate a denoised signal, and wherein the application module is configured to generate the output signal based on the noisy input signal, the denoised signal and the denoising intensity function.
36. The apparatus of claim 27, 28 or 29, wherein the application module is configured to generate the output signal based on the DNN output signal and the noise reduction intensity function.