Automatic speech recognition in sound processing

ASR integration in cochlear implants addresses the challenge of speech perception in noise by generating augmented stimulation patterns for improved phonemic information processing, enhancing speech intelligibility.

WO2025229471A1PCT designated stage Publication Date: 2025-11-06COCHLEAR LIMITED
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/054332
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2025-04-25
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Cochlear implant recipients face challenges in understanding speech in noisy environments due to limited sound information and reduced temporal and spectral resolution, leading to difficulties in perceiving phonemes.

Method used

Integration of automatic speech recognition (ASR) into cochlear implants to enhance phonemic information processing by converting sound signals into predicted phoneme stimulation patterns, which are then used to generate augmented stimulation patterns for improved speech perception.

Benefits of technology

Enhances speech intelligibility in noise for cochlear implant users by providing more precise phonemic information through ASR-based sound processing, improving neural activation patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025054332_06112025_PF_FP_ABST
    Figure IB2025054332_06112025_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are techniques for improving speech perception for hearing device recipients through the use of phonemic information. For example, in certain embodiments, automatic speech recognition (ASR) is used to enhance the perception of phonemic information by a recipient. Such techniques can, for example, provide improved speech intelligibility in noise for hearing device recipients.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATIC SPEECH RECOGNITION IN SOUND PROCESSINGBACKGROUNDField of the Invention[oooi] The present invention relates generally to techniques for use of automatic speech recognition in sound processing.Related Art

[0002] Medical devices have provided a wide range of therapeutic benefits to recipients over recent decades. Medical devices can include internal or implantable components / devices, external or wearable components / devices, or combinations thereof (e.g., a device having an external component communicating with an implantable component). Medical devices, such as traditional hearing aids, partially or fully-implantable hearing prostheses (e.g., bone conduction devices, mechanical stimulators, cochlear implants, etc.), pacemakers, defibrillators, functional electrical stimulation devices, and other medical devices have been successful in performing lifesaving and / or lifestyle enhancement functions and / or recipient monitoring for a number of years.

[0003] The types of medical devices and the ranges of functions performed thereby have increased over the years. For example, many medical devices, sometimes referred to as “implantable medical devices,” now often include one or more instruments, apparatus, sensors, processors, controllers or other functional mechanical or electrical components that are permanently or temporarily implanted in a recipient. These functional devices are typically used to diagnose, prevent, monitor, treat, or manage a disease / injury or symptom thereof, or to investigate, replace or modify the anatomy or a physiological process. Many of these functional devices utilize power and / or data received from external devices that are part of, or operate in conjunction with, implantable components.SUMMARY

[0004] In one aspect, a method is provided. The method comprises: receiving one or more sound signals at a hearing device of a recipient; determining a direct stimulation pattern based on the one or more sound signals; determining a predicted phoneme stimulation pattern based on the one or more sound signals; and generating an augmented stimulation pattern for delivery to the recipient based at least on the direct stimulation pattern and the predicted phoneme stimulation pattern.

[0005] In another aspect, a method is provided. The method comprises: receiving one or more input signals at an implantable medical device; determining one or more predicted phoneme codes based on one or more phonemes of speech present in the one or more input signals; and generating an augmented stimulation pattern for delivery to a recipient of the implantable medical device based on the one or more input signals and the one or more predicted phoneme codes.

[0006] In another aspect, one or more non-transitory computer readable storage media are provided. The one or more non-transitory computer readable storage comprise instructions that, when executed by one or more processors of a hearing device associated with a recipient, are operable to: obtain one or more sound signals that include speech in a first language of a speaker; convert the one or more sound signals into speaker language text representing the speech in the first language of the speaker; translate the speaker language text into recipient language text in a second language of the recipient; determine one or more speaker characteristics associated with the speaker; and generate a translated stimulation pattern for delivery to the recipient based on the recipient language text in the second language of the recipient and the one or more speaker characteristics associated with the speaker, wherein the translated stimulation pattern represents translated speech in the second language of the recipient.

[0007] In another aspect, a hearing device is provided. The hearing device comprises: one or more sound inputs configured to receive at least one sound signal; a memory; and one or more processors configured to: determine a direct stimulation pattern from the at least one sound signal, determine one or more phonemes of speech present in the at least one sound signal, determine a predicted phoneme stimulation pattern based on one or more phonemes of speech present in the at least one sound signal, generate an augmented stimulation pattern based on the direct stimulation pattern and the predicted phoneme stimulation pattern, and generate one or more stimulation signals based on the representation of the augmented stimulation pattern.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Embodiments of the present invention are described herein in conjunction with the accompanying drawings, in which:

[0009] FIG. 1A is a schematic diagram illustrating a cochlear implant system with which aspects of the techniques presented herein can be implemented;[ooio] FIG. IB is a side view of a recipient wearing a sound processing unit of the cochlear implant system of FIG. 1A;[ooii] FIG. 1C is a schematic view of components of the cochlear implant system of FIG. 1 A;

[0012] FIG. ID is a block diagram of the cochlear implant system of FIG. 1A;

[0013] FIG. IE is a schematic diagram illustrating a computing device with which aspects of the techniques presented herein can be implemented;

[0014] FIG. 2A is a block diagram of an automatic speech recognition (ASR)-based sound processor with which aspects of the techniques presented herein can be implemented;

[0015] FIG. 2B is a block diagram of a sound coding module, in accordance with certain embodiments presented herein;

[0016] FIG. 3 is a block diagram of a specific example implementation of an ASR-based sound processor of FIG. 2, according to an example embodiment;

[0017] FIG. 4 is a block diagram of another specific example implementation of the ASR- based sound processor of FIG. 2 with an integrated smart mixing algorithm, according to an example embodiment;

[0018] FIG. 5 is a block diagram of another specific example implementation of the ASR- based sound processor of FIG. 2 with an integrated deep neural network (DNN) based sound coding module, according to an example embodiment;

[0019] FIGs. 6A-6B are block diagrams of specific example implementations of the ASR- based sound processor of FIG. 2 with an integrated translator algorithm, according to an example embodiment;

[0020] FIG. 7 is a diagram of a specific example implementation of the ASR-based sound processor of FIG. 4 using an audio-visual automatic speech recognition module with the smart mixing algorithm, according to an example embodiment;

[0021] FIG. 8 is a diagram of a specific example implementation of the ASR-based sound processor of FIG. 4 using speaker-dependent phoneme templates with the smart mixing algorithm, according to an example embodiment;

[0022] FIG. 9 is a diagram of a specific example implementation of the ASR-based sound processor of FIG. 2 using phoneme predictions of the automatic speech recognition module to control a noise reduction algorithm, according to an example embodiment;

[0023] FIG. 10 is a flowchart of a method according to an example embodiment of the techniques presented herein; and

[0024] FIG. 11 is a flowchart of a method according to an example embodiment of the techniques presented herein.DETAILED DESCRIPTION

[0025] Presented herein are techniques for improving speech perception for hearing device recipients through the use of phonemic information. For example, in certain embodiments, automatic speech recognition (ASR) is used to enhance the perfection of phonemic information by a recipient. Such techniques can, for example, provide improved speech intelligibility in noise for bearing device recipients.

[0026] There are a number of different types of devices in / with which embodiments of the present invention may be implemented. Merely for ease of description, the techniques presented herein are primarily described with reference to a specific auditory prosthesis system in the form of a cochlear implant system. However, it is to be appreciated that the techniques presented herein may also be partially or fully implemented by any of a number of different types of devices, including hearing devices, implantable medical devices, consumer electronic devices (e.g., mobile phones), wearable devices (e.g., smart watches), etc. As used herein, the term “hearing device” is to be broadly construed as any device that acts on an acoustical perception of an individual, including to improve perception of sound signals, to reduce perception of sound signals, etc. In particular, a hearing device can deliver sound signals to a user in any form, including in the form of acoustical stimulation, mechanical stimulation, electrical stimulation, etc., and / or can operate to suppress all or some sound signals. As such, a hearing device can be a device for use by a hearing-impaired person (e.g., hearing aids, middle ear auditory prostheses, bone conduction devices, direct acoustic stimulators, electro-acoustic hearing prostheses, auditory brainstem stimulators, bimodal hearing prostheses, bilateral hearing prostheses, dedicated tinnitus therapy devices, tinnitus therapy device systems,combinations or variations thereof, etc.), a device for use by a person with normal hearing (e.g., consumer devices that provide audio streaming, consumer headphones, earphones, and other listening devices), a hearing protection device, etc. In other examples, the techniques presented herein can be implemented by, or used in conjunction with, various implantable medical devices, such as visual devices (i.e., bionic eyes), sensors, pacemakers, drug delivery systems, defibrillators, functional electrical stimulation devices, catheters, seizure devices (e.g., devices for monitoring and / or treating epileptic events), sleep apnea devices, electroporation devices, etc.

[0027] FIGs. 1A-1D illustrate an example cochlear implant system 102 with which aspects of the techniques presented herein can be implemented. The cochlear implant system 102 comprises an external component 104 that is configured to be directly or indirectly attached to the body of the user, and an intemal / implantable component 112 that is configured to be implanted in or worn on the head of the user. In the examples of FIGs. 1A-1D, the implantable component 112 is sometimes referred to as a “cochlear implant.” FIG. 1A illustrates the cochlear implant 112 implanted in the head 154 of a user, while FIG. IB is a schematic drawing of the external component 104 worn on the head 154 of the user. FIG. 1C is another schematic view of the cochlear implant system 102, while FIG. ID illustrates further details of the cochlear implant system 102. For ease of description, FIGs. 1A-1D will generally be described together.

[0028] In the examples of FIGs. 1A-1D, the external component 104 comprises a sound processing unit 106, an external coil 108, and generally, a magnet fixed relative to the external coil 108. The cochlear implant 112 includes an implantable coil 114, an implant body 134, and an elongate stimulating assembly 116 configured to be implanted in the user’s cochlea. In one example, the sound processing unit 106 is an off-the-ear (OTE) sound processing unit, sometimes referred to herein as an OTE component, that is configured to send data and power to the implantable component 112. In general, an OTE sound processing unit is a component having a generally cylindrically shaped housing 111 and which is configured to be magnetically coupled to the user’s head 154 (e.g., includes an integrated external magnet 150 configured to be magnetically coupled to an intemal / implantable magnet 152 in the implantable component 112). The OTE sound processing unit 106 also includes an integrated external (headpiece) coil 108 (the external coil 108) that is configured to be inductively coupled to the implantable coil

[0029] It is to be appreciated that the OTE sound processing unit 106 is merely illustrative of the external devices that could operate with implantable component 112. For example, in alternative examples, the external component 104 may comprise a behind-the-ear (BTE) sound processing unit configured to be attached to, and worn adjacent to, the recipient’s ear. A BTE sound processing unit comprises a housing that is shaped to be worn on the outer ear of the user. In certain examples, the BTE is connected to a separate external coil assembly via a cable, where the external coil assembly is configured to be magnetically and inductively coupled to the implantable coil 114, while in other embodiments the BTE includes a coil disposed in or on the housing worn on the outer ear of the user. It is also to be appreciated that alternative external components could be located in the user’s ear canal, worn on the body, etc.

[0030] Although the cochlear implant system 102 includes the sound processing unit 106 and the cochlear implant 112, as described below, the cochlear implant 112 can operate independently from the sound processing unit 106, for at least a period, to stimulate the user. For example, the cochlear implant 112 can operate in a first general mode, sometimes referred to as an “external hearing mode,” in which the sound processing unit 106 captures sound signals which are then used as the basis for delivering stimulation signals to the user. The cochlear implant 112 can also operate in a second general mode, sometimes referred as an “invisible hearing” mode, in which the sound processing unit 106 is unable to provide sound signals to the cochlear implant 112 (e.g., the sound processing unit 106 is not present, the sound processing unit 106 is powered-off, the sound processing unit 106 is malfunctioning, etc.). As such, in the invisible hearing mode, the cochlear implant 112 captures sound signals itself via implantable sound sensors and then uses those sound signals as the basis for delivering stimulation signals to the user. Further details regarding operation of the cochlear implant 112 in the external hearing mode are provided below, followed by details regarding operation of the cochlear implant 112 in the invisible hearing mode. It is to be appreciated that reference to the external hearing mode and the invisible hearing mode is merely illustrative and that the cochlear implant 112 could also operate in alternative modes.

[0031] In FIGs. 1A and 1C, the cochlear implant system 102 is shown with an external device 110, configured to implement aspects of the techniques presented. The external device 110, which is shown in greater detail in FIG, 1 E, is a computing device, such as a personal computer (e.g., laptop, desktop, tablet), a mobile phone (e.g., smartphone), a remote control unit, etc. The external device 110 and the cochlear implant system 102 (e.g., sound processing unit 106 or the cochlear implant 112) wirelessly communicate via a bi-directional communication link 126.The bi-directional communication link 126 may comprise, for example, a short-range communication, such as Bluetooth link, Bluetooth Low Energy (BLE) link, a proprietary link, etc.

[0032] Returning to the example ofFIGs. 1A-1D, the sound processing unit 106 of the external component 104 also comprises one or more input devices configured to capture and / or receive input signals (e.g., sound or data signals) at the sound processing unit 106. The one or more input devices include, for example, one or more sound input devices 118 (e.g., one or more external microphones, audio input ports, telecoils, etc.), one or more auxiliary input devices 119 (e.g., audio ports, such as a Direct Audio Input (DAI), data ports, such as a Universal Serial Bus (USB) port, cable port, etc.), and a short-range wireless transmitter / receiver (wireless transceiver) 120 (e.g., for communication with the external device 110), each located in, on or near the sound processing unit 106. However, it is to be appreciated that one or more input devices may include additional types of input devices and / or less input devices (e.g., the short- range wireless transceiver 120 and / or one or more auxiliary input devices 119 could be omitted).

[0033] The sound processing unit 106 also comprises a charging coil 121, a closely-coupled radio frequency transmitter / receiver (RF transceiver) 122, at least one rechargeable battery 123, and an external processing module 124. The external processing module 124 can be configured to perform a number of operations, including sound processing operations that are represented in FIG. ID by an automatic speech recognition (ASR)-based sound processor 210, which will be described in further detail below with reference to FIG. 2. In certain embodiments, the ASR- based sound processor 210 is configured to improve speech perception by augmenting the encoding of phonemic information, among various other features. The external processing module 124 and / or the ASR-based sound processor 210 can be formed by one or more processors (e.g., one or more Digital Signal Processors (DSPs), one or more uC cores, etc.), firmware, software, etc. arranged to perform operations described herein. The ASR-based sound processor 210 can be implemented as firmware elements, partially or fully implemented with digital logic gates in one or more application-specific integrated circuits (ASICs), partially or fully in software, etc. Although FIG. ID illustrates the ASR-based sound processor 210 as being implemented / performed at the external processing module 124, it is to be appreciated that these elements (e.g., functional operations) could also or alternatively be implemented / performed as part of the implantable processing module 158, as part of the external device 110, etc. The option for implantation of the based sound processor 210 as beingimplemented / performed at the external processing module 124, the implantable processing module 158, and / orthe external device 110 is signified by the dashed lines for 210 in FIGs. ID and IE.

[0034] Returning to the example of FIGs. 1A-1D, the implantable component 112 comprises an implant body (main module) 134, a lead region 136, and the stimulating assembly 116, all configured to be implanted under the skin (tissue) 115 of the user. The implant body 134 generally comprises a hermetically-sealed housing 138 in which the RF interface circuitry 140, at least one power source 141 (e.g., one or more batteries, one or more capacitors, etc.), and a stimulator unit 142 are disposed. The implant body 134 also includes the intemal / implantable coil 114 that is generally external to the housing 138, but which is connected to the RF interface circuitry 140 via a hermetic feedthrough (not shown in FIG. ID).

[0035] As noted, the stimulating assembly 116 is configured to be at least partially implanted in the user’s cochlea. The stimulating assembly 116 includes a plurality of longitudinally spaced intra-cochlear electrical stimulating contacts (electrodes) 144 that collectively form a contact array (electrode array) 146 for delivery of electrical stimulation (current) to the recipient’s cochlea. The stimulating assembly 116 extends through an opening in the recipient’s cochlea (e.g., cochleostomy, the round window, etc.) and has a proximal end connected to stimulator unit 142 via lead region 136 and a hermetic feedthrough (not shown in FIG. ID). Lead region 136 includes a plurality of conductors (wires) that electrically couple the electrodes 144 to the stimulator unit 142. The implantable component 112 also includes an electrode outside of the cochlea, sometimes referred to as the extra-cochlear electrode (ECE) 139.

[0036] As noted, the cochlear implant system 102 includes the external coil 108 and the implantable coil 114. The external magnet 150 is fixed relative to the external coil 108 and the intemal / implantable magnet 152 is fixed relative to the implantable coil 114. The external magnet 150 and the intemal / implantable magnet 152 fixed relative to the external coil 108 and the intemal / implantable coil 114, respectively, facilitate the operational alignment of the external coil 108 with the implantable coil 114. This operational alignment of the coils enables the external component 104 to transmit data and power to the implantable component 112 via a closely-coupled wireless link 148 formed between the external coil 108 with the implantable coil 114. In certain examples, the closely-coupled wireless link 148 is an RF link. However, various other types of energy transfer, such as infrared (IR), electromagnetic, capacitive and inductive transfer, may be used to transfer the power and / or data from an external component to an implantable component and, as such, FIG. ID illustrates only one example arrangement.

[0037] As noted above, the sound processing unit 106 includes the external processing module 124. The external processing module 124 is configured to process the received sound signals (received at one or more of the input devices, such as sound input devices 118 and / or auxiliary input devices 119) and convert the received input sound signals into output control signals for use in stimulating a first ear of a recipient or user (i.e., the external processing module 124 is configured to perform sound processing on input signals received at the sound processing unit 106). Stated differently, the one or more processors (e.g., processing element(s) implementing firmware, software, etc.) in the external processing module 124 are configured to execute sound processing logic in memory to convert the received input sound signals into output control signals (stimulation signals) that represent electrical stimulation for delivery to the recipient.

[0038] As noted, FIG. ID illustrates an embodiment in which the external processing module 124 in the sound processing unit 106 generates the output control signals. In an alternative embodiment, the sound processing unit 106 can send less processed information (e.g., audio data) to the implantable component 112, and the sound processing operations (e.g., conversion of input sounds to output control signals 156) can be performed by a processor within the implantable component 112.

[0039] In FIG. ID, according to an example embodiment, output control signals (stimulation signals) are provided to the RF transceiver 122, which transcutaneously transfers the output control signals (e.g., in an encoded manner) to the implantable component 112 via the external coil 108 and the implantable coil 114. That is, the output control signals (stimulation signals) are received at the RF interface circuitry 140 via the implantable coil 114 and provided to the stimulator unit 142. The stimulator unit 142 is configured to utilize the output control signals to generate electrical stimulation signals (e.g., current signals) for delivery to the user’s cochlea via one or more of the stimulating contacts 144. In this way, cochlear implant system 102 electrically stimulates the user’s auditory nerve cells, bypassing absent or defective hair cells that normally transduce acoustic vibrations into neural activity, in a manner that causes the recipient to perceive one or more components of the input sound signals (the received sound signals).

[0040] As detailed above, in the external hearing mode, the cochlear implant 112 receives processed sound signals from the sound processing unit 106. However, in the invisible hearing mode, the cochlear implant 112 is configured to capture and process sound signals for use in electrically stimulating the user’s auditory nerve cells. In particular, as shown in FIG. ID, an example embodiment of the cochlear implant 112 can include a plurality of implantable soundsensors 165(1), 165(2) that collectively form a sensor array 160, and an implantable processing module 158. In certain embodiments, the implantable processing module 158 can comprise an ASR-based sound processor 210, which is configured to improve speech perception by augmenting the encoding of phonemic information, among various other features, as described in further detail below with reference to FIG. 2. Similar to the external processing module 124 described above, the implantable processing module 158 and the ASR-based sound processor 210 can comprise, for example, one or more processors and a memory device (memory) that includes sound processing logic. The memory device may comprise any one or more of: Non- Volatile Memory (NVM), Ferroelectric Random Access Memory (FRAM), read only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. The one or more processors are, for example, microprocessors or microcontrollers that execute instructions for the sound processing logic stored in memory device.

[0041] In the invisible hearing mode, the implantable sound sensors 165(1), 165(2) of the sensor array 160 are configured to detect / capture sound signals 166 (e.g., acoustic sound signals, vibrations, etc.), which are provided to the implantable processing module 158. The implantable processing module 158 is configured to convert received sound signals 166 (received at one or more of the implantable sound sensors 165(1), 165(2)) into output control signals 156 for use in stimulating the first ear of a recipient or user (i.e., the implantable processing module 158 is configured to perform sound processing operations). Stated differently, the one or more processors (e.g., processing element(s) implementing firmware, software, etc.) in the implantable processing module 158 are configured to execute sound processing logic in memory to convert the received sound signals 166 into output control signals 156 that are provided to the stimulator unit 142. The stimulator unit 142 is configured to utilize the output control signals 156 to generate electrical stimulation signals (e.g., current signals) for delivery to the user’s cochlea, thereby bypassing the absent or defective hair cells that normally transduce acoustic vibrations into neural activity.

[0042] It is to be appreciated that the above description of the so-called external hearing mode and the so-called invisible hearing mode are merely illustrative and that the cochlear implant system 102 could operate differently in different embodiments. For example, in one alternative implementation of the external hearing mode, the cochlear implant 112 could use signals captured by the sound input devices 118 and the implantable sound sensors 165(1), 165(2) ofsensor array 160 in generating stimulation signals for delivery to the user. For hearing devices that include an implantable processing module, such as implantable processing module 158, the techniques presented herein may be implemented without an external processor. Accordingly, a hearing device that includes an implant body 134 and lacks an external component 104 may be configured to implement the techniques presented herein.

[0043] FIG. IE is a block diagram illustrating one example arrangement for an external computing device 110 configured to perform one or more operations in accordance with certain embodiments presented herein. As shown in FIG. IE, in its most basic configuration, the external computing device 110 includes at least one processing unit 183 and a memory 184. The processing unit 183 includes one or more hardware or software processors (e.g., Central Processing Units) that can obtain and execute instructions. The processing unit 183 can communicate with and control the performance of other components of the external computing device 110. The memory 184 is one or more software or hardware-based computer-readable storage media operable to store information accessible by the processing unit 183. The memory 184 can store, among other things, instructions executable by the processing unit 183 to implement applications or cause performance of operations described herein, as well as other data. The memory 184 can be volatile memory (e.g., RAM), non-volatile memory (e.g., ROM), or combinations thereof. The memory 184 can include transitory memory or non-transitory memory. The memory 184 can also include one or more removable or non-removable storage devices. In examples, the memory 184 can include RAM, ROM, EEPROM (Electronically- Erasable Programmable Read-Only Memory), flash memory, optical disc storage, magnetic storage, solid state storage, or any other memory media usable to store information for later access. By way of example, and not limitation, the memory 184 can include wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, RF, infrared, other wireless media, or combinations thereof. The memory 184 comprises logic 185 that, when executed, enables the processing unit 183 to perform aspects of the techniques presented.

[0044] In the illustrated example of FIG. IE, the external computing device 110 further includes a network adapter 186, one or more input devices 187, and one or more output devices 188. The external computing device 110 can include other components, such as a system bus, component interfaces, a graphics system, a power source (e.g., a battery), among other components. The network adapter 186 is a component of the external computing device 110 that provides network access (e.g., access to at least one network 189). The network adapter186 can provide wired or wireless network access and can support one or more of a variety of communication technologies and protocols, such as Ethernet, cellular, Bluetooth, near-field communication, and RF, among others. The network adapter 186 can include one or more antennas and associated components configured for wireless communication according to one or more wireless communication technologies and protocols. The one or more input devices187 are devices over which the external computing device 110 receives input from a user. The one or more input devices 187 can include physically-actuatable user-interface elements (e.g., buttons, switches, or dials), a keypad, keyboard, mouse, touchscreen, and voice input devices, among other input devices that can accept user input. The one or more output devices 188 are devices by which the external computing device 110 is able to provide output to a user. The output devices 188 can include a display 190 (e.g., a liquid crystal display (LCD)) and one or more speakers 191, among other output devices for presentation of visual or audible information to the recipient, a clinician, an audiologist, or other user.

[0045] In certain embodiments, the memory 184 also comprises the ASR-based sound processor 210. In some examples, the external computing device 110 can be used to program or otherwise configure the external processing module 124 of the sound processing unit 106 of FIG. ID with the ASR-based sound processor 210 for implementing the techniques presented herein. In some other examples, the external computing device 110 can be used to program or otherwise configure the implantable processing module 158 of the cochlear implant 112 of FIG. ID with the ASR-based sound processor 210 for implementing the techniques presented herein. Likewise, the external computing device 110 can be used to provide updates to the sound processing unit 106 and / or the cochlear implant 112 with respect to the ASR-based sound processor 210 (e.g., retraining machine learning models or neural networks, optimizing parameters or operational settings, enabling or disabling features, adjusting calculations, changing mixing factors or weights, etc.). In addition, the external computing device 110 can be used to store data (e.g., mapping tables or data structures, speaker specific databases, user customized parameters or operational settings, etc.) on behalf of the sound processing unit 106 and / or the cochlear implant 112 with respect to the ASR-based sound processor 210.

[0046] It is to be appreciated that the arrangement for the external computing device 110 shown in FIG. IE is merely illustrative and that aspects of the techniques presented herein can be implemented at a number of different types of systems / devices including any combination of hardware, software, and / or firmware configured to perform the functions described herein. For example, the external computing device 110 can be a personal computer (e.g., a desktop orlaptop computer), a hand-held device (e.g., a tablet computer), a mobile device (e.g., a smartphone), a surgical system, and / or any other electronic device having the capabilities to perform the associated operations described elsewhere herein.

[0047] As noted above, certain hearing devices, such as cochlear implants, operate by electrically stimulating a recipient based on received acoustic sounds (sound signals) to bypass damaged sensory receptors and eliciting neural activation patterns that represent acoustic sounds. While cochlear implants restore a sense of hearing to people with severe to profound deafness, some cochlear implant recipients still struggle with complicated listening situations such as speech perception in noise. One source of these difficulties is the limited sound information that is provided by a cochlear implant.

[0048] For example, conventional cochlear implants generally extract the signal envelope in each of a number of frequency bands that correspond to each of the implanted electrodes, and those envelopes are used to modulate fixed-rate biphasic pulse trains that are provided to the electrodes. The use of only temporal envelopes in a limited number of frequency bands reduces both temporal and spectral resolution of the acoustic sounds. Furthermore, the electrical current delivered to the electrodes spreads through the conductive fluid of the cochlea, limiting channel independence. Consequently, the neural activation patterns evoked by cochlear implants are only a coarse approximation of those evoked by acoustic hearing. The result is that vowels and consonants (“phonemes”) that make up speech are unclear, and the cochlear implant recipient has difficulty understanding speech in noise.

[0049] As used herein, the term “phoneme” refers to perceptually distinct units of speech in a specified language that distinguish one word from another (e.g., the smallest unit of speech for differentiating one word or word element from another). In general, phonemes are any of the abstract units of the phonetic system of a language that correspond to a set of similar speech sounds. Stated differently, a phoneme is a sound or group of different sounds perceived to have the same function by speakers of the language or dialect in question. In English, for example, there are 44 phonemes that make up the language, and these word sounds are divided into 19 consonants, 7 digraphs, 5 sounds that are “r-controlled” or “r-influenced” sounds, 5 long vowels, 5 short vowels, 2 “oo” vowel sounds, and 2 diphthongs. Other languages have different numbers of phonemes and various different types of phonemes, such that these numbers are merely intended to be illustrative and may vary widely across different languages and dialects. Generally, phonemes are critical to speaker pronunciation and speech understanding by recipients.

[0050] Presented herein are techniques for enhancing the ability of a cochlear implant, or other auditory prosthesis recipient, to correctly perceive phonetic information, including phonemes related aspects of human speech, through the use of “automatic speech recognition (ASR)- based sound processing.” As used herein, ASR-based sound processing is a processing technique proposed by the inventors in which automatic speech recognition (ASR) is integrated into the sound processing operations of a cochlear implant, auditory prosthesis, or other hearing device in order to identify phonetic information for subsequent emphasis (enhancement) thereof. A cochlear implant or other hearing device implementing the techniques presented herein is sometimes referred to as including a “ASR-based sound processor” (e.g., a sound processing module comprised of one or more processors, memory, firmware, etc., implementing ASR-based sound processing).

[0051] Automatic speech recognition (ASR) algorithms are artificially intelligent (e.g., machine-learned) algorithms that can recognize unique phonetic information (e.g., phonemes and / or related aspects of human speech) from an audio input. By way of example and not by limitation, ASR algorithms can be sequence-to-sequence neural networks configured to convert speech to text.

[0052] In accordance with certain embodiments presented herein, an ASR-based sound processor comprises an automatic speech recognition module that is configured to convert parts of human speech, identified from an audio input, to a “predicted phoneme stimulation pattern,” as described in more detail below. As one non-limiting illustrative example, if the automatic speech recognition module recognizes the phoneme “ah,” then the output will be a stimulation sequence template (predicted phoneme stimulation pattern) corresponding to the sound “ah.” The predicted phoneme stimulation pattern can then be used to modify a stimulation pattern that is a direct translation of the acoustic environment into to electrical pulses, sometimes referred to herein as “direct stimulation pattern.” In this manner, the ASR-based sound processor generates and outputs a so-called “augmented” stimulation sequence that is used to deliver stimulation to the recipient.

[0053] According to another aspect of the techniques presented herein, a deep neural network (DNN) based sound coding module can be combined with an automatic speech recognition module. In this example embodiment, the ASR-based sound processor comprises an automatic speech recognition module that is configured to identify phonemes in an audio input and output corresponding “predicted phoneme codes,” as described in more detail below. If the DNN- based sound coding module is made aware that a particular phoneme is meant to be conveyedto the recipient, then the DNN-based sound coding module can implement one or more adjustments to improve the representation of that particular phoneme in the stimulation sequence template that is output by the DNN-based sound coding module. In this manner, the ASR-based sound processor generates and outputs an augmented stimulation pattern that is used to deliver stimulation to the recipient.

[0054] In certain example embodiments, as described in more detail below, an automatic speech recognition module can be further combined with a “speaker characteristics” module that is configured to identify various speaker-specific characteristics, including but not limited to duration, pitch, timbre, spectrum, modulations, etc.

[0055] Initially, an automatic speech recognition (ASR)-based sound processor (ASR-based sound processor) 210 in accordance with the techniques presented will be described at a high level with reference to FIG. 2A. Thereafter, several different variations / embodiments of the ASR-based sound processor 210 are described in greater detail with reference to FIGs. 3, 4, 5, 6A-6B, 7, 8, and 9, respectively, according to various example embodiments of the techniques presented herein.

[0056] FIG. 2A is a block diagram of an ASR-based sound processor 210 with which aspects of the techniques presented herein can be implemented. The ASR-based sound processor 210 can be implemented as part of a cochlear implant system, such as the cochlear implant system 102 described above with reference to FIGs. 1A-1D. In some examples, the ASR-based sound processor 210 can be integrated in an external sound processing component, such as the external processing module 124 of the sound processing unit 106 of FIG. ID. In some other examples, the ASR-based sound processor 210 can be integrated in an intemal / implantable sound processing component, such as the implantable processing module 158 of the cochlear implant 112 of FIG. ID.

[0057] In the example of FIG. 2A, the ASR-based sound processor 210 receives an audio / sound signal 205 from an input 202, processes the sound signal 205 to generate an augmented stimulation pattern 215, and provides the augmented stimulation pattern 215 to an output 219. The augmented stimulation pattern 215 can then be used to deliver electrical stimulation to the recipient in the manner described above with respect to FIG. ID.

[0058] As shown in FIG. 2A, the ASR-based sound processor 210 comprises at least a sound coding module 220 and an automatic speech recognition module 230. As described below with reference to FIG. 2B, the sound coding module 220 is configured to generate a stimulationpatern directly from the full sound signal 205 (e.g., a direct stimulation patern), and the automatic speech recognition module 230 is configured to at least identify phonemes from speech present in the sound signal 205. The ASR-based sound processor 210 is configured to use the phonemes identified by the automatic speech recognition module 230 to modify the stimulation patern generated by the sound coding module 220, and thereby generate the augmented stimulation patern 215.

[0059] As noted, FIGs. 3, 4, 5, 6A-6B, 7, 8, and 9 illustrate various embodiments of an ASR- based sound processor, several of which include a sound coding module. FIG. 2B illustrates one example implementation of sound coding module 210, sometimes referred to as a sound processing path, of a cochlear implant in accordance with certain embodiments presented herein.

[0060] In the example of FIG. 2B, the sound coding module 220 can receive audio inputs via three input elements 201, which comprise two microphones 209 and at least one auxiliary input 211 (e.g., an audio input port, a cable port, a telecoil, a wireless transceiver, etc.). If not already in an electrical form, sound input elements 208 convert received / sound signals into electrical signals 253, referred to herein as electrical input signals, that represent the received sound signals. As shown in FIG. 2B, the electrical input signals 253 are provided to a pre-filterbank processing module 254.

[0061] The pre-filterbank processing module 254 is configured to, as needed, combine the electrical input signals 253 received from the sound input elements 201 and prepare those signals for subsequent processing. The pre-filterbank processing module 254 then generates a pre-filtered output signal 255 that, as described further below, is the basis of further processing operations. The pre-filtered output signal 255 represents the collective sound signals received at the sound input elements 201 at a given point in time.

[0062] The sound coding module 210 is generally configured to execute sound processing and coding to convert the pre-filtered output signal 255 into output signals that represent electrical stimulation for delivery to the recipient. As such, the coding module 210 comprises a filterbank module (filterbank) 256, a post-filterbank processing module 258, a channel selection module 260, and a channel mapping and encoding module 262.

[0063] In operation, the pre-filtered output signal 255 generated by the pre-filterbank processing module 254 is provided to the filterbank module 256. The filterbank module 256 generates a suitable set of bandwidth limited channels, or frequency bins, that each includes aspectral component of the received sound signals. That is, the filterbank module 256 comprises a plurality of band-pass fdters that separate the pre-filtered output signal 255 into multiple components / channels, each one carrying a single frequency sub-band ofthe original signal (i.e., frequency components of the received sounds signal).

[0064] The channels created by the filterbank module 256 are sometimes referred to herein as sound processing channels, and the sound signal components within each of the sound processing channels are sometimes referred to herein as band-pass fdtered signals or channelized signals. The band-pass fdtered or channelized signals created by the filterbank module 256 are processed (e.g., modified / adjusted) as they pass through the sound processing path 250. As such, the band-pass fdtered or channelized signals are referred to differently at different stages of the sound processing path 250. However, it will be appreciated that reference herein to a band-pass fdtered signal or a channelized signal may refer to the spectral component of the received sound signals at any point within the processing path (e.g., pre-processed, processed, selected, etc.).

[0065] At the output of the fdterbank module 256, the channelized signals are initially referred to herein as pre-processed signals 257. The number ‘m’ of channels and pre-processed signals 257 generated by the fdterbank module 256 may depend on a number of different factors including, but not limited to, implant design, number of active electrodes, coding strategy, and / or recipient preference (s). In certain arrangements, twenty-two (22) channelized signals are created and the sound processing path is said to include 22 channels.

[0066] The pre-processed signals 257 are provided to the post-fdterbank processing module 258. The post-fdterbank processing module 258 is configured to perform a number of sound processing operations on the pre-processed signals 257. These sound processing operations include, for example, channelized gain adjustments for hearing loss compensation (e.g., gain adjustments to one or more discrete frequency ranges of the sound signals), noise reduction operations, speech enhancement operations, etc., in one or more of the channels. After performing the sound processing operations, the post-fdterbank processing module 258 outputs a plurality of processed channelized signals 259.

[0067] In the specific arrangement of FIG. 2A, the processing path of the sound coding module 210 includes a channel selection module 260. The channel selection module 260 is configured to perform a channel selection process to select, according to one or more selection rules, which of the ‘m’ channels should be use in hearing compensation. The signals selected at channelselection module 260 are represented in FIG. 2B by arrow 261 and are referred to herein as selected channelized signals or, more simply, selected signals.

[0068] In the embodiment of FIG. 2B, the channel selection module 256 selects a subset ‘n’ of the ‘m’ processed channelized signals 259 for use in generation of electrical stimulation for delivery to a recipient (i.e., the sound processing channels are reduced from ‘m’ channels to ‘n’ channels). In one specific example, the ‘n’ largest amplitude channels (maxima) from the ‘m’ available combined channel signals / masker signals is made, with ‘m’ and ‘n’ being programmable during initial fitting, and / or operation of the prosthesis. It is to be appreciated that different channel selection methods could be used, and are not limited to maxima selection.

[0069] It is also to be appreciated that, in certain embodiments, the channel selection module 260 may be omitted. For example, certain arrangements may use a continuous interleaved sampling (CIS), CIS-based, or other non-channel selection sound coding strategy.

[0070] The processing path of sound coding module 210 also comprises the channel mapping module 262. The channel mapping module 262 is configured to map the amplitudes of the selected signals 261 (or the processed channelized signals 259 in embodiments that do not include channel selection) into a set of output signals 263 (e.g., direct stimulation pattern) that represent the attributes of the electrical stimulation signals that are to be delivered to the recipient so as to evoke perception of at least a portion of the received sound signals. This channel mapping may include, for example, threshold and comfort level mapping, dynamic range adjustments (e.g., compression), volume adjustments, etc., and may encompass selection of various sequential and / or simultaneous stimulation strategies.

[0071] In summary, the sound coding module 210 operates to extract the signal envelope in each of a number of frequency bands that correspond to each of the implanted electrodes, and those envelopes are used to generate a so-called “direct stimulation pattern” (e.g., modulation of fixed-rate biphasic pulse trains that are provided to the electrodes) for a broad frequency range of the input sound signals (e.g., sound signal received via the inputs 201). The output of the sound coding module 210 is referred to herein as a “direct stimulation pattern” because, as detailed above, the output is based on a broad frequency range of the input sound signals.

[0072] Generally, the exemplary automatic speech recognition (ASR) based sound coding modules described below with reference to FIGs. 3-9 can be implemented as part of a cochlear implant system (e.g., the cochlear implant system 102 of FIGs. 1A-1D). For example, the ASR- based sound processor of FIGs. 3-9 can be integrated in an external sound processingcomponent (e.g., the external processing module 124 of the sound processing unit 106 of FIG. ID), or the ASR-based sound processor s can be integrated in an intemal / implantable sound processing component (e.g., the implantable processing module 158 of the cochlear implant 112 of FIG. ID).Example Embodiments

[0073] FIG. 3 is a block diagram of one implementation of an ASR-based sound processor 310, according to an example embodiment. In the ASR-based sound processor 310 of FIG. 3, an automatic speech recognition module and a speaker characteristics module are used to modify a stimulation pattern output from a sound coding module.

[0074] More specifically, the ASR-based sound processor 310 comprises a sound coding module 320, an automatic speech recognition module 330, and a speaker characteristics module 340. Each of the sound coding module 320, the automatic speech recognition module 330, and the speaker characteristics module 340 receives an input sound signal 305 (audio) and performs respective processing operations.

[0075] The sound coding module 320, which could operate similar to sound coding module 220, is configured to convert at least one input sound signal 305 into a stimulation pattern 325 (i.e., a stimulation pattern for the full acoustic input). The automatic speech recognition module 330 is configured to process the at least one input sound signal 305 to identify one or more phonemes of speech present in the at least one input sound signal 305, and to generate a predicted phoneme stimulation pattern 335 based on the one or more identified phonemes. The speaker characteristics module 340 is configured to process the input sound signal 305 to identify and output one or more speaker characteristics 345 (e.g., duration, pitch, timbre, spectrum, modulations, etc.).

[0076] The ASR-based sound processor 310 is also configured to combine the predicted phoneme stimulation pattern 335 output by the automatic speech recognition module 330 with the speaker characteristics 345 output by the speaker characteristics module 340 (at operation 350) to generate an adjusted phoneme stimulation pattern 355. That is, the speaker characteristics 345 can be used to adjust the predicted phoneme stimulation pattern 335. The ASR-based sound processor 310 is configured to use the adjusted phoneme stimulation pattern 355 to modify the stimulation pattern 325 output by the sound coding module 320 (at operation 360) to generate an augmented stimulation pattern 365. The ASR-based sound processor 310then outputs the augmented stimulation pattern 365 for use in delivering stimulation to the recipient.

[0077] Thus, in the example of FIG. 3, the output of the automatic speech recognition module 330 is matched to a phoneme template stimulation pattern, and can be mixed with a sound coding output, as well as various speaker-specific characteristics. As noted, the automatic speech recognition module 330 converts the input sound signal 305 into a predicted phoneme stimulation pattern 335.

[0078] In the non-limiting illustrative example mentioned above, the automatic speech recognition module 330 can output a template of a stimulation sequence for the vowel sound “ah” upon that vowel sound being recognized. Preferably, in certain embodiments, this vowel stimulation pattern can be modified to reflect the specific vocal characteristics of the speaker. This would be advantageous since vowels can sound different between different individual speakers, respectively. For example, the speaker characteristics module 340 can be used to extract duration, pitch, timbre, spectrum information, modulations, etc. from the input sound signal 305 to modulate the vowel stimulation pattern. These different outputs would then be combined to produce the augmented stimulation pattern 365 (e.g., by modifying temporal envelopes, increasing a channel for a certain sound and / or decrease other channels, etc.), which is the final stimulation sequence that would be applied to the implanted electrodes of the cochlear implant system. The speaker characteristics module 340 is also useful because speech is not just content and semantic information, but it also conveys emotion (through prosody), background (through accents), health, age, fatigue, etc.

[0079] FIG. 4 is a block diagram of an ASR-based sound processor 410, according to an example embodiment. In the ASR-based sound processor 410 of FIG. 4, a smart mixing algorithm is used to modify a stimulation pattern output from a sound coding module based on outputs from an automatic speech recognition module and a speaker characteristics module.

[0080] More specifically, the ASR-based sound processor 410 comprises a sound coding module 420, an automatic speech recognition module 430, a speaker characteristics module 440, and a smart mixing algorithm 470. Each of the sound coding module 420, the automatic speech recognition module 430, and the speaker characteristics module 440 receives at least one input sound signal 405 (audio) and performs respective processing operations.

[0081] The sound coding module 420, which could operate similar to sound coding module 220, is configured to process the input sound signal 405 to generate a stimulation pattern 425(i.e., a stimulation pattern for the full acoustic input). The automatic speech recognition module 430 is configured to process the input sound signal 405 to identify one or more phonemes of speech present in the at least one input sound signal 405, and to generate a predicted phoneme stimulation pattern 435 based on the one or more identified phonemes. The speaker characteristics module 440 is configured to process the at least one input sound signal 405 to identify and output one or more speaker characteristics 445 (e.g., duration, pitch, timbre, spectrum, modulations, etc.).

[0082] The smart mixing algorithm 470 receives the stimulation pattern 425 from the sound coding module 420, the predicted phoneme stimulation pattern 435 from the automatic speech recognition module 430, and the speaker characteristics 445 from the speaker characteristics module 440, respectively. The smart mixing algorithm 470 is configured to generate an augmented stimulation pattem / sequence 475 based on the stimulation pattern 425, the predicted phoneme stimulation pattern 435, and the speaker characteristics 445. In certain embodiments, the smart mixing algorithm 470 can use the speaker characteristics 445 to adjust the predicted phoneme stimulation pattern 435, and use the adjusted phoneme stimulation pattern to modify the stimulation pattern 425 to generate the augmented stimulation pattern 475. The ASR-based sound processor 410 then outputs the augmented stimulation pattern 475 for use in delivering stimulation to the recipient.

[0083] The sound coding module 420 may perform acceptable on its own when the recipient (i.e., implant recipient) is in, for example, a quiet environment (e.g., in silence, etc.). However, performance of the sound coding module 420 can be reduced in a noisy environment (e.g., speech in noise, etc.), and supplementing the sound coding module 420 with one or both of the automatic speech recognition module 430 and / or the speaker characteristics module 440 provides improved performance, and is better for speech in noise situations in particular. For example, modifications to the stimulation pattern 425 output from the sound coding module 420 can include, but are not limited to, increasing envelopes in the area mapped to that phoneme and / or decreasing envelopes in other areas (e.g., increase channel for that sound and / or decrease other channels), so as to enhance the perceptibility of the phoneme by the recipient. Integrating the speaker-specific characteristics into the processing path also ensures that every person’s voice does not sound the same to the recipient.

[0084] The smart mixing algorithm 470 can mix the stimulation pattern 425, the predicted phoneme stimulation pattern 435, and the speaker characteristics 445 in many different ways to generate the augmented stimulation pattern 475. In one example, an averaging techniquecould be used, whereby averages of the pulse amplitudes of the stimulation pattern 425 and the predicted phoneme stimulation pattern 435, respectively, can be calculated and the average pulse amplitudes can then be used for the augmented stimulation pattern 475. In another example, a weighting technique could be used, whereby the stimulation pattern 425 and the predicted phoneme stimulation pattern 435 can be assigned mixing factors (weights), such that a mixing factor of 0.5 indicates that each of the stimulation pattern 425 and the predicted phoneme stimulation pattern 435 have an equal weight (50% / 50%), a mixing factor of 0.6, 0.7, 0.8, or 0.9 indicates a relatively greater weight for the stimulation pattern 425 or the predicted phoneme stimulation pattern 435, a mixing factor of 0.4, 0.3, 0.2, or 0. 1 indicates a relatively lesser weight for the stimulation pattern 425 or the predicted phoneme stimulation pattern 435, and so forth. Preferably, these mixing factors or weights can be user adjustable to allow for a more customized listening experience depending on individual preferences.

[0085] As noted, the automatic speech recognition module 430 and the speaker characteristics module 440 can be beneficial for use in noisy environments, and particularly for a speech in noise situation. In certain embodiments, the automatic speech recognition module 430 and the speaker characteristics module 440 could be activated or enabled only for use in situations where the recipient (implant recipient) is struggling to understand speech or individual speakers. For example, an “environmental classifier” could be integrated in the sound coding processing path and used as a trigger condition to activate or enable the automatic speech recognition module 430 and the speaker characteristics module 440, such as when the environmental classifier detects a “speech-in-noise” sound class in the input sound signal 405, for example. It should be appreciated, however, that the techniques presented herein could also be used in a quiet environment if desired, which could better convey certain speaker-specific characteristics (e.g., pitch, accents, etc.), for example.

[0086] Thus, in the example of FIG. 4, the output of the automatic speech recognition module 430 is matched to a phoneme template stimulation pattern, and can be mixed with a sound coding output, as well as various speaker-specific characteristics. In this example, the smart mixing algorithm 470 can be a machine learning algorithm or a deep neural network (DNN) that is trained to produce the augmented stimulation pattern 595 based on a mixture of the stimulation pattern 425 for the full acoustic input, the predicted phoneme stimulation pattern 435, and the speaker characteristics 445.

[0087] FIG. 5 is a block diagram of an ASR-based sound processor 510, according to an example embodiment. In the ASR-based sound processor 510 of FIG. 5, a deep neural network(DNN)-based sound coding module is configured to generate an augmented stimulation sequence based on at least one audio input and an output (e.g., a predicted phoneme) from an automatic speech recognition module (and optionally, and output from a speaker characteristics module).

[0088] More specifically, the ASR-based sound processor 510 comprises at least an automatic speech recognition module 530 and a DNN-based sound coding module 590. Each of the automatic speech recognition module 530 and the DNN-based sound coding module 590 receives at least one input sound signal 505 (audio) and performs respective processing operations.

[0089] The automatic speech recognition module 530 is configured to process the at least one input sound signal 505 to identify a phoneme of speech present in the input sound signal 405. However, rather than generating a predicted phoneme stimulation pattern (like in the examples of FIGs. 3 and 4 above), the automatic speech recognition module 530 of FIG. 5 is configured to output a predicted phoneme code 535 that corresponds to the identified phoneme. For example, assuming that there are 44 phonemes in the English language, as described above, the predicted phoneme code 535 can be a number from 1-44 that corresponds to a respective phoneme (e.g., as stored in a mapping table or other data structure), although example embodiments are not limited thereto and other suitable techniques can also be used to implement the phoneme code feature.

[0090] In addition to receiving the input sound signal 505, the DNN-based sound coding module 590 also receives the predicted phoneme code 535 (e.g., a number from 1-44) from the automatic speech recognition module 530. The DNN-based sound coding module 590 is configured to generate an augmented stimulation pattern 595 based on at least the input sound signal 505 and the predicted phoneme code 535. The ASR-based sound processor 510 then outputs the augmented stimulation pattern 595 for use in delivering stimulation to the recipient.

[0091] In another variation of the specific example of FIG. 5, the ASR-based sound processor 510 can also comprise a speaker characteristics module 540, although the speaker characteristics module 540 can be considered optional in the example of FIG. 5 (as indicated by dashed lines in FIG. 5) and is not necessarily required, and thus the speaker characteristics module 540 can be omitted in some other example embodiments. In this variation, the speaker characteristics module 540 also receives the input sound signal 505 and performs respective processing operations. The speaker characteristics module 540 is configured to process theinput sound signal 505 to identify and output one or more speaker characteristics 545 (e.g., duration, pitch, timbre, spectrum, modulations, etc.). In addition to the predicted phoneme code 535, the DNN-based sound coding module 590 also receives the speaker characteristics 545 from the speaker characteristics module 540. In this variation, the DNN-based sound coding module 590 is configured to generate the augmented stimulation pattern 595 based on the input sound signal 505, the predicted phoneme code 535, and the speaker characteristics 545.

[0092] Thus, in the example of FIG. 5, automatic speech recognition techniques are combined with a DNN-based sound coding path, where the “phoneme code” output by the automatic speech recognition module 530 is used as an additional input to the DNN-based sound coding module 590 (which is audio to stimulation direct in this example). The DNN-based sound coding module 590 is trained to convert an input audio sequence to a cochlear implant electrode stimulation pattern directly, using features from the raw audio input and from a parallel ASR- based algorithm. In particular, the DNN-based sound coding module 590 is trained to enhance phonemes and to output corresponding channel envelopes. If the DNN-based sound coding module 590 is made aware of the specific phoneme it is meant to convey, then the DNN-based sound coding module 590 can improve its representation of that particular phoneme in the augmented stimulation pattern 595 that is output for use in delivering stimulation to the recipient.

[0093] FIG. 6A is a block diagram of an ASR-based sound processor 610A, according to an example embodiment. In the ASR-based sound processor 610A of FIG. 6A, a translator algorithm and a text-to-speech algorithm are incorporated so that an audio input in the speaker’s language is converted to a stimulation sequence in the recipient’s language (i.e., the cochlear implant recipient’s language). In this example, the text-to-speech algorithm is applied to the outputs of the translator algorithm and the speaker characteristics module, and then a sound coding module is applied to the output of the text-to-speech algorithm to generate a translated stimulation pattern.

[0094] The ASR-based sound processor 610A comprises an automatic speech recognition module 630, a speaker characteristics module 640, a translator algorithm 650, a text-to-speech algorithm 660, and a sound coding module 620. Each of the automatic speech recognition module 630 and the speaker characteristics module 640 receives at least one input sound signal (audio) 605 and performs respective processing operations.

[0095] The automatic speech recognition module 630 is configured to process the at least one input sound signal 605 to convert the speech present in the at least one input sound signal 605 into speaker language text 635 (e.g., based on the one or more identified phonemes, as described above). The speaker characteristics module 640 is configured to process the input sound signal 605 to identify and output one or more speaker characteristics 645 (e.g., duration, pitch, timbre, spectrum, modulations, etc.).

[0096] The translator algorithm 650 receives the speaker language text 635 from the automatic speech recognition module 630, and is configured to convert the speaker language text 635 into recipient language text 655. The text-to-speech algorithm 660 receives the speaker characteristics 645 from the speaker characteristics module 640 and the recipient language text 655 from the translator algorithm 650, and is configured to generate an output sound signal 665 based on the recipient language text 655 and the speaker characteristics 645.

[0097] The sound coding module 620 receives the output sound signal 665 from the text-to- speech algorithm 660, and is configured to process the output sound signal 665 to generate a translated stimulation pattern 675 (i.e., a stimulation pattern for the speech that has been converted from the speaker’s language to the recipient’s language). In certain embodiments, the text-to-speech algorithm 660 can use the speaker characteristics 645 to adjust the speech in the output sound signal 665, and use the adjusted speech to modify the stimulation pattern used for the translated stimulation pattern 675. The ASR-based sound processor 610A then outputs the translated stimulation pattern 675 for use in delivering stimulation to the recipient.

[0098] Thus, in the example of FIG. 6A, language translation and speech synthesis features can be inserted into the processing pipeline of the ASR-based sound processor 610A. The automatic speech recognition module 630 converts the incoming audio to text in the speaker’s language, the speaker language text 635 is output to the translator algorithm 650 to convert the text in the speaker’s language into text in the recipient’s language, and then the text-to-speech algorithm 660 can perform speech synthesis to generate the output sound signal 665 based on the recipient language text 655 and the speaker characteristics 645, along with enhanced processing by the sound coding module 620 when generating the translated stimulation pattern 675 based on the output sound signal 665. Embedding the automatic speech recognition module 630 into the sound processing pipeline enables the addition of the translator algorithm 650. In certain examples, the translator algorithm 650 and / or the text-to-speech algorithm 660 can be machine-learning based algorithms. The output of the sound coding module 620 is a stimulation sequence that represents the translated audio.

[0099] FIG. 6B is a block diagram of an ASR-based sound processor 61 OB, according to an example embodiment. In the ASR-based sound processor 610B of FIG. 6B, a translator algorithm is incorporated so that an audio input in the speaker’s language is converted to a stimulation sequence in the recipient’s language (i.e., the cochlear implant recipient’s language). In this example, a DNN-based sound coding module is applied directly to the outputs of the translator algorithm and the speaker characteristics module to generate a translated stimulation pattern, and a text-to-speech algorithm can be omitted.[ooioo] In the specific example of FIG. 6B, the ASR-based sound processor 610B comprises the automatic speech recognition module 630, the speaker characteristics module 640, the translator algorithm 650, and a DNN-based sound coding module 690. Each of the automatic speech recognition module 630 and the speaker characteristics module 640 receives at least one input sound signal 605 and performs respective processing operations.[ooioi] The automatic speech recognition module 630 is configured to process the at least one input sound signal 605 to convert the speech present in the at least one input sound signal 605 into the speaker language text 635 (e.g., based on the one or more identified phonemes, as described above). The speaker characteristics module 640 is configured to process the at least one input sound signal 605 to identify and output the one or more speaker characteristics 645 (e.g., duration, pitch, timbre, spectrum, modulations, etc.). The translator algorithm 650 receives the speaker language text 635 from the automatic speech recognition module 630, and is configured to convert the speaker language text 635 into the recipient language text 655.

[0102] The DNN-based sound coding module 690 receives the speaker characteristics 645 from the speaker characteristics module 640 and the recipient language text 655 from the translator algorithm 650, and is configured to generate atranslated stimulation pattern 695 (i.e., a stimulation pattern for the speech present in the text that has been converted from the speaker’s language to the recipient’s language) based on the recipient language text 655 and the speaker characteristics 645. In certain embodiments, the DNN-based sound coding module 690 can use the speaker characteristics 645 to adjust the stimulation pattern that is used for the translated stimulation pattern 675 and corresponds to the recipient language text 655. The ASR- based sound processor 610B then outputs the translated stimulation pattern 695 for use in delivering stimulation to the recipient.

[0103] Thus, in the example of FIG. 6B, language translation and speech synthesis features can be inserted into the processing pipeline of the ASR-based sound processor 610B, in whichautomatic speech recognition techniques are combined with a DNN-based sound coding path. The automatic speech recognition module 630 converts the incoming audio to text in the speaker’s language, and the speaker language text 635 is output to the translator algorithm 650 to convert the text in the speaker’s language into text in the recipient’s language. As noted, the translator algorithm 650 can be a machine -learning based algorithm, for example. However, a text-to-speech algorithm can be omitted (or bypassed) in this example, and the speaker characteristics and the desired text are incorporated into a DNN-based sound coding strategy instead. In this example, the DNN-based sound coding module 690 is trained to convert the recipient language text 655 directly to a cochlear implant electrode stimulation pattern representing the translated stimulation pattern, and can perform speech synthesis when generating the translated stimulation pattern 695 with enhanced processing based on the recipient language text 655 along with the speaker characteristics 645. As noted above, embedding the automatic speech recognition module 630 into the sound processing pipeline enables the addition of the translator algorithm 650. The output of the DNN-based sound coding module 690 is a stimulation sequence that represents the translated audio.

[0104] FIG. 7 is a block diagram of an ASR-based sound processor 710, according to an example embodiment. In the ASR-based sound processor 710 of FIG. 7, a smart mixing algorithm is used to modify a stimulation pattern output from a sound coding module based on outputs from an automatic speech recognition module and a speaker characteristics module. In this example, the automatic speech recognition module is also applied to video input, in addition to at least one audio input, to generate the predicted phoneme stimulation pattern. That is, the automatic speech recognition module is an audio-visual ASR algorithm in this example.

[0105] In the specific example of FIG. 7, the ASR-based sound processor 710 comprises a sound coding module 720, an automatic speech recognition module 730, a speaker characteristics module 740, and a smart mixing algorithm 770. Each of the sound coding module 720, the automatic speech recognition module 730, and the speaker characteristics module 740 receives at least one input sound signal (audio) 705 and performs respective processing operations. In this example, the automatic speech recognition module 730 also receives at least one input video signal 715 for processing along with the input sound signal 705.

[0106] The sound coding module 720 is configured to process the at least one input sound signal 705 to generate a stimulation pattern 725 (i.e., a stimulation pattern for the full acoustic input). The automatic speech recognition module 730 is configured to process the at least oneinput sound signal 705 and the input video signal 715 to identify one or more phonemes of speech present in the input sound signal 705 and shown the input video signal 715, and to generate a predicted phoneme stimulation pattern 735 based on the one or more identified phonemes. The speaker characteristics module 740 is configured to process the input sound signal 705 to identify and output one or more speaker characteristics 745 (e.g., duration, pitch, timbre, spectrum, modulations, etc.).

[0107] The smart mixing algorithm 770 receives the stimulation pattern 725 from the sound coding module 720, the predicted phoneme stimulation pattern 735 from the automatic speech recognition module 730, and the speaker characteristics 745 from the speaker characteristics module 740, respectively. The smart mixing algorithm 770 is configured to generate an augmented stimulation pattern 775 based on the stimulation pattern 725, the predicted phoneme stimulation pattern 735, and the speaker characteristics 745. In certain embodiments, the smart mixing algorithm 770 can use the speaker characteristics 745 to adjust the predicted phoneme stimulation pattern 735, and use the adjusted phoneme stimulation pattern to modify the stimulation pattern 725 to generate the augmented stimulation pattern 775. One or more of the techniques described above with reference to the smart mixing algorithm 470 of FIG. 4 can be used, for example. The ASR-based sound processor 710 then outputs the augmented stimulation pattern 775 for use in delivering stimulation to the recipient.

[0108] Thus, in the example of FIG. 7, the output of the automatic speech recognition module 730 is matched to a phoneme template stimulation pattern, and can be mixed with sound coding as well as various speaker-specific characteristics. Further, the automatic speech recognition module 730 can be improved with visual objective measures in this example. That is, the abovedescribed audio-visual ASR algorithm of FIG. 7 can use both auditory and visual information to identify phonemes. For example, the automatic speech recognition module 730 can make use of facial features of the target speaker to improve speech recognition accuracy. In certain embodiments, the smart mixing algorithm 770 can be a machine learning algorithm or a deep neural network (DNN) that is trained to produce the augmented stimulation pattern 775 based on the stimulation pattern 725 for the full acoustic input, the predicted phoneme stimulation pattern 735, and the speaker characteristics 745. In this example, the visual information can be used to enhance the stimulation sequence that is output by the ASR-based sound processor 710 through knowledge of the sequence of phonemes being produced by the speaker.

[0109] FIG. 8 is a block diagram of an ASR-based sound processor 810, according to an example embodiment. In the ASR-based sound processor 810 of FIG. 8, a smart mixingalgorithm is used to modify a stimulation pattern output from a sound coding module based on outputs from an automatic speech recognition module and a speaker characteristics module. In this example, either a clean audio sample and / or a known speaker database can be used to inform the speaker characteristics module or to provide a “phoneme library” for a particular speaker (e.g., an individual with which the implant recipient has frequent interaction).[oono] In the specific example of FIG. 8, the ASR-based sound processor 810 comprises a sound coding module 820, an automatic speech recognition module 830, a speaker characteristics module 840, and a smart mixing algorithm 870. Each of the sound coding module 820 and the automatic speech recognition module 830 receives an input sound signal (audio) 805 and performs respective processing operations.

[0111] The sound coding module 820 is configured to process the input sound signal 805 to generate a stimulation pattern 825 (i.e., a stimulation pattern for the full acoustic input). The automatic speech recognition module 830 is configured to process the input sound signal 805 to identify one or more phonemes of speech present in the input sound signal 805, and to generate a predicted phoneme stimulation pattern 835 based on the one or more identified phonemes.

[0112] In some example embodiments, as shown in FIG. 8, the speaker characteristics module 840 can receive a clean audio sample 815, and is configured to process the clean audio sample 815 to identify one or more speaker characteristics 845 (e.g., duration, pitch, timbre, spectrum, modulations, etc.) from the speech present in the clean audio sample 815. The clean audio sample 815 can also be provided to a known speaker database 880. In addition, the speaker characteristics module 840 can receive the predicted phoneme stimulation pattern 835 from the automatic speech recognition module 830, and is configured to generate one or more speakerdependent phoneme templates 885 based on the clean audio sample 815 and the predicted phoneme stimulation pattern 835. The speaker characteristics module 840 outputs the one or more speaker-dependent phoneme templates 885 to the smart mixing algorithm 870, along with the speaker characteristics 845. The one or more speaker-dependent phoneme templates 885 can also be stored in the known speaker database 880 (e.g., for future use during conversations with this speaker) by the speaker characteristics module 840.

[0113] In some other example embodiments, as also shown in FIG. 8, the speaker characteristics module 840 can receive the clean audio sample 815 and the predicted phoneme stimulation pattern 835 from the automatic speech recognition module 830, and is configuredto identify the speaker based on the clean audio sample 815 and to retrieve one or more speakerdependent phoneme templates 885 from the known speaker database 880 based on the identified speaker and the predicted phoneme stimulation pattern 835. The speaker characteristics module 840 then provides the one or more speaker-dependent phoneme templates 885 retrieved from the known speaker database 880 to the smart mixing algorithm 870, along with the speaker characteristics 845.

[0114] The smart mixing algorithm 870 receives the stimulation pattern 825 from the sound coding module 820, the predicted phoneme stimulation pattern 835 from the automatic speech recognition module 830, and the one or more speaker-dependent phoneme templates 885 (along with the speaker characteristics 845) from the speaker characteristics module 840, respectively. The smart mixing algorithm 870 is configured to generate an augmented stimulation pattern 875 based on the stimulation pattern 825, the predicted phoneme stimulation pattern 835, the speaker characteristics 845, and the one or more speaker-dependent phoneme templates 885. In certain embodiments, the smart mixing algorithm 870 can use the one or more speakerdependent phoneme templates 885 (and / or the speaker characteristics 845) to adjust the predicted phoneme stimulation pattern 835, and use the adjusted phoneme stimulation pattern to modify the stimulation pattern 825 to generate the augmented stimulation pattern 875. One or more of the techniques described above with reference to the smart mixing algorithm 470 of FIG. 4 can be used, for example. The ASR-based sound processor 810 then outputs the augmented stimulation pattern 875 for use in delivering stimulation to the recipient.

[0115] Thus, in the example of FIG. 8, the output of the automatic speech recognition module 830 is matched to a phoneme template stimulation pattern, and can be mixed with sound coding as well as a speaker-dependent phoneme template stimulation pattern. In particular, the automatic speech recognition module 830 and the speaker characteristics module 840 can be improved with speaker-specific phoneme recognition. For example, the speaker can be recognized from the known speaker database 880 (i.e., a database of common conversation partners of the recipient, such as family, friends, coworkers, etc.), and a library of phonemes can be known for that speaker. Alternatively, the speaker could provide a clean vocal sample for the recipient, which contains most or all of the phonemes that might be spoken by that person. In the English language, for example, there are 44 phonemes, as noted above. Unvoiced consonants tend to be quite similar across speakers, while voiced consonants and vowels tend to carry most of the unique auditory information. These voiced consonants and vowels (or features from voiced consonants and vowels) can be stored in a “speaker-dependent phonemetemplate” within the known speaker database 880 and then used to enhance the stimulation sequence output by the ASR-based sound processor 810. In certain embodiments, the smart mixing algorithm 870 can be a machine learning algorithm or a deep neural network (DNN) that is trained to produce the augmented stimulation pattern 875 based on the stimulation pattern 825 for the full acoustic input, the predicted phoneme stimulation pattern 835, the speaker characteristics 845, and the one or more speaker-dependent phoneme templates 885.

[0116] FIG. 9 is a block diagram of an ASR-based sound processor 910, according to an example embodiment. In the ASR-based sound processor 910 of FIG. 9, phoneme predictions from an automatic speech recognition module can be used to control a noise reduction algorithm.

[0117] In the specific example of FIG. 9, the ASR-based sound processor 910 comprises a sound coding module 920, an automatic speech recognition module 930, a speaker characteristics module 940, and a noise reduction algorithm 990. Each of the sound coding module 920, the automatic speech recognition module 930, and the speaker characteristics module 940 receives an input sound signal 905 (audio) and performs respective processing operations.

[0118] The sound coding module 920 is configured to process the input sound signal 905 to generate a stimulation pattern 925 (i.e., a stimulation pattern for the full acoustic input). In this example, the automatic speech recognition module 930 is configured to process the input sound signal 905 to identify a phoneme of speech present in the input sound signal 905, and output a predicted phoneme 935 based on the one or more identified phonemes. For example, a representation of the identified phoneme itself (e.g., “ah”), or alternatively, a corresponding phoneme code that identifies the respective phoneme present in the speech (e.g., a number from 1-44 based on a mapping table or other data structure), can be output as the predicted phoneme 935. The speaker characteristics module 940 is configured to process the input sound signal 905 to identify and output one or more speaker characteristics 945 (e.g., duration, pitch, timbre, spectrum, modulations, etc.).

[0119] The noise reduction algorithm 990 receives the stimulation pattern 925 from the sound coding module 920, the predicted phoneme 935 from the automatic speech recognition module 930, and the speaker characteristics 945 from the speaker characteristics module 940, respectively. The noise reduction algorithm 990 is configured to generate an augmented stimulation pattern 995 (i.e., a “noise-reduced” stimulation pattern representing the speechpresent in the input sound signal 905 with most or all of the noise removed) based on the stimulation pattern 925, the predicted phoneme 935, and the speaker characteristics 945. In certain embodiments, the noise reduction algorithm 990 can modify the stimulation pattern 925 based on the predicted phoneme 935 and the speaker characteristics 945 to generate the augmented stimulation pattern 995. The ASR-based sound processor 910 then outputs the augmented stimulation pattern 995 for use in delivering stimulation to the recipient.

[0120] Thus, in the example of FIG. 9, noise reduction can be performed using ASR-based phoneme prediction techniques. The predicted phoneme 935 can be explicitly used to control the noise reduction algorithm 990. The noise reduction algorithm 990 can be a machinelearning based algorithm, for example. In this example, the noise reduction algorithm 990 does not necessarily perform “mixing” of the stimulation pattern 925, the predicted phoneme 935, and the speaker characteristics 945, per se (as otherwise described above in connection with the smart mixing algorithm 470 of FIG. 4, for example). Rather, the predicted phoneme 935 and the speaker characteristics 945 are used as additional inputs to inform the noise reduction processing performed by the noise reduction algorithm 990 with respect to the stimulation pattern 925. If the noise reduction algorithm 990 is made aware of the phoneme of the target speaker, then the noise reduction algorithm 990 can more easily remove the noise from the input sound signal 905 that is represented in the stimulation pattern 925 when generating the augmented stimulation pattern 995 for use in delivering stimulation to the recipient.

[0121] FIG. 10 is a flowchart illustrating a method 1000 according to an aspect of the techniques presented here. Method 1000 beings 1010, the method 1000 includes receiving one or more sound signals at a hearing device of a recipient. At operation 1020, the method 1000 includes determining a direct stimulation pattern based on the one or more sound signals (e.g., using a sound coding module). At operation 1030, the method 1000 includes determining a predicted phoneme stimulation pattern based on one or more phonemes of speech present in the one or more sound signals. At operation 1040, the method 1000 includes generating an augmented stimulation pattern for delivery to the recipient based at least one the direct stimulation pattern and the predicted phoneme stimulation pattern.

[0122] FIG. 11 is a flowchart illustrating a method 1100 according to an aspect of the techniques presented herein. Method 1100 begins at 1110 where an implantable medical device receives one or more input signals. At 1120, the implantable medical device determines one or more predicted phoneme codes based on one or more phonemes of speech present in the one or more sound signals. At 1130, the implantable medical device generates an augmentedstimulation pattern for delivery to the recipient based on the one or more sound signals and the one or more predicted phoneme codes.

[0123] Thus, as described above with reference to FIGs. 2, 3, 4, 5, 6A-6B, 7, 8, 9, 10, and 11, presented herein are various techniques for improving speech perception for individuals with cochlear implants in noisy environments by combining automatic speech recognition technology with sound coding modules. By augmenting the encoding of phonemic information, the various approaches described herein have the potential to provide enhanced speech intelligibility in noise for cochlear implant recipients.

[0124] As should be appreciated, while particular uses of the technology have been illustrated and discussed above, the disclosed technology can be used with a variety of devices in accordance with many examples of the technology. The above discussion is not meant to suggest that the disclosed technology is only suitable for implementation within systems akin to that illustrated in the figures. In general, additional configurations can be used to practice the processes and systems herein and / or some aspects described can be excluded without departing from the processes and systems disclosed herein.

[0125] This disclosure described some aspects of the present technology with reference to the accompanying drawings, in which only some of the possible aspects were shown. Other aspects can, however, be embodied in many different forms and should not be construed as limited to the aspects set forth herein. Rather, these aspects were provided so that this disclosure was thorough and complete and fully conveyed the scope of the possible aspects to those skilled in the art.

[0126] As should be appreciated, the various aspects (e.g., portions, components, etc.) described with respect to the figures herein are not intended to limit the systems and processes to the particular aspects described. Accordingly, additional configurations can be used to practice the methods and systems herein and / or some aspects described can be excluded without departing from the methods and systems disclosed herein.

[0127] According to certain aspects, systems and non-transitory computer readable storage media are provided. The systems are configured with hardware configured to execute operations analogous to the methods of the present disclosure. The one or more non-transitory computer readable storage media comprise instructions that, when executed by one or more processors, cause the one or more processors to execute operations analogous to the methods of the present disclosure.

[0128] Similarly, where steps of a process are disclosed, those steps are described for purposes of illustrating the present methods and systems and are not intended to limit the disclosure to a particular sequence of steps. For example, the steps can be performed in differing order, two or more steps can be performed concurrently, additional steps can be performed, and disclosed steps can be excluded without departing from the present disclosure. Further, the disclosed processes can be repeated.

[0129] Although specific aspects were described herein, the scope of the technology is not limited to those specific aspects. One skilled in the art will recognize other aspects or improvements that are within the scope of the present technology. Therefore, the specific structure, acts, or media are disclosed only as illustrative aspects. The scope of the technology is defined by the following claims and any equivalents therein.

[0130] It is also to be appreciated that the embodiments presented herein are not mutually exclusive and that the various embodiments may be combined with another in any of a number of different manners.

Claims

CLAIMSWhat is claimed is:

1. A method comprising : receiving one or more sound signals at a hearing device of a recipient; determining a direct stimulation pattern based on the one or more sound signals; determining a predicted phoneme stimulation pattern based on the one or more sound signals; and generating an augmented stimulation pattern for delivery to the recipient based at least on the direct stimulation pattern and the predicted phoneme stimulation pattern.

2. The method of claim 1, wherein determining a predicted phoneme stimulation pattern based on the one or more sound signals comprises: processing the one or more sound signals using an automatic speech recognition module.

3. The method of claim 1, wherein determining a predicted phoneme stimulation pattern based on the one or more sound signals comprises: identifying one or more phonemes of speech present in the one or more sound signals; and generating the predicted phoneme stimulation pattern based on the one or more phonemes.

4. The method of claim 1, 2, or 3, further comprising: determining one or more speaker characteristics based on at least the one or more sound signals, wherein generating an augmented stimulation pattern for delivery to the recipient further comprises: generating the augmented stimulation pattern based on the direct stimulation pattern, the predicted phoneme stimulation pattern, and the one or more speaker characteristics.

5. The method of claim 4, wherein determining one or more speaker characteristics based on at least the one or more sound signals comprises: extracting one or more features from the one or more sound signals, wherein the one or more features comprise one or more of duration, pitch, timbre, spectrum information, or modulations; and determining the one or more speaker characteristics based on the one or more features extracted from the one or more sound signals.

6. The method of claim 4, wherein generating the augmented stimulation pattern based on the direct stimulation pattern, the predicted phoneme stimulation pattern, and the one or more speaker characteristics comprises: determining a speaker-specific phoneme stimulation pattern based on the predicted phoneme stimulation pattern and the one or more speaker characteristics; and generating the augmented stimulation pattern based on the direct stimulation pattern and the speaker-specific phoneme stimulation pattern.

7. The method of claim 6, wherein determining a speaker-specific stimulation pattern based on the predicted phoneme stimulation pattern and the one or more speaker characteristics comprises: modifying the predicted phoneme stimulation pattern to reflect the one or more speaker characteristics in the speaker-specific phoneme stimulation pattern.

8. The method of claim 4, wherein generating the augmented stimulation pattern based on the direct stimulation pattern, the predicted phoneme stimulation pattern, and the one or more speaker characteristics comprises: generating the augmented stimulation pattern using a smart mixing algorithm configured to mix the direct stimulation pattern, the predicted phoneme stimulation pattern, and the one or more speaker characteristics according to a predetermined mixing factor.

9. The method of claim 8, wherein the smart mixing algorithm is implemented using one or more machine learning algorithms or a deep neural network (DNN) algorithm configured to produce the augmented stimulation pattern based on a mixture of the direct stimulation pattern, the predicted phoneme stimulation pattern, and the one or more speaker characteristics.

10. The method of claim 8, further comprising: receiving one or more input video signals associated with the one or more sound signals; wherein determining a predicted phoneme stimulation pattern based on at least the one or more sound signals comprises: determining the predicted phoneme stimulation pattern based on the one or more sound signals and the one or more input video signals.

11. The method of claim 8, wherein determining one or more speaker characteristics based on at least the one or more sound signals comprises: determining one or more speaker-dependent phoneme templates associated with a speaker based on a clean audio sample obtained from the speaker or using a known speaker database configured to store speaker-dependent phoneme templates associated with respective speakers; wherein the smart mixing algorithm is configured to generate the augmented stimulation pattern further based on the one or more speaker-dependent phoneme templates associated with the speaker.

12. The method of claim 4, wherein generating an augmented stimulation pattern for delivery to the recipient further comprises: adjusting the direct stimulation pattern based on the predicted phoneme stimulation pattern, and the one or more speaker characteristics using a noise reduction algorithm configured to generate the augmented stimulation pattern.

13. The method of claim 1, 2, or 3, wherein generating an augmented stimulation pattern for delivery to the recipient further comprises: adjusting the direct stimulation pattern based on the predicted phoneme stimulation pattern using a noise reduction algorithm configured to generate the augmented stimulation pattern.

14. The method of claim 1, 2, or 3, further comprising: delivering the augmented stimulation pattern to the recipient via one or more electrodes implanted in the recipient.

15. A method comprising : receiving one or more input signals at an implantable medical device; determining one or more predicted phoneme codes based on one or more phonemes of speech present in the one or more input signals; and generating an augmented stimulation pattern for delivery to a recipient of the implantable medical device based on the one or more input signals and the one or more predicted phoneme codes.

16. The method of claim 15, wherein determining one or more predicted phoneme codes based on the one or more input signals comprises: processing the one or more input signals using an automatic speech recognition module configured to identify the one or more phonemes of speech present in the one or more input signals and generate the one or more predicted phoneme codes based thereon.

17. The method of claim 16, wherein generating an augmented stimulation pattern comprises: processing the one or more input signals and the one or more predicted phoneme codes with a deep neural network (DNN)-based sound coding module.

18. The method of claim 17, further comprising: determining one or more speaker characteristics based on the one or more input signals; and processing the one or more input signals, the one or more predicted phoneme codes, and the one or more speaker characteristics to generate the augmented stimulation pattern with the DNN-based sound coding module.

19. The method of claim 15, 16, 17, or 18, further comprising: delivering the augmented stimulation pattern to the recipient via one or more electrodes implanted in the recipient.

20. One or more non-transitory computer readable storage media comprising instructions that, when executed by one or more processors of a hearing device associated with a recipient, are operable to:obtain one or more sound signals that include speech in a first language of a speaker; convert the one or more sound signals into speaker language text representing the speech in the first language of the speaker; translate the speaker language text into recipient language text in a second language of the recipient; determine one or more speaker characteristics associated with the speaker; and generate a translated stimulation pattern for delivery to the recipient based on the recipient language text in the second language of the recipient and the one or more speaker characteristics associated with the speaker, wherein the translated stimulation pattern represents translated speech in the second language of the recipient.

21. The one or more non-transitory computer readable storage media of claim 20, wherein the instructions operable to determine the speaker language text representing the speech in the first language of the speaker comprise instructions operable to: process the one or more sound signals using an automatic speech recognition module to convert the one or more sound signals into a predicted phoneme stimulation pattern associated with the speech in the first language of the speaker.

22. The one or more non-transitory computer readable storage media of claim 20 or 21, wherein the instructions operable to determine the one or more speaker characteristics associated with the speaker comprise instructions operable to: extract one or more features from the one or more sound signals, wherein the one or more features comprise one or more of duration, pitch, timbre, spectrum information, or modulations associated with the speech in the first language of the speaker; and determine the one or more speaker characteristics based on the one or more features extracted from the one or more sound signals.

23. The one or more non-transitory computer readable storage media of claim 20 or 21, further comprising instructions operable to: process the recipient language text in the second language of the recipient and the one or more speaker characteristics associated with the speaker using a text-to-speech algorithm configured to generate one or more output sound signals representing the translated speech in the second language of the recipient;wherein the instructions operable to generate a translated stimulation pattern for delivery to the recipient comprise instructions operable to: process the one or more output sound signals representing the translated speech in the second language of the recipient using a sound coding module to generate the translated stimulation pattern.

24. The one or more non-transitory computer readable storage media of claim 20 or 21, further comprising instructions operable to: process the recipient language text in the second language of the recipient and the one or more speaker characteristics associated with the speaker using a deep neural network (DNN) based sound coding module configured to generate the translated stimulation pattern representing the translated speech in the second language of the recipient.

25. The one or more non-transitory computer readable storage media of claim 20 or 21, further comprising instructions operable to: initiate delivery of the translated stimulation pattern to the recipient via one or more electrodes implanted in the recipient.

26. A hearing device, comprising: one or more sound inputs configured to receive at least one sound signal; a memory; and one or more processors configured to: determine a direct stimulation pattern from the at least one sound signal, determine one or more phonemes of speech present in the at least one sound signal, determine a predicted phoneme stimulation pattern based on one or more phonemes of speech present in the at least one sound signal, generate an augmented stimulation pattern based on the direct stimulation pattern and the predicted phoneme stimulation pattern, andgenerate one or more stimulation signals based on the augmented stimulation patern.

27. The hearing device of claim 26, further comprising: one or more electrodes configured to deliver the one or more stimulation signals.

28. The hearing device of claim 26, wherein the one or more processors are configured to execute an automatic speech recognition process to determine the one or more phonemes of speech present in the at least one sound signal, and to determine the predicted phoneme stimulation patern based on one or more phonemes of speech present in the at least one sound signal.

29. The hearing device of claim 26, 27, or 28, wherein the one or more processors are further configured to: determine one or more speaker characteristics based on the speech present in the at least one sound signal, and generate the augmented stimulation patern based on the direct stimulation patern, the predicted phoneme stimulation patern, and the one or more speaker characteristics.

Citation Information

Patent Citations

  • Audio processing method, electronic equipment and storage medium

    CN114120965A

  • Language translation based on speaker-related information

    US20130144595A1

  • System comprising a cochlear stimulation device and a second hearing stimulation device and a method for adjustment according to a response to combined stimulation

    US20160175591A1

  • Somatic, auditory and cochlear communication system and method

    US20200152223A1

  • System and method for identifying and processing audio signals

    US20220383884A1