New processing techniques
The system addresses the issue of medical devices losing function when external components are removed by using a convolutional based subsystem to process data and generate electrode activation control data, enabling continuous functionality and sensory perception in medical devices like cochlear implants.
Patent Information
- Application Number
- PCT/IB2024/060822
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-11-01
- Publication Date
- 2025-05-08
AI Technical Summary
Existing medical devices, such as cochlear implants, often rely on external components for power and data, which can lead to loss of function when the external device is removed, such as during sleep, resulting in inability to hear ambient sounds.
A system comprising an input subsystem and a convolutional based subsystem that processes audio and visual data using machine learning algorithms to generate electrode activation control data, enabling direct stimulation of electrodes in medical devices like cochlear implants without the need for external components.
The system allows for continuous and consistent functionality of medical devices by processing ambient sound and light data to produce recognizable sounds and images, and generating electrode activation sequences that evoke sensory percepts, such as hearing, without the need for external power or data sources.
Smart Images

Figure IB2024060822_08052025_PF_FP_ABST
Abstract
Description
NEW PROCESSING TECHNIQUESCROSS-REFERENCE TO RELATED APPLICATIONS[oooi] This application claims priority to U.S. Provisional Application No. 63 / 547,315, entitled PROCESSING TECHNIQUES, filed on November 3, 2023, naming Timothy Jean BROCHIER as an inventor, the entire contents of that application being incorporated herein by reference in its entirety.BACKGROUND
[0002] Medical devices have provided a wide range of therapeutic benefits to recipients over recent decades. Medical devices can include internal or implantable components / devices, external or wearable components / devices, or combinations thereof (e.g., a device having an external component communicating with an implantable component). Medical devices, such as traditional hearing aids, partially or fully-implantable hearing prostheses (e.g., bone conduction devices, mechanical stimulators, cochlear implants, etc.), pacemakers, defibrillators, functional electrical stimulation devices, and other medical devices, have been successful in performing lifesaving and / or lifestyle enhancement functions and / or recipient monitoring for a number of years.
[0003] The types of medical devices and the ranges of functions performed thereby have increased over the years. For example, many medical devices, sometimes referred to as “implantable medical devices,” now often include one or more instruments, apparatus, sensors, processors, controllers or other functional mechanical or electrical components that are permanently or temporarily implanted in a recipient. These functional devices are typically used to diagnose, prevent, monitor, treat, or manage a disease / injury or symptom thereof, or to investigate, replace or modify the anatomy or a physiological process. Many of these functional devices utilize power and / or data received from external devices that are part of, or operate in conjunction with, implantable components.SUMMARY
[0004] In an exemplary embodiment, there is a system, comprising an input subsystem configured to receive input; and a convolutional based sub-system in signal communication with the input subsystem, wherein the convolutional based sub-system outputs electrode activation sequence based data.
[0005] In an exemplary embodiment, there is a device comprising a product of machine learning, wherein the product is configured to receive an input based on a captured phenomenon in the visual and / or audio environment, and the product is configured to process the input and automatically develop electrode activation control data.
[0006] In an exemplary embodiment, there is a method, comprising receiving data based on at least one of ambient sound and / or ambient light in an ambient environment of a device or electromagnetic signals received by the device, the data enabling direct usage thereof by a conventional speaker apparatus and / or a conventional video monitor apparatus to produce a recognizable corresponding sound and / or image without further processing and generating electrode activation sequences directly from the received data.
[0007] In an exemplary embodiment, there is a method, comprising obtaining data based on audio and / or visual data, processing the obtained data using one or more algorithms based on machine learning and evoking a sensory percept based on the processed data, wherein the action of processing the data using one or more algorithms one or more of: executes output side processing of a cochlear implant processing pipeline; or executes one or more of: amplitude compression; channel envelope sampling; maxima selection; speech enhancement; binaural processing; or microphone equalization.
[0008] In an exemplary embodiment, there is a method, comprising obtaining output of signal processing of an input signal and providing data based on the obtained output to a convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided data, wherein the output is electrical stimulation control data for a sensory prosthesis.
[0009] In an exemplary embodiment, there is a device comprising microphone and a stimulation output assembly, wherein the device is a hearing prosthesis, the device includes a sound processing path that receives input from the implanted microphone, the device controls the tissue stimulation output assembly based on output from the sound processing path to evoke a hearing percept based on the input from the implanted microphone, and at least one of:the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; the sound processing path includes a product of machine learning trained to minimize feedback; or a processing path that includes at least a portion of the sound processing path and extends to the tissue stimulation output assembly is made entirely of a product of machine learning.
[0010] In an exemplary embodiment, there is a device, comprising: a signal input configured to receive data based on audio data and / or visual data; and circuitry configured to simultaneously functionally execute at least some input side processing and at least some output side processing of the data based on audio data and / or visual data.BRIEF DESCRIPTION OF THE DRAWINGS[ooii] Embodiments are described below with reference to the attached drawings, in which:
[0012] FIG. l is a perspective view of an exemplary hearing prosthesis;
[0013] FIG. 2 presents a functional block diagram of an exemplary cochlear implant;
[0014] FIG. 3A and FIG. 3B and 3C present exemplary systems of communication between devices;
[0015] FIG. 4 presents an exemplary retinal prosthesis;
[0016] FIGs. 5-11 present exemplary functional diagram for exemplary embodiments;
[0017] FIGs. 12-24 present exemplary flowcharts for exemplary methods;
[0018] FIGs. 25-26 present exemplary functional diagram for exemplary embodiments; and
[0019] FIG. 27 presents conceptual electrode activation sequence data.DETAILED DESCRIPTION
[0020] Merely for ease of description, the techniques presented herein are described herein with reference by way of background to an illustrative medical device, namely a cochlear implant. However, it is to be appreciated that the techniques presented herein may also be used with a variety of other medical devices that, while providing a wide range of therapeutic benefits to recipients, patients, or other users, may benefit from setting changes based on thelocation of the medical device. For example, the techniques presented herein may be used to determine the viability of various types of prostheses, such as, for example, a vestibular implant and / or a retinal implant, with respect to a particular human being. And with regard to the latter, the techniques presented herein are also described with reference by way of background to another illustrative medical device, namely a retinal implant. The techniques presented herein are also applicable to the technology of vestibular devices (e.g., vestibular implants), visual devices (i.e., bionic eyes), sensors, pacemakers, drug delivery systems, defibrillators, functional electrical stimulation devices, catheters, seizure devices (e.g., devices for monitoring and / or treating epileptic events), sleep apnea devices, electroporation, etc.
[0021] Also, embodiments are directed to other types of hearing prostheses, such as middle ear implants, bone conduction devices (active transcutaneous, passive transcutaneous, percutaneous), and conventional hearing aids. Thus, embodiments are directed to devices that include implantable portions and embodiments that do not include implantable portions.
[0022] Any reference to one of the above-noted sensory prostheses corresponds to an alternate disclosure using one of the other above-noted sensory prostheses unless otherwise noted, providing that the art enables such.
[0023] FIG. 1 is a perspective view of a cochlear implant, referred to as cochlear implant 100, implanted in a recipient, to which some embodiments detailed herein and / or variations thereof are applicable. Particularly, as will be detailed below, there are aspects of a cochlear implant that are utilized with respect to a vestibular implant, and thus there is utility in describing features of the cochlear implant for purposes of understanding a vestibular implant. The cochlear implant 100 is part of a system 10 that can include external components in some embodiments, as will be detailed below. Additionally, it is noted that the teachings detailed herein are also applicable to other types of hearing prostheses, such as, by way of example only and not by way of limitation, bone conduction devices (percutaneous, active transcutaneous and / or passive transcutaneous), direct acoustic cochlear stimulators, middle ear implants, and conventional hearing aids, etc. Indeed, it is noted that the teachings detailed herein are also applicable to so-called multi-mode devices. In an exemplary embodiment, these multi-mode devices apply both electrical stimulation and acoustic stimulation to the recipient. In an exemplary embodiment, these multi-mode devices evoke a hearing percept via electrical hearing and bone conduction hearing.
[0024] In view of the above, it is to be understood that at least some embodiments detailed herein and / or variations thereof are directed towards a body-worn sensory supplement medical device (e.g., the hearing prosthesis of FIG. 1, which supplements the hearing sense, even in instances when there are no natural hearing capabilities, for example, due to degeneration of previous natural hearing capability or to the lack of any natural hearing capability, for example, from birth). Again, it is noted that at least some exemplary embodiments of some sensory supplement medical devices are directed towards devices such as conventional hearing aids, which supplement the hearing sense in instances where some natural hearing capabilities have been retained, and visual prostheses (both those that are applicable to recipients having some natural vision capabilities and to recipients having no natural vision capabilities). Accordingly, the teachings detailed herein are applicable to any type of sensory supplement medical device to which the teachings detailed herein are enabled for use therein in a utilitarian manner. In this regard, the phrase sensory supplement medical device refers to any device that functions to provide sensation to a recipient irrespective of whether the applicable natural sense is only partially impaired or completely impaired, or indeed never existed.
[0025] The recipient has an outer ear 101, a middle ear 105, and an inner ear 107. Components of outer ear 101, middle ear 105, and inner ear 107 are described below, followed by a description of cochlear implant 100.
[0026] In a fully functional ear, outer ear 101 comprises an auricle 110 and an ear canal 102. An acoustic pressure or sound wave 103 is collected by auricle 110 and channeled into and through ear canal 102. Disposed across the distal end of ear channel 102 is a tympanic membrane 104 which vibrates in response to sound wave 103. This vibration is coupled to oval window or fenestra ovalis 112 through three bones of middle ear 105, collectively referred to as the ossicles 106 and comprising the malleus 108, the incus 109, and the stapes 111. Bones 108, 109, and 111 of middle ear 105 serve to filter and amplify sound wave 103, causing oval window 112 to articulate, or vibrate in response to vibration of tympanic membrane 104. This vibration sets up waves of fluid motion of the perilymph within cochlea 140. Such fluid motion, in turn, activates tiny hair cells (not shown) inside of cochlea 140. Activation of the hair cells causes appropriate nerve impulses to be generated and transferred through the spiral ganglion cells (not shown) and auditory nerve 114 to the brain (also not shown) where they are perceived as sound.
[0027] As shown, cochlear implant 100 comprises one or more components which are temporarily or permanently implanted in the recipient. Cochlear implant 100 is shown in FIG. 1 with an external device 142, that is part of system 10 (along with cochlear implant 100), which, as described below, is configured to provide power to the cochlear implant, where the implanted cochlear implant includes a battery that is recharged by the power provided from the external device 142.
[0028] In the illustrative arrangement of FIG. 1, external device 142 can comprise a power source (not shown) disposed in a Behind-The-Ear (BTE) unit 126. External device 142 also includes components of a transcutaneous energy transfer link, referred to as an external energy transfer assembly. The transcutaneous energy transfer link is used to transfer power and / or data to cochlear implant 100. Various types of energy transfer, such as infrared (IR), electromagnetic, capacitive and inductive transfer, may be used to transfer the power and / or data from external device 142 to cochlear implant 100. In the illustrative embodiments of FIG. 1, the external energy transfer assembly comprises an external coil 130 that forms part of an inductive radio frequency (RF) communication link. External coil 130 is typically a wire antenna coil comprised of multiple turns of electrically insulated single-strand or multistrand platinum or gold wire. External device 142 also includes a magnet (not shown) positioned within the turns of wire of external coil 130. It should be appreciated that the external device shown in FIG. 1 is merely illustrative, and other external devices may be used with embodiments.
[0029] Cochlear implant 100 comprises an internal energy transfer assembly 132 which can be positioned in a recess of the temporal bone adjacent auricle 110 of the recipient. As detailed below, internal energy transfer assembly 132 is a component of the transcutaneous energy transfer link and receives power and / or data from external device 142. In the illustrative embodiment, the energy transfer link comprises an inductive RF link, and internal energy transfer assembly 132 comprises a primary internal coil 136. Internal coil 136 is typically a wire antenna coil comprised of multiple turns of electrically insulated singlestrand or multi-strand platinum or gold wire.
[0030] Cochlear implant 100 further comprises a main implantable component 120 and an elongate electrode assembly 118. In some embodiments, internal energy transfer assembly 132 and main implantable component 120 are hermetically sealed within a biocompatible housing. In some embodiments, main implantable component 120 includes an implantable microphone assembly (not shown) and a sound processing unit (not shown) to convert thesound signals received by the implantable microphone in internal energy transfer assembly 132 to data signals. That said, in some alternative embodiments, the implantable microphone assembly can be located in a separate implantable component (e.g., that has its own housing assembly, etc.) that is in signal communication with the main implantable component 120 (e.g., via leads or the like between the separate implantable component and the main implantable component 120). In at least some embodiments, the teachings detailed herein and / or variations thereof can be utilized with any type of implantable microphone arrangement.
[0031] Main implantable component 120 further includes a stimulator unit (also not shown) which generates electrical stimulation signals based on the data signals. The electrical stimulation signals are delivered to the recipient via elongate electrode assembly 118.
[0032] Elongate electrode assembly 118 has a proximal end connected to main implantable component 120, and a distal end implanted in cochlea 140. Electrode assembly 118 extends from main implantable component 120 to cochlea 140 through mastoid bone 119. In some embodiments electrode assembly 118 may be implanted at least in basal region 116, and sometimes further. For example, electrode assembly 118 may extend towards apical end of cochlea 140, referred to as cochlea apex 134. In certain circumstances, electrode assembly 118 may be inserted into cochlea 140 via a cochleostomy 122. In other circumstances, a cochleostomy may be formed through round window 121, oval window 112, the promontory 123 or through an apical turn 147 of cochlea 140.
[0033] Electrode assembly 118 comprises a longitudinally aligned and distally extending array 146 of electrodes 148, disposed along a length thereof. As noted, a stimulator unit generates stimulation signals which are applied by electrodes 148 to cochlea 140, thereby stimulating auditory nerve 114.
[0034] Thus, as seen above, one variety of implanted devices depends on an external component to provide certain functionality and / or power. For example, the recipient of the implanted device can wear an external component that provides power and / or data (e.g., a signal representative of sound) to the implanted portion that allow the implanted device to function. In particular, the implanted device can lack a battery and can instead be totally dependent on an external power source providing continuous power for the implanted device to function. Although the external power source can continuously provide power, characteristics of the provided power need not be constant and may fluctuate. Additionally,where the implanted device is an auditory prosthesis such as a cochlear implant, the implanted device can lack its own sound input device (e.g., a microphone). It is sometimes utilitarian to remove the external component. For example, it is common for a recipient of an auditory prosthesis to remove an external portion of the prosthesis while sleeping. Doing so can result in loss of function of the implanted portion of the prosthesis, which can make it impossible for recipient to hear ambient sound. This can be less than utilitarian and can result in the recipient being unable to hear while sleeping. Loss of function would also prevent the implanted portion from responding to signals representative of streamed content (e.g., music streamed from a phone) or providing other functionality, such as providing tinnitus suppression noise.
[0035] The external component that provides power and / or data can be worn by the recipient, as detailed above. While a wearable external device is worn by a recipient, the external device is typically in very close proximity and tightly aligned with an implanted component. The wearable external device can be configured to operate in these conditions. Conversely, in some instances, an unworn device can generally be further away and less tightly aligned with the implanted component. This can create difficulties where the implanted device depends on an external device for power and data (e.g., where the implanted device lacks its own battery and microphone), and the external device can need to continuously and consistently provide power and data in order to allow for continuous and consistent functionality of the implanted device.
[0036] FIG. 2 is a functional block diagram of a cochlear implant system 200 to which the teaching herein can be applicable. The cochlear implant system 200 includes an implantable component 201 (e.g., implantable component 100 of FIG. 1) configured to be implanted beneath a recipient’s skin or other tissue 249, and an external device 240 (e.g., the external device 142 of FIG. 1).
[0037] The external device 240 can be configured as a wearable external device, such that the external device 240 is worn by a recipient in close proximity to the implantable component, which can enable the implantable component 201 to receive power and stimulation data from the external device 240. As described in FIG. 1, magnets can be used to facilitate an operational alignment of the external device 240 with the implantable component 201. With the external device 240 and implantable component 201 in close proximity, the transfer of power and data can be accomplished through the use of near-field electromagnetic radiation,and the components of the external device 240 can be configured for use with near-field electromagnetic radiation.
[0038] Implantable component 201 can include a transceiver unit 208, electronics module 213, which module can be a stimulator assembly of a cochlear implant, and an electrode assembly 254 (which can include an array of electrode contacts disposed on lead 118 of FIG. 1). The transceiver unit 208 is configured to transcutaneously receive power and / or data from external device 240. As used herein, transceiver unit 208 refers to any collection of one or more components which form part of a transcutaneous energy transfer system. Further, transceiver unit 208 can include or be coupled to one or more components that receive and / or transmit data or power. For example, the example includes a coil for a magnetic inductive arrangement coupled to the transceiver unit 208. Other arrangements are also possible, including an antenna for an alternative RF system, capacitive plates, or any other utilitarian arrangement. In an example, the data modulates the RF carrier or signal containing power. The transcutaneous communication link established by the transceiver unit 208 can use time interleaving of power and data on a single RF channel or band to transmit the power and data to the implantable component 201. In some examples, the processor 244 is configured to cause the transceiver unit 246 to interleave power and data signals, such as is described in U.S. Patent Publication Number 2009 / 0216296 to Meskens. In this manner, the data signal is modulated with the power signal, and a single coil can be used to transmit power and data to the implanted component 201. Various types of energy transfer, such as infrared (IR), electromagnetic, capacitive and inductive transfer, can be used to transfer the power and / or data from the external device 240 to the implantable component 201.
[0039] Aspects of the implantable component 201 can require a source of power to provide functionality, such as receive signals, process data, or deliver electrical stimulation. The source of power that directly powers the operation of the aspects of the implantable component 201 can be described as operational power. There are two exemplary ways that the implantable component 201 can receive operational power: a power source internal to the implantable component 201 (e.g., a battery) or a power source external to the implantable component. However, other approaches or combinations of approaches are possible. For example, the implantable component may have a battery but nonetheless receive operational power from the external component (e.g., to preserve internal battery life when the battery is sufficiently charged).
[0040] The internal power source can be a power storage element (not pictured). The power storage element can be configured for the long-term storage of power, and can include, for example, one or more rechargeable batteries. Power can be received from an external source, such as the external device 240, and stored in the power storage element for long-term use (e.g., charge a battery of the power storage element). The power storage element can then provide power to the other components of the implantable component 201 over time as needed for operation without needing an external power source. In this manner, the power from the external source may be considered charging power rather than operational power, because the power from the external power source is for charging the battery (which in turn provides operational power) rather than for directly powering aspects of the implantable component 201 that require power to operate. The power storage element can be a long-term power storage element configured to be a primary power source for the implantable component 201.
[0041] In some embodiments, the implantable component 201 receives operational power from the external device 240 and the implantable component 201 does not include an internal power source (e.g., a battery) / internal power storage device. In other words, the implantable component 201 is powered solely by the external device 240 or another external device, which provides enough power to the implantable component 201 to allow the implantable component to operate (e.g., receive data signals and take an action in response). The operational power can directly power functionality of the device rather than charging a power storage element of the external device implantable component 201. In these examples, the implantable component 201 can include incidental components that can store a charge (e.g., capacitors) or small amounts of power, such as a small battery for keeping volatile memory powered or powering a clock (e.g., motherboard CMOS batteries). But such incidental components would not have enough power on their own to allow the implantable component to provide primary functionality of the implantable component 201 (e.g., receiving data signals and taking an action in response thereto, such as providing stimulation) and therefore cannot be said to provide operational power even if they are integral to the operation of the implantable component 201.
[0042] As shown, electronics module 213 includes a stimulator unit 214 (e.g., which can correspond to the stimulator of FIG. 1). Electronics module 213 can also include one or more other components used to generate or control delivery of electrical stimulation signals 215 to the recipient. As described above with respect to FIG. 1, a lead (e.g., elongate lead 118 ofFIG. 1) can be inserted into the recipient’s cochlea. The lead can include an electrode assembly 254 configured to deliver electrical stimulation signals 215 generated by the stimulator unit 214 to the cochlea.
[0043] In the example system 200 depicted in FIG. 2, the external device 240 includes a sound input unit 242, a sound processor 244, a transceiver unit 246, a coil 247, and a power source 248. The sound input unit 242 is a unit configured to receive sound input. The sound input unit 242 can be configured as a microphone (e.g., arranged to output audio data that is representative of a surrounding sound environment), an electrical input (e.g., a receiver for a frequency modulation (FM) hearing system), and / or another component for receiving sound input. The sound input unit 242 can be or include a mixer for mixing multiple sound inputs together.
[0044] The processor 244 is a processor configured to control one or more aspects of the system 200, including converting sound signals received from sound input unit 242 into data signals and causing the transceiver unit 246 to transmit power and / or data signals. The transceiver unit 246 can be configured to send or receive power and / or data 251. For example, the transceiver unit 246 can include circuit components that send power and data (e.g., inductively) via the coil 247. The data signals from the sound processor 244 can be transmitted, using the transceiver unit 246, to the implantable component 201 for use in providing stimulation or other medical functionality.
[0045] The transceiver unit 246 can include one or more antennas or coils for transmitting the power or data signal, such as coil 247. The coil 247 can be a wire antenna coil having of multiple turns of electrically insulated single-strand or multi-strand wire. The electrical insulation of the coil 247 can be provided by a flexible silicone molding. Various types of energy transfer, such as infrared (IR), radiofrequency (RF), electromagnetic, capacitive and inductive transfer, can be used to transfer the power and / or data from external device 240 to implantable component 201.
[0046] FIG. 3 A depicts an exemplary system 210 according to an exemplary embodiment, including hearing prosthesis 100, which, in an exemplary embodiment, corresponds to cochlear implant 100 detailed above, and a portable body carried device (e.g., a portable handheld device as seen in FIG. 2A, a watch, a pocket device, etc.) 2401 in the form of a mobile computer having a display 2421. The system includes a wireless link 230 between the portable handheld device 2401 and the hearing prosthesis 100. In an embodiment, theprosthesis 100 is an implant implanted in recipient 99 (represented functionally by the dashed lines of box 100 in FIG. 3 A).
[0047] In an exemplary embodiment, the system 210 is configured such that the hearing prosthesis 100 and the portable handheld device 2401 have a symbiotic relationship. In an exemplary embodiment, the symbiotic relationship is the ability to display data relating to, and, in at least some instances, the ability to control, one or more functionalities of the hearing prosthesis 100. In an exemplary embodiment, this can be achieved via the ability of the handheld device 2401 to receive data from the hearing prosthesis 100 via the wireless link 230 (although in other exemplary embodiments, other types of links, such as by way of example, a wired link, can be utilized). As will also be detailed below, this can be achieved via communication with a geographically remote device in communication with the hearing prosthesis 100 and / or the portable handheld device 2401 via link, such as by way of example only and not by way of limitation, an Internet connection or a cell phone connection. In some such exemplary embodiments, the system 210 can further include the geographically remote apparatus as well. Again, additional examples of this will be described in greater detail below.
[0048] As noted above, in an exemplary embodiment, the portable handheld device 2401 comprises a mobile computer and a display 2421. In an exemplary embodiment, the display 2421 is a touchscreen display. In an exemplary embodiment, the portable handheld device 2401 also has the functionality of a portable cellular telephone. In this regard, device 2401 can be, by way of example only and not by way of limitation, a smart phone, as that phrase is utilized generically. That is, in an exemplary embodiment, portable handheld device 2401 comprises a smart phone, again as that term is utilized generically.
[0049] It is noted that in some other embodiments, the device 2401 need not be a computer device, etc. It can be a lower tech recorder, or any device that can enable the teachings herein.
[0050] The phrase “mobile computer” entails a device configured to enable human-computer interaction, where the computer is expected to be transported away from a stationary location during normal use. Again, in an exemplary embodiment, the portable handheld device 2401 is a smart phone as that term is generically utilized. However, in other embodiments, less sophisticated (or more sophisticated) mobile computing devices can be utilized to implement the teachings detailed herein and / or variations thereof. Any device, system, and / or methodthat can enable the teachings detailed herein and / or variations thereof to be practiced can be utilized in at least some embodiments. (As will be detailed below, in some instances, device 2401 is not a mobile computer, but instead a remote device (remote from the hearing prosthesis 100. Some of these embodiments will be described below).)
[0051] In an exemplary embodiment, the portable handheld device 2401 is configured to receive data from a hearing prosthesis and present an interface display on the display from among a plurality of different interface displays based on the received data. Exemplary embodiments will sometimes be described in terms of data received from the hearing prosthesis 100. However, it is noted that any disclosure that is also applicable to data sent to the hearing prosthesis from the handheld device 2401 is also encompassed by such disclosure, unless otherwise specified or otherwise incompatible with the pertinent technology (and vice versa).
[0052] It is noted that in some embodiments, the system 210 is configured such that cochlear implant 100 and the portable device 2401 have a relationship. By way of example only and not by way of limitation, in an exemplary embodiment, the relationship is the ability of the device 2401 to serve as a remote microphone for the prosthesis 100 via the wireless link 230. Thus, device 2401 can be a remote mic. That said, in an alternate embodiment, the device 2401 is a stand-alone recording / sound capture device.
[0053] It is noted that in at least some exemplary embodiments, the device 2401 corresponds to an Apple Watch™ Series 1 or Series 2, as is available in the United States of America for commercial purchase as of January 10, 2021. In an exemplary embodiment, the device 2401 corresponds to a Samsung Galaxy Gear™ Gear 2, as is available in the United States of America for commercial purchase as of January 10, 2021. The device is programmed and configured to communicate with the prosthesis and / or to function to enable the teachings detailed herein.
[0054] In an exemplary embodiment, a telecommunication infrastructure can be in communication with the hearing prosthesis 100 and / or the device 2401. By way of example only and not by way of limitation, a telecoil 2491 or some other communication system (Bluetooth, etc.) is used to communicate with the prosthesis and / or the remote device. FIG. 3B depicts an exemplary quasi -functional schematic depicting communication between an external communication system 2491 (e.g., a telecoil, or Bluetooth transceiver), and the hearing prosthesis 100 and / or the handheld device 2401 by way of links 277 and 279,respectively (note that FIG. 3B depicts two-way communication between the hearing prosthesis 100 and the external audio source 2491, and between the handheld device and the external audio source 2491 - in alternate embodiments, the communication is only one way (e.g., from the external audio source 2491 to the respective device)). It is noted that unless otherwise noted, the embodiment of FIG. 3B is applicable to any body worn medical device / implanted device disclosed herein in some embodiments.
[0055] FIG. 3C depicts an exemplary external component 1440. External component 1440 can correspond to external component 142 of the system 10 (it can also represent other body worn devices herein / devices that are used with implanted portions). As can be seen, external component 1440 includes a behind-the-ear (BTE) device 1426 which is connected via cable 1472 to an exemplary headpiece 1478 including an external inductance coil 1458EX, corresponding to the external coil of figure 1. As illustrated, the external component 1440 comprises the headpiece 1478 that includes the coil 1458EX and a magnet 1442. This magnet 1442 interacts with the implanted magnet (or implanted magnetic material) of the implantable component to hold the headpiece 1478 against the skin of the recipient. In an exemplary embodiment, the external component 1440 is configured to transmit and / or receive magnetic data and / or transmit power transcutaneously via coil 1458EX to the implantable component, which includes an inductance coil. The coil 1458X is electrically coupled to BTE device 1426 via cable 1472. BTE device 1426 may include, for example, at least some of the components of the external devices / components described herein.
[0056] FIG. 4 presents an exemplary embodiment of a neural prosthesis in general, and a retinal prosthesis and an environment of use thereof, in particular, the components of which can be used in whole or in part, in some of the teachings herein. In some embodiments of a retinal prosthesis, a retinal prosthesis sensor-stimulator 10801 is positioned proximate the retina 11001. In an exemplary embodiment, photons entering the eye are absorbed by a microelectronic array of the sensor-stimulator 10801 that is hybridized to a glass piece 11201 containing, for example, an embedded array of microwires. The glass can have a curved surface that conforms to the inner radius of the retina. The sensor-stimulator 108 can include a microelectronic imaging device that can be made of thin silicon containing integrated circuitry that convert the incident photons to an electronic charge.
[0057] An image processor 10201 is in signal communication with the sensor-stimulator 10801 via cable 10401 which extends through surgical incision 00601 through the eye wall(although in other embodiments, the image processor 10201 is in wireless communication with the sensor-stimulator 10801). The image processor 10201 processes the input into the sensor-stimulator 10801 and provides control signals back to the sensor-stimulator 10801 so the device can provide processed output to the optic nerve. That said, in an alternate embodiment, the processing is executed by a component proximate with or integrated with the sensor-stimulator 10801. The electric charge resulting from the conversion of the incident photons is converted to a proportional amount of electronic current which is input to a nearby retinal cell layer. The cells fire and a signal is sent to the optic nerve, thus inducing a sight perception.
[0058] The retinal prosthesis can include an external device disposed in a Behind-The-Ear (BTE) unit or in a pair of eyeglasses, or any other type of component that can have utilitarian value. The retinal prosthesis can include an external light / image capture device (e.g., located in / on a BTE device or a pair of glasses, etc.), while, as noted above, in some embodiments, the sensor-stimulator 10801 captures light / images, which sensor-stimulator is implanted in the recipient.
[0059] In the interests of compact disclosure, any disclosure herein of a microphone or sound capture device corresponds to an analogous disclosure of a light / image capture device, such as a charge-coupled device. Corollary to this is that any disclosure herein of a stimulator unit which generates electrical stimulation signals or otherwise imparts energy to tissue to evoke a hearing percept corresponds to an analogous disclosure of a stimulator device for a retinal prosthesis. Any disclosure herein of a sound processor or processing of captured sounds or the like corresponds to an analogous disclosure of a light processor / image processor that has analogous functionality for a retinal prosthesis, and the processing of captured images in an analogous manner. Indeed, any disclosure herein of a device for a hearing prosthesis corresponds to a disclosure of a device for a retinal prosthesis having analogous functionality for a retinal prosthesis. Any disclosure herein of fitting a hearing prosthesis corresponds to a disclosure of fitting a retinal prosthesis using analogous actions. Any disclosure herein of a method of using or operating or otherwise working with a hearing prosthesis herein corresponds to a disclosure of using or operating or otherwise working with a retinal prosthesis in an analogous manner.
[0060] At least some exemplary embodiments according to the teachings detailed herein utilize advanced learning signal processing techniques, which are able to be trained or otherwise are trained to detect higher order, and / or non-linear statistical properties of signals,all by way of example. An exemplary signal processing technique is the so called deep neural network (DNN). At least some exemplary embodiments utilize a DNN (or any other advanced learning signal processing technique) to process a signal representative of captured sound, which processed signal is utilized to evoke a hearing percept. At least some exemplary embodiments entail training data analysis algorithms / developing models to detect subtle and / or not-so-subtle changes, and provide an estimate of future statuses and conditions, etc., or otherwise provide a control regime to process data based on the analysis, or at least identify specific information thereabout. That is, some exemplary methods utilize learning algorithms such as DNNs or any other algorithm that can have utilitarian value where that would otherwise enable the teachings detailed herein to analyze sound and / or light based data and implement a processing strategy and / or develop a processing strategy.
[0061] Thus, at least some exemplary embodiments entail training signal processing algorithms to process signals indicative of captured sound. That is, some exemplary methods utilize learning algorithms or regimes or systems such as DNNs or any other system that can have utilitarian value where such would otherwise enable the teachings detailed herein to analyze captured sound. It is noted that the aforementioned discussion focuses on sound. However, the teachings detailed herein can also be applicable to captured light (or other phenomena that can be captured. In this regard, the teachings detailed herein can be utilized to analyze or otherwise process a signal that is based on captured light, and evoke a sensory percept, such as a vision percept, based on the processed signal. Note also that embodiments are not necessarily limited to captured phenomena. In an embodiment, an audio and / or visual signal can be streamed or otherwise provided to one or more exemplary systems and the system can process that data. More on this below.
[0062] A “neural network” is a specific type of machine learning system. Any disclosure herein of the species “neural network” constitutes a disclosure of the genus of a “machine learning system.” Moreover, any disclosure herein of the species “machine learning” constitutes a disclosure of the genus of “artificial intelligence.” While embodiments herein focus on the species of a neural network, it is noted that other embodiments can utilize other species of machine learning systems accordingly, or the broader genuses noted, any disclosure herein of a neural network constitutes a disclosure of any other species of machine learning system that can enable the teachings detailed herein and variations thereof. To be clear, at least some embodiments according to the teachings detailed herein are embodiments that have the ability to learn without being explicitly programmed. Accordingly, with respectto some embodiments, any disclosure herein of a device or system constitutes a disclosure of a device and / or system that has the ability to learn without being explicitly programmed, and any disclosure of a method constitutes actions that results in learning without being explicitly programmed for such.
[0063] Some of the specifics of the DNN utilized in some embodiments will be described below, including some exemplary processes to train such DNN. First, however, some of the exemplary methods of utilizing such a DNN (or any other system that can have utilitarian value) will be described.
[0064] It is noted that in at least some exemplary embodiments, the DNN or the product from machine learning, etc., or the results thereof, etc., is utilized to achieve a given functionality as detailed herein. In some instances, for purposes of textual economy, there will be disclosure of a device and / or a system that executes an action or the like, and in some instances structure that results in that action or enables the action to be executed. Any method action detailed herein or any functionality detailed herein or any structure that has functionality as disclosed herein corresponds to a disclosure in an alternate embodiment of a DNN or product or results from machine learning, etc., that when used, results in that functionality, unless otherwise noted or unless the art does not enable such.
[0065] Embodiments thus can include devices, systems, and methods that can provide sound and / or light processing, and can utilize artificial intelligence to implement at least some of these teachings herein. Any product of machine learning disclosed herein corresponds to an alternate disclosure of a product of artificial intelligence and vis-a-versa, unless otherwise noted.
[0066] FIG. 5 depicts an exemplary conceptual functional black box schematic associated with an embodiment, where a data 510 is the input into a DNN based device 520 that utilizes a trained DNN or some other trained learning algorithm or trained learning system (or the results thereof -a “product of machine learning” as used herein covers a trained learning algorithm or trained learning system as used in operational mode after training has ceased and covers a product that is developed as a result of training (e.g., a chip or firmware or software - again, this will be described in greater detail below, but briefly, much attention is paid to a convolutional encoder-decoder, and if an algorithm results therefrom, it is a product of a convolutional encoder-decoder)), and outputs data 530. In this embodiment, data 510 could be based on sound data and / or light data (audio and / or visual data), which means that it couldbe pre-processed or could be a raw signal (e.g., a raw sound signal, which could be analogue, where the DNN based device converts the analogue signal to digital data). The data 510 could be a digital signal. In an embodiment, there can be an input subsystem that provides the data 510 to the DNN based device 520. This input subsystem could convert analogue sound signals for example to digital signals. The input subsystem could also simply be the conduit by which the data based on audio and / or light data is provided to device 520.
[0067] By “data based on X data,” this can be the raw signal / raw data, or can be data that is developed from the raw signal, or data developed from data that was developed from the raw signa, and so on. That is, it includes the base / original data and modifications thereof.
[0068] In this exemplary embodiment, the output 530 contains an electrode activation sequence. This is a sequence that can be provided to a device, such as a stimulation device of an assembly that evokes a sensory percept using electrical stimulation from electrodes (hence an electrode activation sequence). FIG. 27 shows by way of example a graphic conceptually presenting electrode activation sequences (not exhaustive, for example, as noted above, phase information, etc., can also be included) for a 22 channel cochlear implant, where the implant will output current based on those sequences (height of the lines corresponds to amplitude of the current when activated, timing is the time that that “pulse” will be active). This is a simplified sequence for purposes of explanation and discussion and other embodiments can have different output. The point is that when referring to electrode activation sequence data, etc., this is the kind of data that such covers. Note that this can be in digital and / or analogue format. For example, the amplitude, timing, etc., can be based on digital data, or can result from a varying analogue signal. Moreover, the data may be instantaneous and transitory. In an embodiment, the output is a digital signal that controls the electrode driver or can be used by a device that controls the electrode drive at that time. The dataset is for that stimulation on that channel and no more, for example. There could be no temporal data in the output, because as soon as it is outputted, it is used. Put another way, electrode activation sequencing is analogous to the output of an electronic ignition that controls sparkplugs. It can be temporary and transitory, although can be recorded for future analysis. Conversely, a dataset could include the sequences for each channel, for a “stimulation block,” which block covers the stimulation from all channels, which can include no from one or more channels, where the cochlear implant is configured to run through all the channels before again stimulating on a prior channel (where there could be a channel skipped if there is no stimulation for that channel). And if, for example, each stimulation from an electrodechannel can take up a 2 ms sub block, each stimulation block would be 44 ms in length, and that could be a gap between each block. The point is that in this exemplary embodiment, stimulation channel 15 / the electrode 15 will not receive current again until at least 44 ms later, which is the length of the stimulation block. Accordingly, the electrode activation sequence based data could be the block or it could be a sub block or could be a plurality of blocks.
[0069] Embodiments thus include a machine learning device that “learns” how to encode or otherwise develop the code for electrode activation sequencing, or otherwise develop data that is directly usable by a controller that controls an electrode driver / or the electrode driver itself. This as opposed to, for example, learning how to process data to develop data that remains in the audio domain and then converting to the electrode domain / electrical domain, utilizing hardcoding or otherwise fixed or hard hardware or firmware or software (and typically, that is readily identifiable as such). Put another way, in an exemplary embodiment, there is a product of machine learning that receives data that is in the audio domain and processes such to output data in the electrical or the electrode domain.
[0070] In this exemplary embodiment, device 520 can have the functionality of a sound processor or a light processor of a hearing or vision prostheses, respectively, and can have the addition of coding (and other functionalities as will be described below).
[0071] It is noted that in at least some exemplary embodiments, the input 510 comes directly from a microphone, while in other embodiments, this is not the case. Input 510 can correspond to any input that can enable the teachings detailed herein to be practiced providing that the art enables such. Thus, in some embodiments, there is no “raw sound” input into the DNN device. Instead, it is all pre-processed data. Any data that can enable the DNN device or other machine learning algorithm or system / product thereof to operate can be utilized in at least some exemplary embodiments.
[0072] By way of comparison, again keeping with the opening theme that the teachings herein will often be described in the context of a cochlear implant, conventional cochlear implant (CI) sound processing techniques utilize an input filter-bank, which converts incoming audio signals to the time-frequency domain. Further signal processing such as by way of example, speech enhancement, noise reduction, auditory scene classification, automatic gain control, spatial beamforming, binaural processing, and finally conversion to stimulation pulse trains for the implanted electrodes, is all done in this domain. In some CCI(conventional CI), Fast Fourier Transform (FFT)-based filter-banks decouple magnitude and phase information, and in the case of CI processing, the phase information is not returned or otherwise the output signal avoids having phase content. Conversely, in some embodiments, the DNN devices executes processing of the input sound signal or otherwise data based on the input sound signal without implementing FFT. In an embodiment, the phase information is not decoupled / never decoupled (in the processing) from magnitude and vice versa. In an embodiment, the output signal includes phase content / includes content where magnitude and phase remain coupled, or otherwise the output is based on data where the magnitude and phase remain coupled and is otherwise not decoupled. This is contrasted to the CCI arrangement. In conventional sound processing arrangements for cochlear implants, all beamforming is executed in the input side by way of example, where there is magnitude and phase coupling. Often, noise cancellation, or at least speech enhancement, is executed on the input side as well. Here, beamforming and noise cancellation for example rely upon the magnitude and phase coupling, which ceases to exist after filter binning / envelope development. In an embodiment, spatial features that are relied upon to provide sound enhancement, such as speech enhancement, also rely on the phase and magnitude coupling, and thus utilize the signal that exists before the filter binning.
[0073] In an exemplary embodiment in accordance with the teachings detailed herein, as opposed to the conventional arrangement, the loss function is in the domain of the electrode sequencing. The method by which the loss is calculated can be a difference metric between the electrode activation sequence at the output of the proposed invention, and a target electrode activation pattern for a particular input. By restricting the domain that the loss function is calculated within to the domain that is delivered via electrical stimulation, as is done in some embodiments, the available computational resources within the proposed invention are reserved only for information relevant to stimulation. As will be detailed below, in an exemplary embodiment, the noise reduction and / or speech enhancement and / or spatial processing or identification of primacy functionalities herein are executed within the electrode sequencing domain as opposed to the channel envelope domain or the frequency bin domain or even the domains upstream of such, such as the broader audio domain (in a sense, the channel envelope domain is a subset or a species of the genus of the audio domain - the electrode sequencing domain is a separate genus from the audio genus).
[0074] Embodiments can include a processing block, such as that which can be achieved by the product of machine learning detailed herein, such as the DNN detailed herein by way ofexample, where there is no hardcoded arrangement. That is, in an exemplary embodiment, the processing can be executed without a hardcoded input filter-bank, at least with respect to the signal processing that functionally would correspond to the above-noted FFT based processing. Instead, embodiments can utilize a data driven approach. In an exemplary embodiment, the product of machine learning implements a learned optimal filter-bank, or more accurately, seeks to duplicate the results of the optimal filter-bank (in part - as will be detailed below, the product does more than this), which can be learned by a neural network or otherwise some form of artificial intelligence system. In an exemplary embodiment, this “learned filter-bank” can be in the form of one or more ID or 2D convolution layers. These layers can be in some embodiments, encoder or the “convolutional encoder.” In some embodiments, the output of the convolutional encoder, or the “latent space” is then processed by other neural network layers, which can comprise the “bottleneck layers” of the neural network architecture. These bottleneck layers could be any form of neural network layer, providing that such enables the teachings herein. By way of example, fully connected, recurrent, gated recurrent units (GRU), Long Short-Term Memory units (LSTM), convolutional layers, attention mechanisms, etc. In some embodiments, a convolutional “decoder” can be used, which is typically comprises ID or 2D transposed convolutional layers. This is done to output a signal that is in a dimension different from that of the input dimension (and thus not returned back to the input dimension(s)), for example. Herein, this neural network structure is often referred to as a convolutional encoder-decoder (CED), a U- Net and a convolutional autoencoder. Any reference herein to a deep neural network or machine learning or an artificial intelligence system corresponds to a reference to any one or more of these.
[0075] Embodiments can use the CED to capture both local and global features (more on this below). In some embodiments, this is used to have the network extract hierarchical features on a variety of time-scales. Additionally, in some embodiments, this is used to efficiently down-sample the input signal or features to a learned latent representation, expanding the temporal context in the feature space and reducing complexity in the bottleneck layers. This learned feature space can be used to enable the network to self-determine a meaningful feature space. This as opposed to the programmer pre-determining a set of input features. Thus, in some embodiments, at least some processing (the processing done by the product of machine learning for example) is done without predetermined input features. In someembodiments, the CED structures used herein share parameters across their inputs. This can reduce the memory footprint of the pertinent networks.
[0076] Embodiments can thus include an encoder-decoder (ED) structure that has a decoder that instead of returning the latent space representation back to the input domain, transforms the latent space representation into a domain different from the input domain. In the case of audio signal processing, input audio is transformed into the latent space by the convolutional encoder, bottleneck layers process the latent space, and then the convolutional decoder transforms the latent space not back to audio, but instead into electrical pulse data / electrode activation sequence data (this can be done at the decoder). That is, embodiments include converting audio input / audio data to electrode activation sequences (or data directly related thereto - more on this below) at the decoder. This as distinguished from, for example, converting the latent space to channel envelopes, such as the channel envelopes of a cochlear implant, where in the context of a CI, post processing would be needed to process the channel envelopes to generate electrode activation sequences / data directly for such in a separate process outside the CED. The teachings herein can avoid this post processing (at least some of the post processing), by having the CED do this in one fell swoop in some embodiments.
[0077] Embodiments include convolutional encoder-decoder neural networks specifically designed for sensory prosthesis applications, such as CI and retinal implant / bionic eye (BE) applications. In some embodiments, the audio domain is transformed to the latent space using a convolutional encoder, bottleneck layers process the latent space, and then the latent space is transformed to the electrode space (e.g., CI or BE space). In an embodiment, the dimension of the output corresponds to the number of active implanted electrodes (N), which can be by way of example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 65, 70, 75, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400 or 450 or more or any value or range of values therebetween in 1 increment (e.g., 83, 303, 10 to 31, etc.). Note that there could be additional electrodes, such as the return electrode(s), which would not be given a channel in some embodiments, and thus the dimensions of the output would be lower than the total number of electrodes (because not all are stimulating electrodes).
[0078] FIG. 6 shows an exemplary block diagram of a convolutional encoder-decoder tailored specifically for electrode output devices, such as cochlear implants or bionic eyes. Here, this diagram is for a cochlear implant, where the input is raw audio (but need no be so -any data that is based on an audio signal can be used providing that the art enables such if such can have utilitarian value), and the output is electrode activation sequences (by way of example - other data can be the output - here, any data that is in the “electrode space” or “electrical output space” can be the output). That is, in an embodiment, electrode activation sequences are generated directly from an audio input using a convolutional encoder-decoder. In an embodiment, one or more or all of the amplitude compression, channel envelope sampling, maxima selection, acoustic-to-electric mapping, and further signal processing such as noise reduction, speech enhancement, binaural processing, and microphone equalization can be encapsulated within one deep learning framework (such as a product of machine learning / a product of artificial intelligence).
[0079] In an embodiment, there is a system, such as system 701 shown in FIG. 7. The system 701 includes an input subsystem 705 configured to receive input, concomitant with the above. This can be a microphone or a camera. This can be a device that converts analogue input to digital output (an AD converter for example). This can be a transceiver or a receiver that receives an electromagnetic signal (a fibre optic signal is an EM signal). Other configurations of input subsystem 705 can be used providing that the art enables such, unless otherwise noted. The output of the input subsystem 705, output 710 (which can correspond to input 510 discussed above and / or the X input features or raw audio input of FIG. 6 discussed above), is the input into the convolutional based subsystem 720. (Herein, the subsystem is often referred to as the convolutional subsystem. Any reference to one corresponds to a disclosure of the other in the interests of textual economy. A convolutional subsystem is just that, whereas as convolutional based subsystem is a system that may not be a convolution system, but is based on the results thereof (e.g., a chip that has an algorithm thereon that is the result of a trained CED). Thus, the convolutional based subsystem 720 is in signal communication with the input subsystem 705. (Note that subsystems can be part of a single assembly, such as a BTE device of a cochlear implant, and can be supported on a common PCB for example, or tiered PCBs that are in signal communication with each other, etc., or separate PCBs that are in signal communication with each other, etc.) The output 710 can be a raw audio signal such as that provided by a microphone, whether or not subjected to analog-to-digital converting, or could be the output from a streamed audio signal, such as by way of example only and not by way of limitation, from the wired output of an Internet radio. In an exemplary embodiment, the input subsystem could be a radio receiver or a transceiver, which converts a received radio wave into an audio signal. This is all by way of example.
[0080] The convolutional based subsystem 720 can correspond to any of those detailed above and below. In an exemplary embodiment, this is a CED (and as noted herein, any disclosure of a CED corresponds to an alternate disclosure of another type of machine learning or otherwise a product of machine learning providing that the art enables such, unless otherwise noted, and vice versa all in the interest of textual economy). Some additional specific teachings of this will be provided in greater detail below, but briefly, in an exemplary embodiment, the convolutional based subsystem can be the results of or otherwise can correspond to a trained neural network having this architecture.
[0081] The output of the convolutional based subsystem is output 730, which can correspond to output 530 above and / or the output N electrode activation sequences of FIG. 6 noted above.
[0082] Accordingly in an embodiment, there is a system, such as system 701, which includes an input subsystem, such as subsystem 705, configured to receive input, and a convolutional based sub-system, subsystem 720, in signal communication with the input subsystem. In this exemplary embodiment, consistent with the teachings above, the convolutional based subsystem outputs electrode activation sequence based data. This can be an actual electrode activation sequence, or can be a dataset that is usable by an electrode driver apparatus, such as an electrode array driver of a CI, or an electrode driver of a BE (e.g., a digital dataset that can be read by a digital portion of an electrode array driver, where the electrode array driver drives the electrode array based on that digital dataset without further processing) - the sequence based data can be zeros and ones only understandable by the electrode driver apparatus, and could also be distinct and identifiable current pulses. And in this regard, by way of example, FIG. 8 shows an exemplary system 801 that is a tissue stimulation system (e.g., a cochlear implant) that includes the subsystems and the communication regime of the embodiment of FIG. 7, and also an electrode driver apparatus 840, such as an electrode array driver of a cochlear implant, or an apparatus that includes a separate controller that translates the input (as opposed to converting to another domain, such as from the audio domain to the electrical domain) and controls the driver accordingly, which driver 840 outputs electrical signals 850 to respective electrodes (N electrodes).
[0083] In an embodiment, the electrode activation sequence based data (which includes electrode activation sequence(s)) has data that corresponds to or is directly translatable to an amplitude of a current to be applied to an electrode. In an embodiment, the electrode activation sequence based data has data that corresponds to or is directly translatable to aninitial activation timing of a current to be applied to an electrode. In an embodiment, the electrode activation sequence based data has data that corresponds to or is directly translatable to length of time of current application to be applied to an electrode. In an embodiment, the electrode activation sequence based data has data that corresponds to or is directly translatable to a relative activation time relative to one or more other electrodes. In an embodiment, electrode activation sequence based data (which can be electrode activation sequences) can include active electrode designation (electrode number), return electrode designation (electrode number), pulse amplitude, phase duration of activation (uniform our could be per channel), interphase gap of activation (uniform or could be per channel), pulse shape (uniform our could be per channel), interpulse interval (uniform our could be per channel), etc. All of these can be specified in the electrode activation sequence. Note that in some strategies, some of these, such as for example, phase duration and interphase gap are fixed, and therefore not determined by the CED. In some embodiments, in these instances, the "sequences" that are output by the CED can be combined with the fixed phase duration and interphase gap to comprise a complete electrode activation sequence (or not, the driver could be “hard controlled” to have the fixed data). In other instances, the CED could output the entire electrode activation sequence. This can be utilitarian in embodiments where there is desire to change the pulse shape or the degree of focusing for different pulses.
[0084] In an embodiment, the electrode activation sequence based data has no data that corresponds to or is directly translatable to frequency content. In an embodiment, the electrode activation sequence based data has no data that corresponds to or is directly translatable to channel envelope (or frequency envelope) content. This is not to say that such data could not be extracted there from if the algorithm is known or otherwise reverse engineering is undertaken. But that is not translation. One would have to know what channels correspond to what frequencies. Granted, it might be that the channel envelope could be estimated from the electrode activation sequence based data, or more accurately, that channel envelopes that would be akin to a finite element analysis could be obtained from the electrode activation sequence based data (there would be gaps between timing for a given channel, and the space between the gaps would be filled in utilizing a linear approximation or some other curve fitting approximation - that is not translation - that is the creation of a model that approximates a channel envelope).
[0085] In an exemplary embodiment, the electrode activation sequences are sequences that are directly readable by the electrode driver or a control unit that controls the electrode driver.Another way of looking at this is that such a control unit or such an electrode driver cannot operate directly from the channel envelopes and / or the frequency bin outputs or otherwise data based directly thereon. That is, the electrode activation sequence based data can be utilized directly by the electrode driver or the control unit thereof. It is all that is needed by such components to apply the current at the given amplitude and / or at for the given length of time so as to evoke a hearing percept utilizing electrical stimulation of the cochlea based on the captured sound.
[0086] Note that any given data of the sequence based data can be outputted serially for each channel or in a block that has two or more or all channel data, or any of the N channel numbers herein. Conceivably, the output could have multiple amplitudes for a given channel to be outputted serially and the system could interleave this with other current applications for other channels. The output could have two or three or four or five or more activations for any one or more of the channels.
[0087] Briefly, it is noted that in some embodiments, the output of the CED is output that has been adjusted or otherwise is output that takes into account threshold and / or comfort levels of the recipient. Conversely, in an exemplary embodiment, there is a downstream component downstream from the CED that takes the output and adjusts for threshold and / or comfort levels of the recipient. In an exemplary embodiment, this downstream component is the downstream component that is controlled based on fitting data sets that are loaded into the hearing prostheses as a result of a so-called fitting procedure where the prostheses is fitted to the recipient. In an exemplary embodiment, the output of the CED is normalized on a scale of 0 to 1, or more accurately, the amplitude portion of the electrode activation sequence data is normalized on a scale of 0 to 1. The component downstream of the CED can then take the normalized amplitude for one or more or all channels and adjusts such to accommodate the threshold levels and the comfort levels for those channels for that particular recipient. By way of example only and not by way of limitation, in an exemplary embodiment, the difference between the threshold level and the comfort level would be a certain value. The output of the CED, or more accurately, the amplitude data in the output would be multiplied by the difference between the threshold and comfort level or a value corresponding to such that can have utilitarian value with respect to the teachings detailed herein, and then the threshold level is added to that value. This will always result in a current level that is at or between the threshold level and comfort level. This is but an exemplary embodiment in other ways of implementing the results of fitting the cochlear implant to the recipient in thestimulation strategy for the cochlear implant can be implemented providing that the art enables such unless otherwise noted herein.
[0088] Returning back to the system of FIG. 7, in an embodiment, the system 701 is an ambient environment stimulus processor system, such as that of a cochlear implant or that of a retinal implant / BI, by way of example, where there is a sound capture device or a light capture device that is part of the system or is in signal communication therewith, where the output from these capture devices is the input into subsystem 720 or subsystem 705. There is thus a stimulus capture device, which can correspond to an image sensor, such as a digital image sensor or a sound sensor (e.g., CCD, CMOS, microphone (conventional or electret, etc.)) Any transducer that can enable the teachings herein can be used providing that the art enables such.
[0089] In an embodiment, the system does not necessarily “react” to an ambient environment. Instead of providing stimulus based on the ambient environment, the system provides stimulus based on an audio signal that can be a streamed signal that is the output from an MP4 player for example or Internet radio or from a television, etc., or can be based on a video signal that is streamed from a television or an internet video device, etc.
[0090] Figure 9 presents an exemplary system based in part on the above teachings. As seen, microphone 902, which can correspond to the microphone of the hearing prosthesis presented in figure 1, can execute the action of capturing sound. Upstream processing component 906 converts the analog signal captured by the microphone into a digital signal. Thus, collectively, microphone 902 and converter 906 can execute the action of capturing sound and pre-processing the signal resulting from sound capture (note there can be an amplifier upstream of the CED). This is often referred to as front end processing. The DNN / CED does not do this in at least some embodiments, and does so in others. Unlike a conventional sound processing arrangement, there is no audio filter (or at least no tangible filter or identifiable structural filter) that breaks up the now digitized signal into different filter bins or otherwise establishes different signals for respective frequencies. That said, if there is filtering, it is not based on Fourier Transform. In an embodiment, there is a correlation between magnitude and phase with respect to the data provided to CED 720. In an embodiment, what is fed to the CED 720 is not broken out by frequency amplitudes from a filter, or at least not a filter based on FFT principles.
[0091] The output from ADC 906, or from any intervening component of the system that meets the above caveats, is fed into the neural network inputs so as to be obtained by device 720 (the CED in this example). In at least some exemplary embodiments, the network of 720 will have already been loaded with pre-taught weights (more on this below). The neural network of device 720 then determines the electrode channel energies that can be utilitarianly stimulated based on the audio input, and also develops recipient threshold / comfort level mapping in some embodiments, and not in others. This as opposed to providing electrode channel energies to a separate device outside the DNN / artificial intelligence based device. Instead, the CED produces electrode amplitude data and timing data (and other data in some embodiments), and otherwise produces electrode activation sequence data, which is provided to electrode driver 840, which converts such to electrode currents therein (and then outputted to a lead assembly extending to the electrode array implanted in a cochlea to stimulate tissue with electrical current and evoke a hearing percept in a tonotopically linked manner).
[0092] In an exemplary embodiment, device 720 is a microprocessor or computer chip or otherwise a system that includes the product from / based on the machine learning. In an exemplary embodiment, there is no circuitry outside device 720 (or at least downstream therefrom) that has logic circuits that takes the output from the device 720 and applies weights to the outputs associated with the recipient threshold and / or comfort levels. Again, instead, this is done by device 720. Device 720 can execute the functionality of a mapping section processor of a cochlear implant. Thus, the device 720 is a DNN map section (amongst other things). That said, in an embodiment, device / subsystem 720 does everything thus described but does not do the final adjustment to take into account threshold and comfort levels. This can be done by a section downstream of device 720. That is, by way of example, the electrode activation sequences can be adjusted prior to being utilized to provide stimulation to the recipient to evoke a hearing percept. Put another way, the output of device 720 can be utilized as is, but the output might be below the threshold level and / or might be above the comfort level, and would not necessarily be adjusted to accommodate those two levels. That would be the only difference in this exemplary embodiment relative to that which is been described above.
[0093] Indeed, in an exemplary embodiment, a cochlear implant can be obtained, and device 720 can be inserted in between the ADC 906 and the electrode driver (or controller therefor). In an embodiment, the conventional audio filter and the mapping section thereof can be replaced / is replaced. In an embodiment, the section can be bypassed. Here, in anembodiment, there is a cochlear implant that has two signal paths - one being the conventional path, and the other being the CED path, as seen by way of example in FIG, 10, where there are two signal processing paths, one based on the CED, and the other using the traditional Cochlear Implant signal processing path (simplified for the purposes of discussion - additional details are provided below by way of example). Some additional details of both are described below. In an embodiment, the system can switch paths from one to the other. Also, in an embodiment, the output from the conventional path (bottom) can be used to “train” the CED (e.g., by way of example, when the input into microphone 902 has low noise / is of good quality, or such as where microphone 902 is the external microphone of a totally implantable hearing prosthesis, and the implantable microphone also captures sound and the signals are compared / output is compared to train the CED). More on this below. It is noted that in some embodiments, the bottom processing path (the conventional path) can be part of a single processor. In an embodiment, a processor could be modified to include the features associated with device 720, or otherwise can include a separate processor that communicates with the processor of the cochlear implant / the cochlear implant sound processor, to execute the actions associated with device 720. (It is noted that in an alternate embodiment, processor 720 is replaced with a non-processing device, or includes non-processing devices, such as a chip or the like that is a result of a machine learning algorithm or machine learning system, etc. Any disclosure herein of a processor corresponds to a disclosure in an embodiment of a non-processor device or a combined processor-non-processor device where the non-processor is a result of machine learning.)
[0094] Figure 9A shows an exemplary embodiment of a system that has a separate component that adjusts the amplitudes of the output from the CED to accommodate the fitting settings or otherwise to implement the fitting settings and thus evoke hearing percepts that are within the constraints of the threshold and comfort levels of the implant. In this exemplary embodiment, there is still the channel amplitudes and timing output 730 output from the CED 720. But here, the CED has not taken into account the threshold and comfort levels or otherwise the fitting settings for the recipient. Thus, the electrode activation sequences or the normalized electrode activation sequences which can have an amplitude ranging from 0 to 1 by way of example. The data for each channel that includes an amplitude that is greater than zero, or data for all channels for that matter (in which case any action would simply result in a zero value for the channels that have no amplitude or zero amplitude) is provided to the amplitude adjustment based on fitting settings block 933, and adjusted so that the output 955which is provided to the electrode driver 840 or otherwise a controller for the electrode driver falls within the threshold and comfort levels determined during the fitting session for that recipient or otherwise the amplitudes result in a hearing percept evoked by the prostheses that corresponds to a fitted hearing percept based on the fitting settings developed in the fitting session (providing that nothing has changed of course, which can be the case after time elapses from a fitting session). In some embodiments, block 933 can be a chip or otherwise can be circuitry that takes the input amplitude for one or more channel(s) and multiplies that by the applicable to those channels and then adds thereto the current corresponding to the threshold level for there is one or more channel(s). The output thus has a value between (and inclusive) the threshold level and the comfort level and does not exceed the comfort level and does not fall below the threshold level (unless the amplitude is zero, in which case it is below the threshold level because there is no current at all). In an exemplary embodiment, the chip or circuitry of the amplitude adjustment block 933 can be the controller for the electrode driver mentioned herein. In this regard, the controller could be a device that makes the final adjustments to the electrode activation sequence based data or otherwise directs the electrode driver to drive the electrodes to evoke a hearing percept based on the output from the CED but within the threshold and comfort levels or otherwise based on the fitting regime applied to the prosthesis in accordance with the fitting settings developed during the fitting sessions to fit the prostheses to an individual recipient.
[0095] In an exemplary embodiment, block 933 also includes memory or otherwise has access to memory in which are stored the fitting settings that are utilized by the circuitry or the chip, etc., of element 933. In this regard, figure 9B presents a fitting settings storage block 966 which can be a semiconductor based memory or any other type of memory that can enable the teachings detailed herein, which memory will store the fitting settings after the prosthesis is turned off and which will render the fitting settings in memory accessible by the amplitude adjustment chip or processor or circuitry as would be utilitarian. Here, in this embodiment, the amplitude adjustment chip or processor or circuitry accesses the memory 966 and utilizes the fitting settings stored therein to determine how one or more channels will be adjusted with respect to the amplitude thereof to ensure that the electrode driver drives the electrodes to within the threshold and comfort levels. Corollary to this is the embodiment of figure 9C which utilizes memory 966 accessible by the CED 720. Here, this arrangement can apply to embodiments where the CED output addresses the threshold and comfort settings and otherwise produces an output that is fitted to the individual human. That is, the CED canautomatically factor in the fitting settings when evaluating the input data so that the output will be electrode activation sequence based data that is already “fitted” to the particular recipient. And while memory 966 is shown as a separate component from the CED, in another embodiment, the CED can be trained already to take into account the fitting setting, although such might result in the CED that needs to be retrained if the fitting settings must be changed whereas in the embodiment of figure 9C and figure 9B, adjustments can be made to the fitting settings without having to retrain the CED.
[0096] Corollary to all the above is that the fitting settings are directly attributable to the output of the CED. This is contrasted to, for example, output corresponding to data that is other than the electrode sequence based data, such as channel envelopes or filter bin data, etc.
[0097] Figures 9B and 9C show input 977 into the fitting settings block 966. This could be input from a jack receptacle of the hearing prostheses, where, an audiologist or the like attaches a male jack portion to the prostheses so that the fitting settings can be downloaded to the prostheses. Alternatively, this can be done by a recipient. Note that instead of a malefemale connection, in an alternate embodiment, a magnetic inductance connection can be utilized or some other form of wireless communication can be utilized to communicate the fitting settings to block 966 as represented by arrow 977.
[0098] Another point. Irrespective of whether or not the data output by the CED is adjusted to take into account fitting, that output is output that is readily usable by the electrode driver or by a controller that controls the electrode driver (again, some form of translation may be needed). The fact that it is utilitarian to adjust that data so that the output of the electrode driver and thus the electrodes is fitted to the recipient does not change the fact that the output from the CED is still electrode activation sequence based data. Put another way, by rough analogy, the electrical analog signal to a speaker is still an electrical analog signal usable by that speaker even if the amplitude in the signal result in the speaker outputting sound that is painfully loud or on the other side, because the speaker to output sound that is sufficiently low in volume that a person cannot hear that sound. Thus, whether or not there is an intervening circuitry portion that adjust the amplitude does not change the fact that the signal that that portion receives is electrode sequence control data. If the output from the CED can be utilized to drive the electrode driver, such qualifications are met.
[0099] Briefly, in a conventional processing system the audio filter splits sound, depending on the embodiment into multiple frequency bands. With respect to embodiments directedtowards hearing prostheses, the splitting emulates the behavior of the cochlea in a normal ear, where different locations along the length of the cochlea are sensitive to different frequencies. The envelope of each filter output controls the amplitude of the stimulation pulses delivered to a corresponding electrode. With respect to hearing prostheses, electrodes positioned at the basal end of the cochlea (closer to the middle ear) are driven by data based on the data of the high frequency bands, and electrodes at the apical end are driven by data based on the data of the low frequencies. The outputs of the audio filter are a set of signal amplitudes per electrode channel or plurality of electrode channels, where the electrode channels are respectively divided into corresponding frequency bands. This is not done in embodiments utilizing the product of machine learning (although this can be done for purposes of training). Instead, the DNN / product of machine learning / the convolutional based subsystem, does this. And there is no specific circuitry dedicated solely to doing this. Instead, the DNN / product of machine learning / the convolutional based subsystem, does this, and one of ordinary skill cannot point to any specific circuitry that does this or any specific software or firmware or otherwise hard coding dedicated solely to doing this.[ooioo] Note that some embodiments do utilize outputs of an audio filter. In an embodiment, there is a method that includes obtaining outputs from an audio filter, such as a set of signal amplitudes per channel or plurality of electrode channels, where the channels are respectively divided into corresponding frequency bands. This can be done as a result of the FFT based filters detailed herein. These outputs can be provided to the product of machine learning, and / or these outputs can be further processed, such as for example, to develop channel envelopes, which then can be provided to the product of machine learning. Accordingly, there are embodiments where there are distinct processing circuitry that do perform one or more of the processing functionalities detailed herein as opposed to simply mimicking or otherwise obtaining a result the corresponds to the functionality.[ooioi] In an embodiment, there are devices systems and / or methods that develop FFT magnitude and / or FFT phase, or other filter bank outputs such as those outputs from a normal hearing model, by way of example only and not by way limitation. In an embodiment, the output of the ACE ™ sound processing model can be provided to the product of machine learning, or otherwise can correspond to the input into the product of machine learning, or an intervening “component” thereof can be the input into the product of machine learning, and then the machine learning can execute one or more the actions detailed herein based on the input to obtain the electrode activation sequence based data or the control data in accordancewith the teachings detailed herein. And note that embodiments that obtain the phase from the processing can utilize this obtained phase establish a coupling between magnitude and phase, which coupling is lost when utilizing normal FFT filters in normal cochlear implant signal processing arrangements. Put another way, it could be that the shortcomings of traditional FFT filtering can be overcome providing at least at a phased based input is provided to the product of machine learning. Thus, embodiments include separately filtering for magnitude and phase, more accurately, developing magnitude and phase separately, and providing what is developed to the product of machine learning so that the output of machine learning is based on data that has a coupling between magnitude and phase.
[0102] The audio filter is located on the input side as that phrase is used in the art. This can entail management of the input, such as the utilization of feedback elimination algorithms where a portion of the signal from the microphone (or signal based thereon) is canceled and / or data signal cancellation, where again, a portion of the signal from the microphone (or signal based thereon) is cancelled or otherwise attenuated. In an exemplary embodiment, the aforementioned canceling can be utilized to achieve noise reduction, and therefore, such cancellation that occurs on the input side corresponds to an input side operation. Speech enhancement and beamforming / directional sound capture techniques are also input side processes. Of course, as noted above, the prefiltering also entails the management of the input. The utilization of signal / data compression, etc., so as to enable the mapping processor block and / or the sampling and selection block to perform in a more efficient manner and / or in a power conservancy mode is also included in the input side signal management (but the mapping and sampling and selection is output side - more on this in a moment). In an embodiment, the DNN / product of machine learning / the convolutional based subsystem, performs at least some front end / input side processing (or functionally achieves such even if the processing is not per se present owing to the mechanics of a DNN / product thereof - this is the relationship of the CED with the conventional path when the CED replaces part of the path).
[0103] The mapping processor block of the conventional processing path compresses the filter envelopes to determine the current level of each pulse. For the purposes of this disclosure, the mapping processor block is between the input side and the output side. Some literature refers to the map processor block as output side processing but as a technical matter, it should be considered the bridge from the input side to the output side. As used herein, it is not output side but instead an intermediate side processing block.
[0104] The sampling and selection block samples the output of the audio filter, such as the filter envelopes, and determines the timing and pattern of the stimulation on each electrode. In general terms, sampling and selection block selects certain electrode channels as a basis for stimulation, based on the amplitude and / or other factors. Still in general terms, sampling and selection block determines how stimulation will be based on the channels corresponding to the divisions established by the audio filter. In at least some exemplary embodiments, the actions of the sampling and selection block are executed by a so-called sound processor with respect to a hearing prosthesis. In an embodiment, the DNN / product of machine learning / the convolutional based subsystem performs this processing of the sampling and selection processor block (or functionally achieves such per the caveat above).
[0105] In an exemplary embodiment, the mapping processor block can be where the threshold and comfort levels are taken into account. In this regard, the mapping processor block can adjust the amplitudes of what would be the electrical pulses if not for the fitting settings or otherwise the adjustment to take into account threshold and comfort levels. Alternatively, in an exemplary embodiment, this can be done after the sampling and selection processor block described below. Indeed, this can be part of the electrode control processor block also as will be described below.
[0106] Thus, it can be seen that in an embodiment, the DNN / product of machine learning / the convolutional based subsystem, performs at least some back end / output side processing (or functionally achieves such even if the processing is not per se present owing to the mechanics of a DNN / product thereof - this is the relationship of the CED with the conventional path when the CED replaces part of the path). Also, it can be seen that in an embodiment, the DNN / product of machine learning / the convolutional based subsystem, performs at least some input side processing (or functionally achieves such even if the processing is not per se present owing to the mechanics of a DNN / product thereof - this is the relationship of the CED with the conventional path when the CED replaces part of the path),
[0107] In an exemplary embodiment, there is a chip or a processor that is programmed and configured or otherwise contains code or circuitry or switches, etc., to execute one or more of the functionalities detailed herein associated with element 720. In an embodiment, there is no circuitry solely dedicated to electrode activation sequencing.
[0108] Some additional features of the device 720 are described below as utilized in a cochlear implant, but for now, it is noted that at least some embodiments can include methods, devices, and / or systems that utilize a DNN device or otherwise a product of machine learning inside a cochlear implant system and / or along with such a system for the analysis of sound and / or the generation of electrical stimulation patterns to be applied inside the cochlea. Any disclosure herein of a DNN or a product of machine learning or an artificial intelligence device or a machine learning device corresponds to an alternate disclosure of a convolutional encoder decoder and / or any other type of machine learning device that can enable the teachings detailed herein, or otherwise a form of artificial intelligence that can be utilized to implement the teachings detailed herein and vice versa, all in the interest of textual economy unless otherwise noted providing that the art enables such.
[0109] Thus, as can be seen, in an exemplary embodiment, there is a device comprising a sensory prosthesis, such as a hearing prosthesis. The hearing prosthesis includes an input subsystem configured to receive input based on sound (e.g., a microphone, a WIFI device configured to receive streamed audio signals, etc.) and an output subsystem configured to stimulate tissue (e.g., the electrode array, a mechanical actuator that imparts vibration to tissue, etc.), based on input into the input subsystem to evoke a hearing percept. In an exemplary embodiment, a neural network interposed between the input subsystem and the output subsystem. In an exemplary embodiment, the neural network is part of an output side processing arrangement of an audio processing system of the hearing prosthesis, that outputs electrode activation sequence based data, as opposed to mere channel energies or channel envelopes. That said, in an exemplary embodiment, the neural network can also be part of the input side processing arrangement of an audio processing system of the hearing prosthesis.
[0110] Accordingly, in an exemplary embodiment, the neural network automatically determines electrode channel energies to be stimulated and the timing thereof (and in some embodiments, other things as will be described below) based on input received by the input subsystem and provides the determined electrode channel energies and timing (and some other information in some embodiments as described below) downstream for ultimate utilization by the output subsystem to evoke a hearing percept.[oom] It is noted that in this embodiment, the CED / DNN outputs data that already accommodates or otherwise meets recipient threshold and / or comfort levels (i.e., the actual recipient). By way of example only and not by way of limitation, for a given electrodechannel, which results in the application of electrical energy at a given location within a cochlea, a recipient has a threshold level and a comfort level, or more specifically, the prosthesis has a mapped threshold level and a comfort level for that electrode channel. The prosthesis in general and the CED specifically (or other product of machine learning) will adjust the input energy to accommodate those levels. For example, if the input energy has a very high level of energy, the prosthesis might lower the amount of energy so that it does not exceed a comfort level for that electrode channel, whereas for another electrode channel, the comfort level might be higher, and thus the prosthesis would not necessarily lower the amount of energy for that electrode channel. Put another way, because the output of the DNN is electrode channel specific even though it is not static with respect to which electrode channels will be utilized for given frequencies, a given frequency will be applied to a recipient but constrained by different and / or threshold levels for a electrode channel relative to another electrode channel. In an exemplary embodiment, subsystem 720 performs the adjustment / mapping. But as seen above, in another embodiment, a separate device does so.
[0112] In view of the above, again returning to the system 801 above, the system can be a sound processor system (where this does not mean that there is a processor, only that the system functions as a sound processor). Also in view of the above, in an embodiment, the input subsystem is configured to receive data based on sound captured by a sound capture device (this does not require that the sound capture device be part of the system, but in some embodiments it is part of the system. In an embodiment, the convolutional based subsystem outputs (i) electrode activation sequences or (ii) data sets convertible to electrode activation sequences. This as opposed to a DNN subsystem that outputs channel envelopes (such as those for a CI) or only channel energies. With regard to the data sets convertible to electrode activation sequences, this could be a digital signal or an analog signal that provides a coded set of information that can be decoded by a component downstream to execute electrode stimulation based on that signal. That is, in this exemplary embodiment, no further signal processing is executed after the generation of the electrode activation sequences or the data sets convertible to electrode activation sequences.
[0113] In an exemplary embodiment, the output of the convolutional based subsystem is not provided directly to the electrode driver or a control unit of the electrode driver or a control unit for the electrode driver (the various permutations by way of example depending on the system architecture). Instead, the output of the subsystem 720 is provided to a transmitter and / or a transceiver system, directly or indirectly, which then takes that output and develops adigital and / or analog signal based thereon for transcutaneous transmission from the external component to the internal component, such as by way of example only and not by way of limitation, via the inductance link detailed above and / or via an MI radio link for example or via a Bluetooth signal for that matter or some other radio frequency link, all by way of example only and not by way of limitation. What is received by the implantable component is a signal transmitted across the link that has embodied therein, the electrode activation sequences or the data sets detailed above. The implantable component then takes the received signal, or more accurately, extracts the data content out of the received signal, and then provides that data to the electrode driver apparatus or otherwise to a control unit that controls the electrode driver apparatus, so that the electrode driver apparatus can stimulate the electrodes to evoke a hearing percept. In this regard, everything from the point where the data sets or the sequences are provided to the transceiver or the transmitter component of the external component to the point where the electrodes are activated to evoke a hearing percept is executed by structure of a standard conventional cochlear implant. Again, in an exemplary embodiment, it can be that the sound processing components or otherwise the sound processing pipeline, at least with respect to the output side, are replaced in whole or in part with the subsystem 720. To be clear, in an embodiment, the DNN / CED can output the data that is coded for transmission. The DNN / CED can “learn” how to do this. Or this can be hard coded within the DNN / CED). That said, in an exemplary embodiment, this can be done by a chip or a processing component or software or firmware located outside the DNN and / CED. That is, this can be executed by a hard coded algorithm that exists outside the DNN / CED. This is consistent with the teachings detailed above with respect to having a cochlear implant configured to directly control stimulation by the electrode array to stimulate tissue to evoke a hearing percept based on the data output by the DNN / CED. The fact that some translation is needed to “encode” the output so that it can be transmitted to the implant still meets the direct control example just detailed above.
[0114] The above said, in an exemplary embodiment, the transceiver and / or transmitter transmits electrode activation sequences or comparable data / the components of such, where the energy level within the transmission varies over time so as to directly power the implanted component and directly control the implanted component so that the implanted component receives the signal and the circuitry therein is both powered and activated in accordance with that signal, which signal is the electrode activation sequence data.
[0115] In an exemplary embodiment, bottleneck layers of the convolutional based subsystem develop current levels and / or current timing (and other parameters in some embodiments as detailed herein as will be detailed below) for electrode channels based on data received by the input subsystem, or otherwise develop components of the electrode activation sequence, including but not limited to active electrode, phase duration, interphase gap, return electrode, interpulse interval, etc.
[0116] In an embodiment, there is a cochlear implant, comprising a sound capture device and the system of FIG. 801 by way of example. This is seen in FIG. 11, where there is a functional diagram of a cochlear implant 1101 utilizing the microphone 902 detailed above which is in signal communication with the system 801 as detailed above. In this exemplary embodiment, this can be a totally implantable cochlear implant where the microphone 902 is implanted with the rest of the components, and system 801 is also implanted. In an alternate embodiment, this can be a so-called partially implantable cochlear implant, where the microphone 902 and the system 801 are located in the external component, such as one or part of or with the BTE device, and the output 850 is provided to the transceiver or the transmitter device as noted above, and then transmitted to the implantable component, which receives the signal and then decodes the signal and otherwise provides an output to an electrode array feedthrough in the implantable component which is in signal communication with an electrode array 1131 located in the cochlea by way of an electrode lead assembly 1121 as shown by way of example. With regard to this latter embodiment, the intervening transceiver and / or transmitter system is not shown. In the just described embodiment, the electrode driver is located in the external component and is otherwise part of the external component, such as the BTE device, whereas in an alternate embodiment, the electrode driver could be located in the implantable component, in which case the transmitter and / or the transceiver system detailed above would be located upstream of the electrode driver of the system 801 detailed above.
[0117] In any event, in this exemplary embodiment, the cochlear implant provides data based on sound captured by the sound capture device to the input subsystem, and the input subsystem provides data based on the provided data to the convolutional based subsystem, concomitant with FIG. 11. Here, the convolutional based sub-system provides output based on the data provided to the convolutional based subsystem, wherein the cochlear implant is configured to directly control stimulation by the electrode array to stimulate tissue to evoke a hearing percept based on the data output by the convolutional based subsystem. By directlycontrol, this means that there is no processing, or at least no substantive processing, that occurs after the output of the convolutional based subsystem. There can be data manipulation and data transposition so as to take the output from the convolutional based subsystem and place it into a form that is readable by one of the components downstream, such as by way of example only and not by way of limitation, the electrode driver. By rough analogy, this paragraph can be utilized to directly teach one of ordinary skill who only speaks and reads the French language if this paragraph is translated into the French language. Or perhaps a more apt analogy is the German language, where nouns and verbs will be completely reorganized with respect to their order, but the underlying substance remains. A good enough translation will convey this teaching in a manner sufficient to implement the exact teaching. Conversely, if, for example, the output was further processed so that timing included in the output was changed, that would not be direct control. Also, the concept of direct control is different from receiving a process signal, and then converting it to a data set for control of the driver for example.
[0118] Note that in some embodiments, it is not necessary that the cochlear implant is so configured to directly control stimulation based on the data output by the convolution of subsystem. It could be that the output could simply enable such if such was implemented. In an exemplary embodiment, for some reason or another, it could be utilitarian to have an electrode driver apparatus that is not amenable to being directly controlled by the output from the subsystem 720, or, more likely, the electrode driver is not compatible with the output from the subsystem 720. This could be the case with respect to modifications to existing cochlear implants or otherwise utilizing legacy equipment with respect to the electrode driver, etc. This is unlikely, but certainly a distinct possibility, so thus in an exemplary embodiment, it could be that the output from the subsystem 720 enables such direct control, whether or not such actually occurs in the cochlear implant (to be clear, if it is stated that the cochlear implant is directly controlled, it is so - that means that the cochlear implant has the equipment to do the direct control).
[0119] In an embodiment, there is a cochlear implant system, comprising one or a plurality of sound capture devices. This could be a totally implantable hearing prosthesis with an implanted microphone and an external microphone. This could be the plurality of microphones of a beamforming system where the external component for example includes two or three or more microphones. This could be a bilateral and / or bimodal system that has a cochlear implant on the left side, and a cochlear implant on the right side, and the plurality ofmicrophones are the respective microphones of the implant on the left side and the implant on the right side, or a cochlear implant on one side and a hearing prosthesis that is different therefrom, such as a conventional hearing aid, on the other side, where the cochlear implant system receives data based on the respective microphones of those respective components (e.g., by way of MI radio or a wired transmission - this could be applicable to the other embodiments just noted as well). As with the cochlear implant described above, this cochlear implant system could include the system 801, and an electrode array (or a plurality of such, depending on the exemplary embodiment - if the system is a bilateral system, there could be a second electrode array, but if the system is a multimodal system, for example, despite the nomenclature of a “cochlear implant system,” there could also be a conventional microphone of a conventional hearing aid, or a bone conduction actuator of a bone conduction device, or an actuator of a middle ear stimulation device, etc.).
[0120] In this exemplary embodiment, the cochlear implant system provides data based on sound captured by the plurality of sound capture devices to the input subsystem, wherein the input subsystem provides data based on the provided data to the convolutional subsystem, and the cochlear implant system controls the electrode array to evoke a hearing percept in response to the sound captured by the plurality of sound capture devices while maintaining a coupling between magnitude and phase thereof. This as distinguished from what results from a FFT filterbank for example, where phase information is removed and never returned to the signal / in the processing pipeline. That is, the signal processing capability of the network architectures presented herein can exceed those of traditional signal processing approaches, such as for speech enhancement, noise reduction, and generative applications. For example, the quality of the noise reduction through the convolutional encoder-decoder is much higher than that of a typical Weiner filter. This high signal quality can be attributed to direct processing of the raw audio waveform or data based therefrom, rather than converting to the Fourier domain, which decouples magnitude and phase information. In a conventional CI, the phase information is removed and never returned to the signal. Here, there is no removal of the phase information, and no decoupling of this phase information from the magnitude. That said, in an exemplary embodiment, if there is utilitarian value with respect to decoupling the phase from the magnitude, embodiments include such.
[0121] Thus, in an exemplary embodiment, there is a learned filter bank (or more accurately, the product of machine learning has learned to function as a filter bank - the processing effectively functions as a filter bank - this is described in greater detail below - is functionalbecause there is no specific structure that would be recognized by one of skill in the art as the filter bank that can be pointed to per se or otherwise there are no electronics that would be readily identifiable as a filter bank) that can filter frequencies of incoming sound or incoming light, while maintaining a coupling between the magnitude of those frequencies and the phase of those frequencies.
[0122] Accordingly, in an exemplary embodiment, the processing is based on phased awareness, where the phase info is present and is utilized in the processing, including as will be detailed below, processing to reduce noise and / or enhance speech understanding or otherwise produce a high quality speech output or at least one that is better than that which would otherwise be the case. To this end, embodiments attempt to mimic normal hearing in the auditory portion of the nervous system. Again, these embodiments can be directed to a cochlear implant, which directly stimulates the nervous system of the human being and bypasses the mechanical or structural bridge between the waves of pressure that are generated in the cochlea and the nerves within the cochlea or otherwise the nerves that provide electrical impulses or more accurately, provide a conduit for the electrical impulses that are transduced by the cochlea. That is, the electrical output of the cochlear implant electrode array, as based on the teachings detailed herein, results in mimicking normal hearing of a human, at least better than that which would otherwise be the case. It is not exact. Or if it is exact, it is more by fortuitous accident. The teachings detailed herein provide a close enough approximation to nerve activation for normal hearing to provide superior results over the conventional teachings.
[0123] In an embodiment, there is a cochlear implant, comprising a sound capture device, the system 801 for example, and an electrode array. Again, the cochlear implant provides data based on sound captured by the sound capture device to the input subsystem, wherein the input subsystem provides data based on the provided data to the convolutional based subsystem. Here, the cochlear implant is configured to stimulate tissue to evoke a hearing percept based on the electrode activation sequence based data outputted by the system 801, or at least the convolutional based subsystem 720. Also, in this embodiment, there is no dedicated circuitry solely dedicated to electrode activation sequencing. This as contrasted to a conventional cochlear implant, which has dedicated circuitry solely dedicated to electrode activation sequencing. With respect to the former, there could be circuitry that has a dual role by way of example, where here, the convolutional based subsystem performs many roles, and in some embodiments, all at once. And thus, there is nothing solely dedicated to electrodeactivation sequencing (which does not mean that the sequences are developed per se - this is circuitry that is utilized to develop electrode activation sequencing which can be broken up into separate components or otherwise modules - here, there is no specific circuitry that relates solely to the development of electrode activation sequencing as one of skill would understand vis-a-vis the dedicated circuitry for such - in some embodiments, any such circuitry even if it does not develop the entire sequence would be excluded). In some embodiments, the DNN / CED / product of machine learning executes sampling and selection for example, and amplitude mapping, or one or two of those but not one or two others of those,. In an exemplary embodiment, there is no dedicated circuitry solely dedicated to FFT filtering and / or sampling and selection and / or amplitude mapping for that matter.
[0124] Still with respect to the implementation of the teachings herein to a cochlear implant, in an exemplary embodiment, the system includes a sound coding path for a cochlear implant and the path is entirely made up of a convolutional encoder-decoder network. Note that this does not mean that the sound coding path does not have an analog to digital converter for example, and / or other conventional components that massage or manipulate the data so that the data can be coded. This is with respect to the sound coding path as one of ordinary skill in the cochlear implant arts would understand. A sampling and selection processing block that is separate from the convolutional encoder-decoder network is excluded in some embodiments.
[0125] Thus, it can be seen that speech coding strategies could benefit from the teachings herein. In one embodiment, the sound coding path could be replaced entirely by the convolutional encoder-decoder network. The input could be raw audio, or features derived from raw audio, and the output could be electrode stimulation sequences, all by way of example. In an embodiment, all aspects of the cochlear implant stimulation strategy could be performed by the neural network.
[0126] Embodiments include methods. In an exemplary embodiment, there is a method 1200, which is represented by the exemplary flowchart of figure 12, that includes method action 1210, which includes the action of obtaining output of signal processing of an input signal, wherein the signal processing is conventional audio signal processing of the input signal.
[0127] This can be executed in accordance with any of the teachings above applicable to the alternative path of figure 10 that bypasses the DNN subsystem. This can be executed in accordance with any of the conventional audio signal processing that is implemented in a cochlear implant by way of example as of October 13, 2023, as is commercially available on that date and FDA approved for general use in the United States of America on that date and / or as is available and is approved for general use in the United Kingdom, the Federal Republic of Germany, the Republic of France, the Commonwealth of Australia and / or the People’s Republic of China on that date.
[0128] Method 1200 further includes method action 1220, which includes the action of providing data based on the obtained output to a convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided data, wherein the output is electrical stimulation control data for a cochlear implant (e.g., an electrode activation sequence, or the datasets detailed herein). FIG. 13 presents an exemplary training framework for a convolutional encoder-decoder based processing strategy, by way of example. Here, the electrode activation patterns from the typical cochlear implant processing, and the CED, are compared to one another utilizing any comparison regime that will have utilitarian value, where the loss function minimizes difference between the patterns. Note that with respect to the action 1220 detailed above, the act of providing the feedback from the loss function block to the CED represented in figure 13 corresponds to the action of providing data based on the obtained output to the CED and training the CED based on the provided data. This is because the comparison is based on the obtained output from the typical cochlear implant processing, and the feedback to the CED is based on the comparison. Thus, the data provided to the CED is based on the obtained output from the typical cochlear implant processing. It need not be the actual data, but it only need be based thereon. In an exemplary embodiment, the CED can be provided with the actual output, and the CED can evaluate that output and compare the output to what the CED developed based on the same input, thereby training the CED. This too would be providing data based on the obtained output, because it is the obtained output. Thus, in an embodiment, the training includes applying a loss function that minimizes differences between electrode activation patterns of the output and an output of the convolutional encoder-decoder.
[0129] Note that this is different from, for example, providing data based on channel envelopes by way of example. This is different than comparing channel envelopes. Here, it is the electrical stimulation control data that is utilized for the comparison.
[0130] In an exemplary embodiment, the convolutional encoder-decoder trains itself on one or more or all aspects of a cochlear implant stimulation strategy during the training, depending on the embodiment.
[0131] Corollary to figure 13 is that in an exemplary embodiment, the method can include obtaining output from the CED based on the input signal input into the conventional processing path (the signal can be split as shown, or data based on the input signal can be provided to the CED - that is, in some embodiments, the same signal is provided to both the CED and the conventional processing path, and in other embodiments, different signals are provided, but the signals are based on the same signal so as to obtain an apples to apples comparison or at least as close of the comparison as would be utilitarian). The output from the CED and the output from the conventional processing can be compared in a manner concomitant with that detailed above, where the output is the electrical stimulation control data for the cochlear implant. But note that this could be for a BI in an alternate embodiment, or any other sensory prostheses to which the teachings detailed herein can have utilitarian value. Accordingly, in an exemplary embodiment, there is a method, such as method 1400, as seen in figure 14, which includes method action 1410, which includes the method action 1410, which includes obtaining output of signal processing of an input signal, wherein the signal processing is conventional audio and / or visual signal processing of the input signal and method action 1420, which includes providing data based on the obtained output to a convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided data, wherein the output is electrical stimulation control data for a sensory prosthesis. In this exemplary embodiment, the conventional audio and / or visual signal processing is conventional audio processing and the output is electrical stimulation control data for a cochlear implant. In an exemplary embodiment, the conventional audio and / or visual signal processing is conventional video processing, and the output is electrical stimulation control data for a bionic eye. Still, again by way of example, with respect to a cochlear implant, in an embodiment, the signal processing of the conventional pipeline is a signal processing of a Nucleus 6, 7 and / or 8™ as those are in existence as of October 13, 2023, in any one or more of the just noted jurisdictions / approved for general use in any one or more of the above noted jurisdictions.
[0132] In an embodiment, the convolutional encoder-decoder is trained, simultaneously, in one or more, two or more, three or more, four or more, five or more, 6 or more, 7 or more, or 8 or more or 9 or more or all of: microphone equalization; automatic gain control; envelopeextraction; loudness growth functions; noise reduction; speech enhancement; maxima selection; envelope sampling; acoustic-to-electric mapping and / or pulse sequencing. This is concomitant with a CEC that can execute one or more of the functions of the output side processing of the cochlear implant processing strategy, including the electrode sequence control data. All by way of example.
[0133] By way of example, the training can be based on comparison on the stimulation rates of one or more or all electrodes (or channels). This can entail comparing the number of pulses per second (e.g., the conventional has 333 pulses per second on electrode 5 and the CED has 322 per second on that electrode, so the training attempts to change the “thinking” of the CED so that the output of the CED will have a number closer to 333 pulses per second on that electrode, which includes having 333 pulses per second on that electrode). This can alternatively and / or in addition to this entail comparing current levels. By way of example, if the current level for electrode 8 is 457 pA, and the CED comes up with 488 pA, the training attempts to change the thinking of the CED so that the output of the CED will have a number closer to 457 pA. This can alternatively and / or in addition to this comparing the pulse widths (e.g., the conventional has a pulse width of 12 ps for example, in the CED has a pulse width of 17 ps, so the training attempts to change the thinking of the CED so that the output of the CED will have a number closer to 12 ps). All of this by way of example only and not by way of limitation. But again, in an exemplary embodiment, the loss functions can include comparing two or three or more or any of these to each other and training the CED based on the given loss function.
[0134] FIG. 15 shows a diagram of a convolutional encoder-decoder based low power CI strategy training framework by way of example. Here, there are two genuses of loss functions: the above-noted loss function and the penalization of the number of pulses used. More specifically, by way of example and not by way of limitation, embodiments can use the teachings herein for low-power speech coding strategies. For example, in an embodiment, to implement or otherwise achieve a low-power stimulation strategy, the number of pulses and / or the pulse length, etc., used to encode a sound could be incorporated into the loss function. The target stimulation strategy during training could be derived from typical, non- DNN based coding strategy.
[0135] In an exemplary embodiment, the output of the non-DNN based coding strategy could be passed to a neural activation model (an electrical hearing model that models neural activation in response to electrical stimulation in the cochlea), and this neural activationpattern would become the target during training. During training, the convolutional encoderdecoder strategy can produce an electrode stimulation sequence or data therefor which would be passed through the same neural activation model as the target (it could be a duplicate model - it need not literally be only one model that is used - it could be a clone). The loss function could thus have two functions: 1) minimize the difference between neural activation patterns, and 2) penalize the number of pulses used and / or the length of time of the pulses used, consistent with FIG. 15.
[0136] This can have utilitarian value with respect to partially implantable cochlear implants (or Bis that have an external camera). In CI, around 40% of the power is used to transmit pulse instructions from the coil to the implanted receiver. If an embodiment can achieve a similar / suitable / effective neural activation patterns for half the number of pulses or half the pulse length (by way of example), one could conserve 20% of power, which would add hours to the battery of the external device and / or the battery autonomy of the implant. In an embodiment, a key to such is that that the processing strategy is deployed via a trainable CED.
[0137] Accordingly, in an exemplary embodiment, there is the method 1200 detailed above, wherein the training includes penalizing for simulated energy quantities used by the convolutional encoder-decoder. In an exemplary embodiment, there is the method 1200 detailed above, wherein the training includes penalizing for pulse number per second or per stimulation channel per given output amplitude (current) and / or pulse length (total and / or mean and / or median and / or mode) per second or per stimulation channel (again total and / or mean and / or median and / or mode) per given output amplitude (current) used by the convolutional encoder-decoder. In an embodiment, the reduction is at least and / or equal to 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80%, or any value or range of values therebetween in 1% increments relative to that which would otherwise be the case, all other things being equal.
[0138] While the embodiment presented with respect to figure 15 utilizes a typical cochlear implant processing regime to set the standard with respect to the embodiment of reducing the energy utilized to affect the results, in an alternate embodiment, such as that represented by way of example in figure 15 A, a trained CED could be utilized to set the standard training an untrained CED. And note that in this exemplary embodiment, while there are two blocks presented for the CED, in an exemplary embodiment, this can be the same CED, where the CED trains itself and the CED could output to two different models. Still, the idea could bethat the initial CED could ultimately be replaced by the now trained and previously untrained CED. Thus, a modification of method 1200 could include a revised method action 1210, which is obtaining output of signal processing of an input signal (regardless of how the processing is executed) and then executing method action 1220 (providing data based on the obtained output to a convolutional encoder-decoder and training the convolutional encoderdecoder based on the provided data), wherein the output is electrical stimulation control data for a sensory prosthesis. Here, in an embodiment, the signal processing is executed by a CED or otherwise a product of machine learning as opposed to (or in addition to) signal processing by a conventional audio and / or visual signal processing of the input signal. Note that in an exemplary embodiment, the loss function for the arrangement of figure 15A can also be the minimize station of the difference between neural activation patterns.
[0139] In an embodiment, there is a method 1600, as represented by the flowchart of FIG. 16. This includes method action 1610, which entails executing method 1200. The method 1600 also includes method action 1620, which includes providing the output to a neural activation model, wherein the output is a first output. This would happen before method action 1220 is executed (note that there is no order in the listing of the actions unless otherwise noted - any presentation in a given order is simply a matter of genera accounting, but note that also corresponds to an embodiment that has that order, but in the absence of a statement that such is the order of action, there is no requirements for order providing that the art enables such). Method 1600 also includes method action 1630, which includes obtaining first data based on first neural activation model output based on the first output provided to the neural activation model. This is represented by way of example of the top rightmost arrow in figure 15. Method 1600 also includes method action 1640, which includes the action of obtaining second output from the convolutional encoder-decoder (represented by the right side arrow leaving the CED of FIG. 15), and method action 1650, which includes the action of providing the second output to the neural activation model. Method 1600 also includes method action 1660, which includes obtaining second data based on second neural activation module output based on the second output.
[0140] In an exemplary embodiment, the action of training includes minimizing difference between the neural activation patterns in the first data and the second data by way of example only and not by way of limitation. Here, the models as noted above attempt to emulate or otherwise protect neural activation when electrical stimulation is provided thereto, such as by way of example with respect to electrical stimulation from a cochlear implant electrode arraylocated in a cochlea. The minimizing of the deference can be a absolute determination or can be a good enough determination. By way of example only and not by way of limitation, the action of minimizing can be where the neural activation patterns are effectively the same with respect to that which would evoke a hearing percept between the two (thus the hearing percept is effectively the same). Conversely, it could be that the activation patterns are different, and potentially in an identifiable way that could be meaningful, but the difference is not so sufficient as to warrant further training or otherwise improves upon that which previously was the case. In this regard, the difference has been minimized. In an exemplary embodiment, the difference is minimized so that amplitudes of the neural activation patterns are no more than and / or equal to 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0.9, 0.8. 0.7. 0.6, 0.5, 0.4, 0.3, 0.2, or 0.1%, or any value or range of values therebetween in 0.01% increments different from each other (utilizing the neural activation patterns of the conventional processing arrangement as the denominator). In an exemplary embodiment, the action of training includes penalizing for simulated number of electrical pulses used by the convolution encoder-decoder and / or the length of those pulses. In an exemplary embodiment, the penalization can be weighted. The length can be weighted less than or greater than the number of pulses, and / or vice versa. In an exemplary embodiment, one is weighted and / or the other is weighted so that the normalized weight is such that one is less than and / or equal to 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, or 5%, or any value or range of values therebetween 1% increments of the other.
[0141] In an embodiment, there is a device, such as a cochlear implant or a bionic eye, by way of example only and not by way of limitation, that includes a product of machine learning. This can be the CED or another DNN or for example a chip based on the CED or DNN, such as a chip with firmware thereon or circuitry thereon that is developed to replicate the action of a trained CED or DNN, etc., as detailed herein and variations thereof. In an embodiment, the product is configured to receive an input based on a captured phenomenon in the visual and / or audio environment. By way of example only and not by way of limitation, the input could be input that is based on output from a microphone or a digital camera by way of example only and not by way of limitation.
[0142] Consistent with the teachings detailed herein, the product of the machine learning is configured to process the input and automatically develop electrode activation control data. In an exemplary embodiment, this is the electrode activation sequences that are utilized bythe electrode driver of a cochlear implant or the electrode driver of a bionic eye. In an exemplary embodiment, this can be the data sets detailed above that are utilized by a control device that controls the driver. In an exemplary embodiment, the electrode activation control data can be the coding that is provided to the transceiver or transmitter apparatus that is utilized to transcutaneously transmit the power and the control data to the implanted component. In an exemplary embodiment, the electrode activation control data can be the coded signal that is developed by the transceiver and / or transmitter apparatus for transfer to the implanted component.
[0143] As noted above, in an exemplary embodiment, the device can be a sensory prosthesis that includes electrodes, concomitant with a cochlear implant. In an exemplary embodiment, the device is configured to activate the electrodes based on the developed electrode activation control data to evoke a sensory percept. This can be direct activation. This can be done based solely on the developed electrode activation control data. This can be done based directly on that data as noted above.
[0144] In embodiments where the device is a cochlear implant, or a bionic eye, by way of example, there is no processing (signal processing) in the signal processing pipe of the device after the product. That said, in some embodiments where there is signal processing in the signal processing pipe of the device after the product, the processing does not change the data contained in the control data with respect to content. By way of example only and not by way of limitation, there is none of microphone equalization; automatic gain control; envelope extraction; loudness growth functions; noise reduction; speech enhancement; maxima selection; envelope sampling; acoustic-to-electric mapping and / or pulse sequencing. This is concomitant with a CEC that can execute one or more of the functions of the output side processing of the cochlear implant processing strategy, including the electrode sequence control data. All by way of example.
[0145] As noted above, embodiments can include an implant that includes a number of electrodes, such as by way of example any of the numbers detailed herein or additional numbers for that matter. In an embodiment, there are at least N electrode channels (any of the numbers or ranges detailed herein). In an exemplary embodiment, the electrode activation control data that is developed by the product of machine learning is dimensionalized relative to the number of channels of the at least N channels.
[0146] In an exemplary embodiment, such as where the device is a hearing prosthesis, the electrode control data is devoid of audio content. This as opposed to, for example, output from a product of machine learning that corresponds to channel envelopes, such as where the hearing prostheses post processes the channel envelopes. In this regard, the signals provided to the electrode driver are devoid of audio content. There is no output audio however processed from the product of machine learning. And comparing this to conventional processing, where amplitude envelopes (channel envelopes) are determined after frequency binning. One can consider envelopes as slow moving changes in the magnitudes of filterbank outputs. The conventional processing then uses the envelopes to develop what can be called electrodograms. Electrodograms include the conversion (as opposed to mere translation) of those envelopes to stimulation sequences that are sent to the stimulation dedicated devices. It is this conversion that is executed by the CED / product of machine learning. And note that while embodiments herein have focused on a CED that outputs electrode activation sequences, in other embodiments, the output might not be such. In an exemplary embodiment, the output could simply be the results of the conversion to the electrodeograms. Further processing downstream could be implemented to obtain the electrode activation sequences to be used to stimulate the recipient.
[0147] Consistent with the teachings detailed above, at least some input side processing and all output side processing is executed by the product of machine learning with respect to the device under discussion. In an exemplary embodiment, the only signal processing that occurs outside the product of machine learning is the front end processing, such as that established by an analog to digital converter and / or an amplifier associated with a microphone of the device.
[0148] In an embodiment, the product receives a raw audio signal and outputs electrode stimulation sequences and / or the other control data detailed herein. In an exemplary embodiment, the product receives a pre-processed audio signal, and the output is still the same as just mentioned. In an exemplary embodiment, the product receives data based on raw audio and the output is still the same as just mentioned.
[0149] In an embodiment, again, there is a device, such as any of the devices detailed herein (CI, BE, or here, a chip or processor. This device comprises a signal input (this could be a jack or any other electrical component (or optical component) that can enable signal input) configured to receive data based on audio data and / or visual data. The device also includes circuitry (such as that of a chip or of a processor - this could be part of / the CED forexample) configured to simultaneously functionally execute at least some input side processing and at least some output side processing of the data based on audio data and / or visual data. In an exemplary embodiment, the circuitry is configured to simultaneously functionally execute some input side processing and all output side processing of the data based on the audio data and / or visual data. Note that the CED will execute actions, at least some of them, sequentially. However, as a block of circuitry (between the input and the output), the functionality is simultaneous.
[0150] First, with respect to the phrase “functionally execute,” this means that the execution does not require a clear distinct action or componentry that does the executed action, only that the result is functionally the same. Again, this is being done by a CED where the circuitry thereof and otherwise the architecture thereof does not lend itself to pointing to a specific “location” or specific substructure that executes a specific function. In many instances, it is not exactly known how the CED processes the input to result in the output. This is understood in the art.
[0151] With the above in mind, it is noted that in some exemplary embodiments, not all output side processing is executed simultaneously with the input side processing. Instead, in an exemplary embodiment, only some of the output side processing is executed simultaneously. Any one or more of the identified output side processings can be executed without this temporal constraint. In an embodiment, the output side processing executed by the circuitry that is executed simultaneously with the at least some input side processing includes determination of timing and / or pattern of stimulation on one or more electrodes in signal communication with the device. In this regard, in an exemplary embodiment, the device could be part of a cochlear implant for example, and the device can be in signal communication with the electrode driver of the electrode array and thus in signal communication with the electrode array. In an exemplary embodiment, the output side processing executed by the circuitry that is executed simultaneously with the at least some input side processing includes channel envelope development and / or frequency bin division.
[0152] In an exemplary embodiment, the circuitry under discussion and otherwise the circuitry or devices and / or systems, etc. disclosed herein can mimic hard-coded filter banks in the input side processing. In an exemplary embodiment, the circuitry or devices and / or systems disclosed herein can mimic any one or more hard coded processing actions and / or features disclosed herein. Thus, by way of example, there is a device and / or system and / or method and / or circuitry and / or processor, etc., that executes / configured to execute, withouthard coding, one or more, two or more, three or more, four or more, five or more, 6 or more, 7 or more, or 8 or more or 9 or more or all of: microphone equalization; automatic gain control; envelope extraction; loudness growth functions; noise reduction; speech enhancement; maxima selection; envelope sampling; acoustic-to-electric mapping and / or pulse sequencing. In an embodiment, any one or more of these is done simultaneously, with or without simultaneous partial input side processing. In an embodiment, the convolutional encoder-decoders used herein avoid the need and thus enable embodiments where there is no hard-coded filter bank.
[0153] Note that by mimicking, the mimicking is done without the hard coding. This is concomitant with the teachings detailed herein. Embodiments can accomplish the goal or otherwise the function of hard coded processing features without hard coding.
[0154] In an embodiment, there is thus a cochlear implant, comprising a microphone, an electrode array and a device such as that detailed above and / or below and variations thereof (or a system or subsystem, etc.), wherein the cochlear implant controls stimulation of tissue with the electrode array based on results of the input side processing and the output side processing. In an embodiment, concomitant with the teachings above, the output of the device is electrode activation sequence data. Further, as noted above, in an embodiment, there is no dedicated circuitry dedicated to input side processing outside of front-end processing and / or there is no dedicated circuitry dedicated to output side processing. Certainly there is no dedicated circuitry solely dedicated to those things. Instead, in an exemplary embodiment, the input side and the output side processing aside from those caveats is implemented by the CED by way of example. And in an exemplary embodiment, such as where there is a cochlear implant or a bionic eyes, there is a digital camera or other type of camera and / or a microphone, a electrode assembly, such as an electrode array located in the cochlea, where the device detailed herein is utilized, and there is no circuitry solely dedicated to the electrode activation sequencing.
[0155] In an embodiment, there is a method, such as that represented by the algorithm of FIG. 17, presenting an algorithm for method 1700, that includes a method action 1710, which includes receiving data based on at least one of ambient sound and / or ambient light in an ambient environment of a device or electromagnetic signals received by the device, the data enabling direct usage thereof by a conventional microphone apparatus and / or a conventional video monitor apparatus to produce a recognizable corresponding sound and / or imaging, respectively without further processing. In an exemplary embodiment, this could bereceiving an analog and / or a digital signal from a camera and / or a microphone. In an exemplary embodiment, this could be receiving a data set in which is included data based on at least one of ambient sound and / or ambient light received by the device. In this regard, embodiments can include remote processing. Accordingly, in an exemplary embodiment, the action of receiving can be executed by server. With respect to the electromagnetic signals, this could be by way of example streamed audio and / or video data. This could be data received by a radio transceiver or a receiver for example. And again, the action of receiving could also be executed by a system or other apparatus that is remote from the device, such as a server. The requirement here is that the received data be based on the exemplary ambient sound in an ambient environment of a device. Indeed, the device need not receive that ambient sound and / or ambient light according to this exemplary method, while in other embodiments, the device does receive that ambient sound and / or ambient light.
[0156] The data that is received is data that enables direct usage thereof by a speaker apparatus and / or a monitor apparatus to produce a recognizable corresponding sound and / or image without further processing. This could be for example the data that is output from an analog to digital converter. This would not be the results of mapping and / or frequency bin division, because additional processing would be needed for this to be used by conventional speaker apparatuses to produce a recognizable corresponding sound. The speaker apparatus would have to be a specialized speaker apparatus that can accommodate the divided frequencies. With respect to the results of the mapping, the channel enveloping would result in the signal that would require the application or modulation of a sine wave(s) to be combined there with to have a usable signal. Granted, this is not technically difficult to achieve, but that is not a conventional speaker apparatus. There would have to be further processing of the signal to recombine the signal so that output could be utilized by a microphone apparatus if only to recombine the channels into a single signal for that conventional microphone apparatus to be utilized. In an exemplary embodiment, the received data is a received audio signal from a microphone. That said, the received data could be data that results from preprocessing of that received audio signal. The data could have beamforming data included therein or otherwise could have been manipulated to account for beamforming, providing that the above noted requirements with respect to the conventional microphone or monitor is met. The data could have noise cancellation applied thereto already, again providing that the above noted requirements with respect to conventional microphone or monitor is met.
[0157] The method also includes method action 1720, which includes the action of generating electrode activation sequences directly from the received data. As noted above, that received data has the conventional microphone or monitor qualifications. That is the data that is being used, not data that is modified after it is received. Here, the operative phrase is “directly.” This has been detailed above and otherwise herein, but briefly, the generation of the sequences results from the action of generating without any intervening processing. Note that this does not mean that there is no translation as detailed above or no conversion from analog to digital, because the received data is now the digital data. The device would not necessarily encompass the microphone and / or the front end processing, but could. This device could be by way of example the CED detailed herein or otherwise the product of machine learning. Put another way, the received data could already have the analog-to-digital conversion or otherwise the front end processing. The method does not require receiving the output from a sound capture light capture device, but instead receiving data that is based on ambient sound and / or ambient light or electromagnetic signals received by the device. Thus, an analog-to-digital conversion would still result in data based on the ambient sound, etc.
[0158] In an exemplary embodiment, the action of generating electrode activation sequencies is executed in real time vis-a-vis the action of receiving data, and the method further comprises evoking a hearing percept based on the generated electrode activation sequences. In an exemplary embodiment, the action of generating electrode activation sequences is executed within 50, 45, 40, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0.9, 0.8, 0.7, 0.6 or 0.5, 0.4, 0.3, or 0.2 microseconds or any value or range of values therebetween in 0.01 ps increments. In an exemplary embodiment, the actions of receiving data and generating electrode activation sequences are executed by a body-worn cochlear implant without converting data based on the received data to a Fourier domain. In an exemplary embodiment, the action of generating electrode activation sequences is executed by a product of a trained neural network (which thus includes a trained neural network).
[0159] In this exemplary embodiment, all output side processing is executed by data-driven developed algorithms. In this exemplary embodiment, all non-front end input side processing is executed by data-driven developed algorithms. In an exemplary embodiment, the circuitry or devices and / or systems and / or methods by way of example execute / are configured to execute, one or more, two or more, three or more, four or more, five or more, 6 or more, 7 ormore, or 8 or more or 9 or more or all of the following by data-driven developed algorithms: microphone equalization; automatic gain control; envelope extraction; loudness growth functions; noise reduction; speech enhancement; maxima selection; envelope sampling; acoustic-to-electric mapping and / or pulse sequencing. In an embodiment, any one or more of these is done simultaneously, with or without simultaneous partial input side processing.
[0160] Embodiments include networks (e.g., the CED) that efficiently capture local and global features, enabling hierarchical feature extraction across various time scales. Embodiments can down-sample input-signals (e.g., with the CED) to create a learned latent representation, expanding temporal context and reducing complexity in bottleneck layers. This learned feature space can allow, for example, autonomous feature discovery. This can reduce and / or eliminate a need for predefined input features (and embodiment can be devoid of such predefined input features). Additionally, parameter sharing across inputs reduces memory usage in convolutional encoder-decoder structures, concomitant with the use thereof for deployment in real-time wearable systems such as CI and BE. Thus, in an exemplary embodiment, the actions of receiving data and generating electrode activation sequences are executed by a body-worn cochlear implant. The action of generating electrode activation sequences is based on both global and local features and / or learned feature representation in at least some embodiments, as will be described in greater detail below.
[0161] In an embodiment, there is an exemplary method as represented by way of example by the algorithm presented in figure 18, which shows an algorithm for method 1800, which includes method action 1810, which includes the action of executing method 1700. Method 1800 further includes method action 1820, which includes the action of generating electrode activation sequences indirectly from the received data. This can correspond to, for example, the utilization of the device of FIG. 10. This can be done for redundancy purposes by way of example only. In an embodiment, this can be done if the machine learning based pipeline fails. In this regard, in an exemplary embodiment, the processing pipeline could utilize a conventional filter that divides the input signal, as well as the other conventional processing detailed above. The development of the electrode activation sequences would only be indirectly from the received data because of this intervening processing, for example.
[0162] FIG. 19 shows an exemplary flowchart for an exemplary method, method 1900, which includes method action 1910, which includes the action of obtaining data based on audio and / or visual data. Method 1900 also includes method action 1920, which includes processing the obtained data using one or more algorithms based on machine learning. Also,method 1900 includes method action 1930, which includes evoking a sensory percept based on the processed data. In an exemplary embodiment, the action of processing the data using one or more algorithms one or more of: (i) executes output side processing of a cochlear implant processing pipeline or (ii) executes one or more of: amplitude compression; channel envelope sampling; maxima selection; speech enhancement; binaural processing; or microphone equalization.
[0163] In an embodiment where “ii” is executed, at least 2, 3, 4, 5 or 6 of the actions are executed. In an embodiment, the processing generates electrode activation sequences directly from the data based on audio data, as distinguished from, for example, sampling and mapping from output of an enveloping device and / or a filter bin device. In an embodiment, the processing generates the activation sequences directly from the data utilizing a convolutional encoder-decoder. In an embodiment, the data obtained in method action 1910 is data based on audio data. The processing of method action 1920 is processing of a sound coding path for a cochlear implant (again, an embodiment is to utilize the CED for the sound coding path of a cochlear implant). In this exemplary embodiment, the path is entirely made up of a convolutional encoder-decoder network.
[0164] Embodiments can be applicable to totally implantable cochlear implants. Such devices have an implanted microphone. This implanted microphone can pick up body noise which is undesired as well as the desired ambient environment sound that travels through the skin to the implanted microphone. The implanted microphone can be prone to patientspecific high-frequency roll-off in the transfer function for example. (Note that a totally implantable cochlear implant can typically also be usable with an external microphone. Often, the external microphone will be utilized when the battery or other power source of the implanted device is not sufficient to operate the implant in the totally implantable mode. In an exemplary embodiment, the external microphone is much less prone to body noise because it is located outside the body.) In an exemplary embodiment, electrode activation sequences or data based thereon or data that is usable to directly develop electrode activation sequences without further processing can be generated by the convolutional encoder-decoder based on input from the internal microphone / implanted microphone. During the training process as represented by way of example only and not by way of limitation by the schematic of figure 20, the convolutional encoder-decoder could learn to emulate the external microphone signal. In an exemplary embodiment, the CED could learn to remove body noise and / or restore missing high frequency information. In some instances, this network could generate highfrequency information that was not included in the input signal. In these instances, there can be utilitarian value to convolutional encoder-decoder architectures, because they are able to generate information.
[0165] And note that the embodiment of figure 20 has a different paradigm than at least some of the other embodiments detailed herein in that the training is not based on a fully processed signal or even a signal that is processed beyond front-end processing in some embodiments. The embodiment of figure 20 compares the outputs from an external microphone and an internal microphone. Granted, some embodiments utilize a more thorough processing pipeline, where there is a conventional cochlear implant processing pipeline that starts with input from the external microphone and the result thereof is compared to the processing that results utilizing the CED where the CED input is into the internal microphone. In either of these arrangements however, the audio is captured utilizing different components. This is a different paradigm from many of the embodiments detailed above, where the audio is captured by the same component. Put another way, the audio in with respect to figure 20 can correspond to a raw ambient environment sound. The sound is captured by the respective microphones, and the energy thereof is transduced to result in an output signal, where the signals are compared to each other in accordance with the teachings detailed herein and / or where the signals are then processed according to the teachings detailed herein and the results are compared. To be clear, in some embodiments, the machine learning is not being utilized to replace the cochlear implant processing or otherwise supplement or mimic the cochlear implant processing. Instead, in some embodiments, the machine learning is utilized to account for the various body noises and / or the transfer functions which would exist irrespective of the use of a conventional cochlear implant processing regime or the DNN and based processing regime.
[0166] Thus, in an embodiment, there is a device, such as a hearing prosthesis such as a cochlear implant, that includes a microphone and a tissue stimulation output assembly (electrode driver and lead and electrode array, etc.). The device includes a sound processing path that receives input from the implanted microphone. The device controls the tissue stimulation output assembly based on output from the sound processing path to evoke a hearing percept based on the input from the implanted microphone. This can be by providing the tissue stimulation output assembly a digital data set that the output assembly can read and from that digital data set control current and / or voltage to the electrodes based thereon toevoke a hearing percept based on the data set. The control could also be by way of output of electrode activation sequences to the stimulation output assembly.
[0167] In an exemplary embodiment, at least one of (i) the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output (ii) the sound processing path includes a product of machine learning trained to minimize feedback or (iii) a processing path that includes at least a portion of the sound processing path and extends to the tissue stimulation assembly is made entirely of a product of machine learning. Some exemplary examples of these will now be described.
[0168] In an exemplary embodiment, the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output. As noted above, this can be based on a comparison of the output signals from the respective microphones for the same noise or sound in the ambient environment. In an exemplary embodiment, this is executed in a controlled manner where the sound of the ambient environment is clear and otherwise clean. In an embodiment, a little to no background noise is present. The sound could be speech or could be other sounds that are utilitarian or more accurately more utilitarian to the recipient than others, such as for example, music if such is desired by the recipient. The idea is that the ambient sound will be clear in at least some embodiments, or otherwise is clear as possible. Training could take place within a quiet room for example. The sources of the sound could be located the way such would be located when the recipient is utilizing the prostheses to hear these important sounds (for example, or right in front of the recipient with respect to speech where the idea is that the recipient is going to be looking and facing directly toward someone to whom he or she is speaking and / or whom is speaking to him or her). That said, in an exemplary embodiment, the training takes place, or more accurately, the data collection takes place (the data can be data logged in compared or otherwise evaluated later including days or weeks later, and the training can thus take place days or weeks later as well) in a more casual environment or otherwise a less controlled environment. In an exemplary embodiment, the training takes place during the normal course of usage of the hearing prostheses in the recipient’s normal day. Any regime of data collection that can enable the teachings detailed herein can be utilized in at least some exemplary embodiments.
[0169] And note that while the embodiment above focuses on the concept of having data collection from a microphone that is implanted in a live human, in another embodiment, theimplanted microphone can be implanted in a simulated implanted environment. This can be placed into a cadaver or could be placed into a model that mimics the skin tissue and / or body noises, etc. Still, in an exemplary embodiment, the idea is that data collection will occur utilizing the external microphone and the implanted microphone simultaneously so that the respective microphones capture the same ambient sound albeit at different locations at the same time. And it is data based on the output of the microphones that is utilized to train the CED. The training could be based on the exact output from the microphones, or could be based on one the output of the microphones that are subject to ADC conversion. There could be additional processing that occurs to one or both of the signals from the microphones. As long as the processing is processing that occurs after the CED is up and running, there is an apples to apples comparison and thus the processing will not interfere or otherwise skew the training. Indeed, in an exemplary embodiment, there can be processing of the signal from one of the microphones and not the other. As long as that is the case when the prosthesis is utilized, the apples to apples comparison can have utilitarian value.
[0170] Thus, in an embodiment, referring back to method 1200 by way of example, the input signal is based on audio captured by an external microphone exposed to an air environment (e.g., the external microphone of the external component / on the BTE). The method can further comprise providing a second input signal to the convolutional encoder-decoder, wherein the second input signal is based on audio captured by an implantable microphone in at least a simulated implanted environment (which could be the actually implanted in a human (hence the “at least”) and obtaining a second output from the convolutional encoderdecoder based on the provided second input. Here, the training can include comparing the obtained output to the second output, wherein the provided data is based on the comparison.
[0171] But note that in embodiments, the action of training to minimize the difference between data based on implanted microphone output and data based on the external microphone output can be executed after the processing by the conventional processing pipeline and the processing by the CED. In fact, in an exemplary embodiment, it could be that the processing of the output from the external microphone could be done by a trained CED and the output thereof can be utilized to train the CED to process the output from the implanted microphone. Again, the idea here is that the effects of body noise and / or the transfer function changes, etc., can be mitigated by utilizing machine learning. It could be that the CED is already trained on how to process output from an external microphone. Here, the CED or another CED is being trained on how to process output from the implantedmicrophone. Any device system and / or method that will enable training of the CED and thus obtaining a product of machine learning that can improve the utilization of an implanted microphone can be utilized in at least some exemplary embodiments. In an embodiment, it is the trained CED that is utilized to process input from the external microphone during use of the cochlear implant by a recipient, meanwhile being trained to process input from the implanted microphone. That said, the trained CED can be utilized to process input from the external microphone during use of the cochlear implant by the recipient, and then an another CED can be trained to process data from the implanted microphone utilizing the results of the processing of the output of the external microphone by the trained CED. In an embodiment, there can be periodic updates to the trained CED (for this or other embodiments). Here, the updates could train the CED to better process the signal from the implanted microphone. That said, to separate CEDs can be utilized, one for processing output from the external microphone and one for processing output from the implanted microphone. The CED for processing output from the external microphone can be utilized to train the CED for processing the implanted microphone data.
[0172] In an exemplary embodiment, again as noted above, the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output. Here, the sound processing path includes a sound processing path of a conventional implantable hearing prosthesis downstream of the product of machine learning. In this exemplary embodiment, the product of machine learning is utilized to pre-process the output from the implanted microphone, and the output thereof is provided to the conventional implantable hearing prostheses processing path. The product of machine learning here is not outputting the electrode sequences for example as was the case in some of the other embodiments herein. Instead, the product of machine learning is outputting a cleaned signal or more accurately, a functionally cleaned signal that has the body noise removed or otherwise minimize or otherwise reduced and / or addresses the transfer function, or otherwise at least improves the result relative to that which would be the case in the absence of addressing the transfer function. The output here is thus an audio signal, which signal can be utilized for example to operate the above conventional microphone apparatus to generate a sound that is recognizable in a manner analogous to the teachings detailed above.
[0173] Note that these embodiments are not limited to a cochlear implant. In some embodiments, the hearing prostheses can be a bone conduction device (e.g., totallyimplantable or partially implantable) or a middle-ear implant (totally implantable or partially implantable). (The tissue stimulating output assembly would thus include by way of example, the vibrator of the bone conduction device or the actuator of the middle ear implant.) Thus, in an embodiment, the sound processing path includes a product of machine learning trained to minimize feedback. This can result from training where processing of a signal without feedback is compared to processing of a signal with feedback and the loss function is to drive the results of processing of the signal with the feedback towards the results of processing of the signal without the feedback. And to be clear, in an exemplary embodiment, with respect to the implantable microphone, the loss function can be to drive the signal or processing of the signal from the implantable microphone towards the signal or processing of the signal from the external microphone.
[0174] Still, embodiments can be applied with respect to a totally implantable cochlear implant. The sound processing path can include a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output and this implant can include a processing path that includes at least a portion of the sound processing path and that extends to the tissue stimulation assembly includes a conventional cochlear implant sound processor. This is an example of utilizing machine learning to adapt to the implantable microphone signal with a conventional cochlear implant.
[0175] In an embodiment where the device under discussion includes a processing path (as distinguished from the sound processing path - more on this in a moment) that includes at least a portion of the sound processing path and that extends to the tissue stimulation assembly is made entirely of a product of machine learning. That is, recall that in the embodiment of the device under discussion, the device includes a sound processing path that receives input from the implanted microphone. The sound processing path at this point is not fully defined. All that is required is that there be a sound processing path and this path receives the input from the implanted microphone. Here, the processing path includes at least a portion of the sound processing path. This processing path must extend to the tissue stimulation output assembly and must be made entirely of a product of machine learning. Recall that the tissue stimulation output assembly can be the electrode driver or can be a device that controls the electrode driver. The output thereto is from the product of machine learning.
[0176] In an embodiment, the tissue stimulation assembly is configured to receive output from the processing path and controls electrical output of the tissue stimulation assembly based directly on the received output. In an embodiment, by way of example, the processing path outputs electrode stimulation sequences to the tissue stimulation assembly or data for such.
[0177] Briefly, it is noted that the action of generating electrode activation sequences indirectly from the received data that is received in method action 1710 detailed above can also be executed in the context of training. In this regard, method 1800 covers an exemplary embodiment of training a DNN by way of example only and not by way of limitation, because the results of the DNN is compared to the results from the conventional processing, and if that conventional processing generates electrode activation sequences indirectly from the received data, that meets method 1800.
[0178] In an embodiment, the cost function that is calculated during the training of the network is influenced only by information that makes it through the cochlear implant (sampled and compressed magnitude envelopes in distinct frequency bands), or otherwise the cochlear implant processing pipeline. This can be significantly less information than the information contained in a raw audio signal. In an embodiment, the information based on a byte to byte comparison, is 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85 or 90% or more or any value or range of values therebetween in 1% increments less than the information contained in the raw audio signal. In an embodiment, by constraining the information space that the neural network needs to map, the convolutional encoder decoder only focuses on relevant information. As a simple example, the CI only encodes information in bands between 150 - 8000 Hz. Any computational resources placed into processing sound information above 8000 Hz or below 150 Hz is wasted. By enforcing an output in the electrode stimulation domain, we enforce that no computational resources are wasted. In an embodiment, the data provided for training is limited to frequencies at and / or above 80, 90, 100, 110, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 210, 220, 230, 240 or 250 or any value or range of values therebetween in 1 Hz increments and at or below 11,000, 10,000, 9,500, 9,000, 8,500, 8000, 7500, 7000, 6500, 6000, 5500 or 5000 or any value or range of values therebetween in 1 Hz increments. In an embodiment, the convolutional encoder-decoder structure can be trained to emulate CI speech processing strategies that would not be feasible on a real-time device. For example, experimental CI speech processing strategies exist that use normal hearing auditory models asan input side. In an embodiment, returning to for example method 1200, the signal processing is a complex signal processing not usable in real time on a Nucleus 5, 6, 7 and / or Nucleus 8 ™ as of October 13, 2023, approved for use / available in any of the jurisdictions detailed herein. That is, in an exemplary embodiment, a processing regime or otherwise a processing strategy that is computationally involved and otherwise unfeasible for use in one commercially available hearing prostheses or otherwise hearing prostheses that would have a market for use in a seven day a week manner, such as the markets for the Nucleus examples given above or otherwise for any BTE based cochlear implant or hearing prostheses for that matter could be utilized to develop the electrode activation sequences. That is, the resulting electrode activation sequences would be, at least in some scenarios, superior, including far superior, then that developed by the processing strategies utilized for the body worn or the commercially available hearing prostheses. Because the DNN is mimicking the functionality of those complicated strategies to achieve a comparable results, but not implementing the strategies, the end result can be achieved or otherwise a result that is close to the end result or otherwise better than that which would result from training utilizing the conventionally commercially available processing strategies can be achieved without having to go through extensive processing or otherwise utilize the computational power that would be the case with the sophisticated strategy. By way of example only and not by way of limitation, in an exemplary embodiment, a processing strategy that would require a high performance personal computer, and otherwise cannot run on a conventional cochlear implant, or more accurately, utilizing the computational power of a conventional cochlear implant, such as any of those detailed herein, could be utilized to develop the electrode activation sequences, etc., for the purposes of training. By way of example only and not by way of limitation, a processing strategy that could be run only on a Dell Alienware ™ with an Intel Core i9 ™ 14thGeneration Intel Core with 32 GB or 64GB of RAM or more processing power can be utilized in some exemplary embodiments to execute the action of signal processing associated with method action 1210. This is something that cannot be done with the electronics of even an advanced behind-the-ear device of a cochlear implant. Accordingly, in an exemplary embodiment, there is a method of obtaining output of signal processing of an input signal utilizing a processing strategy architecture that requires at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1250, 1500, 1750, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 or 10,000 times or more or any value or range of values therebetween in 1 times incrementprocessing power and / or memory than that which is the case with respect to any one or more of the just noted commercially available cochlear implants.
[0179] In an embodiment, the convolutional encoder-decoder can be updated via weight updates. The weights of the convolutional encoder-decoder network could be updated, either through connection to a PC, or via the cloud, in order to change stimulation strategies.
[0180] In an embodiment, for purposes of comparison, the processing strategy could be run on a Dell ™ laptop with an Intel Core i9 Microprocessor with at least a 2.8 GHz clock frequency, at least 16 by 1024 KB L2 cache, at least 22.00 MB L3 cash, a TDP of at least 160 W, a DMI 3.0 I / O bus and a 4 x DDR4-2666 memory. This processing strategy could be run in real time using that computer. Conversely, the processing strategy would not be able to be run in real time on any one or more of the commercially available hearing prostheses noted herein. Alternatively, the processing strategy would take at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1500, 2000, 2500, 3000, 4000 or 5000 times or more or any value or range of values therebetween in 1 times increment longer if run on the commercially available hearing prostheses than the time it would take to run on one or more of the just noted Dell ™ products.
[0181] An exemplary sound processing strategy can be the SAM Coding Strategy, which is detailed in IEEE Transactions on Biomedical Circuits and Systems in the article entitled Making Use of Auditory Models for Better Mimicking of Normal Hearing Processes With Cochlear Implants: The SAM Coding Strategy, authored by Tamas Harczos, Student Member, IEEE, Anja Chilian, and Peter Husar, Senior Member, IEEE, in a manuscript received April 24, 2012 and revised August 10, 2012 and accepted September 10, 2012. This is one of many strategies that can be used to train the DNN that could not be used in real time on / with at least one of the conventional cochlear implants noted herein. Accordingly, in an exemplary embodiment, there is a method of training and / or a DNA and that is trained utilizing a sound processing strategy that is comparable to this just noted strategy or even more computationally taxing by at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 times or any value or range of values therebetween in one time increments.
[0182] Embodiments include noise reduction and / or speech enhancement. Embodiments can include utilizing convolutional encoder-decoders for noise reduction and / or speech enhancement. In a noise reduction and speech enhancement DNN training framework, asource audio waveform (e.g., speech plus background noise) can be transformed to a processed audio waveform, and that processed audio waveform can be compared to a target clean speech audio waveform (e.g., speech without noise) in the loss function. In this embodiment of the convolutional encoder-decoder processing strategy, the electrode activation patterns, etc., at the output of the convolutional encoder decoder are compared to the electrode activation patterns of clean speech. This training framework enables the convolutional encoder-decoder to incorporate noise reduction within its processing.
[0183] Accordingly, in an exemplary embodiment, referring to method 1200, the action of obtaining output of signal processing of an input signal is such that the input signal is an audio input without noise or otherwise a clean audio waveform. In an exemplary embodiment, this can be a clean speech audio waveform. This can be considered an “Audio In” that has base content and only base content. The base content is the sound that one wants to hear, which could be speech etc. Method 1200 is expanded so that there is a second input signal that is an audio input with noise or otherwise an unclean audio waveform. In an exemplary embodiment, this can be clean speech plus noise audio waveform. With respect to figure 21, this can be considered an “Audio In” that has base content plus noise. The noise can be artificially generated or otherwise added. The actor of the method could be very premeditated about the noise so as to provide different noise scenarios for the purposes of training. In this expanded method, the method includes obtaining a second output of signal processing of a second input signal, which would correspond to the output of the CED processing. Here, the second output is electrical stimulation control data for the sensory prostheses. The expanded method further includes providing second data based on the obtained second output to the convolutional encoder-decoder and training the convolution all encoder decoder based on the provided second data. Here, the second data could be the data that is developed with respect to the loss function of minimizing the difference between processing results without noise and with noise. That is, the electrode activation patterns from the typical cochlear implant processing, and the CED, are compared to one another utilizing any comparison regime that will have utilitarian value, where the loss function minimizes difference between the patterns. Note that the action of providing the feedback from the loss function block to the CED represented in figure 21 corresponds to the action of providing data based on the obtained output to the CED and training the CED based on the provided data. Thus, the training includes applying a loss function that minimizes differencesbetween electrode activation patterns of the output and an output of the convolutional encoder-decoder.
[0184] Figure 22 presents an exemplary processing diagram where CDs are utilized for both the base audio in the base plus noise audio. Here, there could be a already trained CED. The CED could have been trained already and otherwise has the capabilities of conventional cochlear implant processing. By way of example only and not by way of limitation, any of the teachings detailed herein with respect to training a CED can have already been executed. Thus, the audio in is provided to this already trained CED. After that, the training and other actions associated with figure 22 can correspond to those of figure 21 just described. Here, in an exemplary embodiment, there is no conventional cochlear implant signal processing utilized. And note while there are two CED blocks presented, this very well could be the same CED, where the CED trains itself. For example, the CED could output a first data set based on the audio in with only the base content. Then the CED could output a second data set based on the audio in with the base plus noise. That said, in an embodiment, it could be that the CED could output the different outputs at the same time. In any event, the comparison between the two can be made and so on.
[0185] To round things out, figure 23 and 24 provide additional exemplary conceptual processing flows. There can be utilitarian value with respect to comparing the typical cochlear implant processing for the base signal plus and the base plus noise signal. There could also be utilitarian value with respect to comparing the results of the CED in the different scenarios. Indeed, this could be a way of simultaneous training for noise and non- noisy signals. Note that the concept of figure 23 and / or figure 24 could also be applicable to the implanted microphone by way of example, as will be modified accordingly to implement the training profit herein.
[0186] In view of the above, it can be seen that in an embodiment, there is a hearing prosthesis where the developed electrode control data has reduced noise content. That is, in an embodiment where the product of machine learning processes input and automatically develops electrode activation control data, the outputted control data has reduced noise content relative to that which would otherwise be the case. In an embodiment, the product is configured to process the input and automatically reduce noise content in the input in the development of the electrode activation control data. In an embodiment, the product automatically reduces noise content at the same time that it automatically develops the electrode activation control data.
[0187] Referring back to method 1700, in an exemplary embodiment, the received data can be based on ambient sound speech and distracting noise to a user of a hearing prosthesis that generates the electrode activation sequences. In an embodiment, the method can include automatically removing noise content from the received data when generating the electrode activation sequences. This is contrasted to, for example, removing the noise content when developing a first data set and then utilizing that data set to develop the electrode activation sequences. In an exemplary embodiment, the amount of noise that is in the data can be less than, greater than and / or equal to 110, 105, 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 dB, or any value or range of values therebetween in one dB increments. In an exemplary embodiment, the noise reduction results in a reduction by at least and / or equal to and / or no more than 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 95 or 90% or more or any value or range of values therebetween in 1% increments. In an exemplary embodiment, the training for noise reduction includes training with noise at any one or more of the aforementioned noise levels. In an exemplary embodiment, the noise can be a pure sine wave added to speech content.
[0188] In an exemplary embodiment, there is no transitory and no non-transitory signal or data between the action of receiving data and the action of generating the electrode activation sequences that has more noise content than elsewhere between the action of receiving data and generating the electrode activation sequences.
[0189] Referring back to the system detailed above that utilizes a convolutional based subsystem, in an embodiment, the system is a sound processor system and the convolutional based subsystem is trained to reduce noise in the received input. In an embodiment, the training is based on a loss function of minimizing a difference between electrode activation patterns that respectively result from input having desired base content and input having the desired base content plus noise content.
[0190] Again, referring to the above, there is a learned filter bank that can filter frequencies of incoming sound or incoming light, while maintaining a coupling between the magnitude of those frequencies and the phase of those frequencies. This can have utilitarian value with respect to obtaining a quality of noise reduction beyond that which would otherwise be the case, such as where the phase and magnitude are decoupled. In an exemplary embodiment, the quality of noise reduction is such that at least and / or equal to 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70% or more or any value or range of values therebetween in 1% increments more noise reduction can be achieved utilizing the teachings detailed hereinversus a conventional filter bank arrangement such as that found in any of the hearing prostheses detailed herein by way of example with respect to the average amount (mean and / or median) of noise reduction with respect to normal usage for a day or a week or a month for example, and / or with respect to a controlled test where a controlled noise is generated simultaneously with a controlled speech pattern for example that would be recognized as providing an adequate test of the noise cancellation subsystem, all other things being equal.
[0191] In an exemplary embodiment, this can result in relatively high quality speech content perception based on the output relative to that which would otherwise be the case. In an exemplary embodiment, the speech is at least and / or equal to 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 80, 90, 100, 125, 150, 175 or 200% or more or any value or range of values therebetween in 1% increments more clear utilizing the teachings detailed herein versus a conventional filter bank arrangement such as that found in any of the hearing prostheses detailed herein by way of example with respect to the average amount (mean and / or median) of normal usage for a day or a week or a month for example, and / or with respect to a controlled test where a controlled noise is generated simultaneously with a controlled speech pattern for example that would be recognized as providing an adequate test of the noise cancellation subsystem, all other things being equal.
[0192] Returning back to the concept of keeping phase data coupled to the magnitude, in an exemplary embodiment, this can have utilitarian value because the processing regimes detailed herein utilize it phase info to extract information about what is noise and what is not noise, and what is the target sound. Put another way, in an exemplary method, and otherwise in an exemplary device and / or system that implements such method, there is the action of obtaining input based on an ambient sound and automatically processing that input based on the teachings detailed herein to utilize the phase content that is coupled to the magnitude of the pertinent portions of the sound input to reduce noise and / or identify a target sound that is to be targeted at the exclusion of other sounds. In an exemplary embodiment, the target sound is enhanced and the nontarget sounds are reduced (which includes eliminated). In an exemplary embodiment, this is done at least in part based on data where the magnitude and phase are not decoupled. And note that in an exemplary embodiment, this can occur in what would be considered the output side processing as opposed to the input side processing.
[0193] FIG. 25 shows an exemplary block diagram of a multi -mi crophone noise reduction system, which takes signals from multiple microphones as inputs, and generates denoisedelectrode stimulation sequences or data therefor for the respective electrode arrays for the left and right ears. This concept can also be applied to a bionic eye for example, where the input would be light from a left and right camera, etc.
[0194] More particularly, embodiments include multi -mi crophone speech enhancement and spatial filtering. In an example of multi -mi crophone speech enhancement, the convolutional encoder-decoder receives several audio waveforms at its input. This is the respective microphone signals and a normalized cross correlation (NCC) between the microphone signals. These are encoded by the convolutional encoder to the latent space, where the input is processed by the separator. In this embodiment, the separator calculates a mask to remove noise in the latent space. The separator could be implemented using a wide variety of neural network architectures (not just the CED, as is the case with all embodiments herein unless otherwise noted, providing that the art enables such). The latent space is then decoded by transposed convolutional layers to electrode activation sequences for the left ear and the right ear. This is an example of binaural noise reduction, which applies spatial filtering to remove spatially separated noise.
[0195] Briefly, the idea behind the schematic of figure 25 is that there is a recipient that has a binaural hearing prostheses (sometime referred to as a bilateral system) in general, and a binaural cochlear implant system in particular, where there is a cochlear implant implanted in the left side cochlea and a cochlear implant implanted in the right side cochlea of the human’s head. The two implants include respective behind-the-ear devices or off the ear devices attached to the left and right side of the person’s head. Each of these two components have respective microphones (one or more microphones). In an exemplary embodiment, the cochlear implants “talk” with each other, or more accurately communicate with each other, and develop stimulation based on input from the other cochlear implant. Here, for example, a cochlear implant receives the input from its own microphones, and the microphones of the other cochlear implant. This can be done utilizing radiofrequency communication or wired communication. Bluetooth commit indication and / or MI radio can be utilized. Any device system and / or method that can enable communication of the data based on sound captured by given microphone to the other component of the prostheses can be utilized in at least some exemplary embodiments, providing that the art enables such, unless otherwise noted.
[0196] With this in mind, with respect to FIG. 25, the “left front” and the “left rear” are the left front microphone and the left rear microphone (the microphones on the left side of a human’s head, where for example a BTE or an OTE has two or more microphones). Theright front and right rear correspond to the microphones on the right side. The NCC LF-RF is the normalized cross correlation between the left front and the right front microphones. The NCC LR-RR is the normalized cross correlation between the left rear and right rear microphones. The NCC LF-LR is the normalized cross correlation between the left front and left rear microphones, and the NCC RF-RR is the normalized cross correlation between the right front and right rear microphones. Note that the microphones could also be implanted microphones and this could be a totally implantable cochlear implant by way of example
[0197] The product of artificial intelligence or otherwise machine learning can be located in one or both of the BTE’s or OTE’s of the respective cochlear implants. It could be that the two components share resources, where communication between the two is quick enough and the band with his large enough to enable division of the computational actions between the two cochlear implants.
[0198] Embodiments can utilize the NCC data as a way of enforcing timing differences between two signals. Such can reduce inefficiency and otherwise enforce feature recognition. This can enable feature extraction or otherwise the use of features which are not available with respect to conventional sound processing or even the utilization of products of machine learning to analyze only one set of data from a single microphone as opposed to microphones that are spatially separated or otherwise located on the opposite sides of a person’s head, such as here. In an exemplary embodiment, the microphones are spaced at least and / or equal to 1, 1.5, 2, 2.5, 3, 3.5, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65 or 70 mm or more or less than and / or equal to or greater than 100, 150, 200, 250, 300, 350, 400 or 450 mm from each other more or any value or range of values therebetween in 0.1 mm increments (in that no part of the structure of the microphone is located within those distances from the structure of another microphone).
[0199] In an embodiment, the separator network can use two or more GRU layers. The encoder and decoder can be both 4 layers. In the encoder, the number of kernels can increase from 16 to 128 by powers of two in each consecutive layer, and in the decoder, the kernels can decrease from 64 to 16 in powers of two, before generating 2 x (# of electrodes) outputs for electrode activation sequences, all by way of example. FIG. 26 shows an exemplary block diagram of a multi -mi crophone noise reduction CDD according to an exemplary embodiment. This can use any one or more of the teachings herein to achieve enablement to achieve a real-time or near real time system.
[0200] Thus, as can be seen, in an embodiment, there is a product of machine learning that executes noise reduction in the same network as that which is utilized to output the electrode activation sequences, or otherwise the data from which the electrode activation sequences can be directly developed. This is concomitant with any of the embodiments herein, or any of the functionalities detailed herein associated with the product of machine learning or otherwise executed (functionally executed, if only by mimicking or “working” on an input to produce an output based on learned attributes that may be or are different from how the conventional sound processing operates), where such functionalities are executed by the same network as that which is utilized to output the electrode activation sequences or data based thereon or the data that is then used to produce the electrode activation sequences.
[0201] Note that the embodiment of figures 24 and 25, while introduced in the concept of a binaural cochlear implant system, can also be executed or otherwise implemented utilizing a single-sided / uniaural cochlear implant system. Such a system could have two or more microphones, such as that utilized for beamforming for example. The system could implement the teachings associated with figures 24 and 25 except that there would be no comparisons or otherwise data from the microphones of the other side of the head. Note also that the embodiments of figures 24 and 25 could also be implemented in a bilateral arrangement where the two cochlear implants or to hearing prostheses are separate and they do not communicate with one another. This would effectively be the same as the aforementioned unilateral cochlear implant system that has two or more microphones, such as two or more microphones that are part of a behind the ear device (BTE device) and / or part of the off the ear device (OTE device) sound processor by way of example.
[0202] In embodiments, the data that is utilized to extract spatial information (or more accurately, to functionally mimic the extraction of spatial info concomitant with the teachings detailed herein), is based on the data where the phase is linked to the magnitude as noted above. In an embodiment, the teachings detailed herein, or more accurately, the products of machine learning utilize the phase difference components of the signal to determine where certain sounds are coming from and where certain sounds are not coming from. By way of example only and not by way of limitation, in an exemplary embodiment, the teachings detailed herein are utilized to enhance sounds that emanate from in front of the recipient. This can be a proxy for speech, where, for example, typically, a human being looks to or towards the person who is speaking to him or her. This can be especially the case with respect to a person with hearing deficiencies, because such person could attempt to alsoattempt to “read lips” or otherwise follow the body expressions of the speaker to further extrapolate what the speaker is saying.
[0203] In view of the above, referring to the above exemplary systems, in an exemplary embodiment, there is an input subsystem that is part of the system that is configured to receive data based on sound captured by a plurality of separate sound capture devices. In an exemplary embodiment, there can be two, three, four, five, six, seven, eight, nine, 10, 11 or 12 or more or any value or range of values therebetween in one increment sound capture devices, such as by way of example, microphones. These can be spaced between two or more hearing prostheses, such as a binaural hearing prostheses system, or could be part of the same hearing prostheses, one or more or all of which are supported by the BTE or the OTE device thereof. In this exemplary embodiment, there is a convolutional based subsystem that executes noise reduction based comparison of respective data based on sound captured by the respective sound capture devices, or at least a subset of the respective sound capture devices. In an exemplary embodiment, the convolutional subsystem executes speech enhancement and spatial filtering based on comparison of the respective data based on sound captured by the respective sound capture devices. Note that the data that is provided to the input subsystem could be the raw signals from the microphones or could be pre-processed signals, and also can be data sets that include data corresponding to the input from the microphones as well as normalized cross correlation data thereof. Again, it could be that the convolutional based subsystem performs the normalization cross correlation between the various microphone signals. Accordingly, in an exemplary embodiment, the received data based on sound captured by a plurality of separate sound capture devices is contained in a single signal or in a plurality of singles, and there could be other signals are there can be additional data in the signals, such as the normalization data by way of example. The received data of course can include the normalized cross correlation data for spatially distinct sound capture devices or the convolutional based subsystem normalizes the crosscorrelation of spatially distinct sound capture devices. With respect to the methods detailed herein, such as by way of example the method where the electrode activation sequences are generated directly from received data, this received data could include the audio content including speech sub content from at least two different microphones of the hearing prostheses, which microphones capturing ambient sound, which ambient sound would include the speech of course. Further, the action of generating electrode activation sequences includes enhancing the speech sub content. This as opposed to, for example, enhancing the speech sub content prior to generating the electrodeactivation sequences. Put another way, the electrode activation sequences that are generated are generated to enhance the speech content (and reduce noise depending on the embodiment) as opposed to being generated to simply create electrode activation sequences based on data speech content has already been enhanced.
[0204] In an exemplary embodiment, the speech sub content is enhanced by at least and / or equal to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 dB or more or any value or range of values therebetween in 0.1 dB increments, and corollary to this is that the noise sub content of the audio content is reduced by any one or more of those values just detailed.
[0205] With respect to the products of machine learning detailed herein, in an exemplary embodiment, the product is configured to receive input based on captured phenomenon in the audio environment in accordance with the teachings herein. In this exemplary embodiment, the captured phenomenon can include audio content including speech sub content from at least two different microphones of a hearing prosthesis (which could be the binaural system or a unilateral system for example), which microphones capture the captured phenomenon. Here, the product of machine learning can include a separator (at least functionally concomitant with the teachings herein) configured to calculate a mask to reduce noise from the audio content in the latent space of the product.
[0206] Further with respect to the various products of machine learning, in an embodiment, the separator is configured to enhance the speech sub content.
[0207] In an exemplary embodiment, the product of machine learning does not output frequency divided signals (e.g., different channels having different frequency bands so that collectively the output covers a wide range of the frequency that is input - for example, 10 frequency channels each of which have a width of 500 Hz collectively making up 5000 Hz of the frequency content of the input (in reality, the frequency bins will be such that certain frequency ranges will be more concentrated than others because certain frequency ranges are more important than others, such as the frequency ranges that contain the frequencies that are normally associated with speech versus higher frequencies for example (above 5000 Hz for example))) and also does not even output channel envelopes. Again, in an exemplary embodiment, the output is more processed than that which results from enveloping and otherwise directed towards the unique features associated with a cochlear implant or a bionic eye. That is, for example, frequency binning and channel enveloping are actions that can be utilized in other types of hearing prostheses, such as by way of example, conventionalhearing aids or even a bone conduction devices or hearing prostheses that utilize electromagnetic actuators or piezoelectric actuators, or otherwise output devices that are not channel specific for example. Again, the output is an output that could not be utilized by a conventional monitor or a conventional speaker, even if processed utilizing like datatypes (adding a carrier signal such as a sine wave to the outputs of the envelopes). In at least the exemplary embodiments for a cochlear implant or a bionic eye, the output is an output that is solely usable by those types of prostheses, or at least a prostheses that stimulates utilizing electrical output as opposed to mechanical output pressure or force output (such as that which is the case with respect to a conventional hearing aid). Indeed, output corresponding to channel envelopes could be usable by a microphone, albeit the output would be nonsensical, hence why the application of a sine wave to the envelopes to create the negative portion of the signal would then result in a more meaningful result.
[0208] To be clear, in at least some exemplary embodiments, the decoder is not converting the latent space to frequency channels and / or channel envelopes. There is no separate process outside of the product of machine learning in general, and the convolutional encoder decoder in particular, that processes any channel envelopes so as to ultimately obtain the electrode activation sequences or otherwise data that corresponds to such.
[0209] Indeed, in an exemplary embodiment, there are no channel envelopes that can be identified in the sound processor utilizing at least some of the teachings detailed herein. In an exemplary embodiment, there are not frequency bins that can be identified in the sound processor utilizing at teach some of the teachings detailed herein. With respect to figure 10, it could be that there is another sound processor that has such (there, the alternate path) but that is different than a sound processor utilizing the product of machine learning herein that outputs data corresponding to the electrode activation sequences.
[0210] In an exemplary embodiment, any denoising that occurs and / or any speech enhancement or spatial filtering processing that occurs happen simultaneously with the development of the electrode activation sequence data.
[0211] In view of the above, it can be seen that in an exemplary embodiment, there is a product of machine learning, as well as methods and / or systems, where signal processing functionality such as speech enhancement, noise reduction, auditory scene classification, automatic gain control, spatial beamforming, binaural processing and / or conversion of data to stimulation pulse trains for electrodes of the cochlear implant occurs outside the time-frequency domain. In an exemplary embodiment, there is no conversion of audio signals or otherwise sound signals to the time-frequency domain. Again, in an exemplary embodiment, the sound processing strategy does not utilize Fourier transformation. In exemplary embodiments, the shortcoming of traditional sound processing where the output of various filters in the filter bank has the same temporal context is avoided. This enables the utilization of both global and local features. Moreover, the number of features that can be addressed in the sound processing is not limited by the number of channels in a filter bank, because there is no filter bank, at least not in the product of machine learning. For example, there are no limitations vis-a-vis the number of FFT bins and / or the number of implanted cochlear implant electrodes or otherwise active electrodes implanted in the cochlea for example. Embodiments can effectively result in the use of more features than sound processors that are so limited. By way of example only and not by way of limitation, there can be more than and / or at least equal to 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 1500 or 2000 or more features or any value or range of values therebetween in one feature increment as opposed to, for example, less than 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 features or any value or range of values therebetween in one feature increment associated with the conventional sound processing. Here, in an exemplary embodiment, there are more features that are inputted into the signal processing algorithms than those of the conventional system.
[0212] In an exemplary embodiment, there is automated feature analysis where the analysis of different features is for all intents and purposes executed simultaneously. This as opposed to, for example, conventional sound processing. In an exemplary embodiment, there is effectively simultaneous evaluation of a number of features corresponding to any one or more of the above just noted number of features.
[0213] Corollary to this is that in at least some exemplary embodiments, there is parameter sharing within the trained neural network.
[0214] Embodiments can enable input features that are data driven, as opposed to hardcoded arrangements with respect to that which is the case for conventional systems. Embodiments can thus utilize neural network optimized hardware accelerators, which can enable the adoption of data driven algorithms on portable devices, such as cochlear implants. Embodiments can utilize learned feature representation on edge devices, as opposed to fixing the feature representation through hardcoded filter banks.
[0215] Further with regards to the ability to utilize global features, real hearing in a human and mammals for that matter will treat the same sound loudness differently over time (over seconds). For example, a sound can be perceived initially as having a loudness that is louder than that which is perceived by the human after the onset of the sound (e.g., after 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.75, 1, 1.25, 1.5, 1.75 or 2 seconds or any value or range of values therebetween in 0.01 second increments by way of example, depending on the mammal and the age thereof, etc.). The point is, that a sound having magnitude X dB will initially be perceived at Y dB and then less than Y dB even though X has not changed. In an exemplary embodiment, depending on the initial loudness and the time frame from the onset, the decrease in perceived loudness could be less than greater than and / or equal to 5, 10, 15, 20, 25 or 30% or any value or range of values therebetween in 1% increments. In a conventional sound processing system, this decrease in loudness will not be experienced by the recipient of the hearing prostheses utilizing such. Indeed, this is because, in at least some processing systems, of the discrete nature of the signal processing, where the processing occurs in tiny segments on the order of one to five millisecond range for example. The processing does not account for past history or otherwise take into account the past history. The processing also does not predict the future or otherwise estimate what the future will be based on the past, and thus does not take into account that possible future with respect to the processing.
[0216] Conversely, the teachings detailed herein utilize a more holistic approach. The processing can take into account the past, and the processing can also estimate or otherwise attempt to predict the future. By way of example only and not by way of limitation, a sound that is extending in loudness would be predicted to continue the ascension of the loudness. The teachings detailed herein can predict the future loudness or at least attempt to predict the future loudness and take that into account with respect to the processing of the sound at the current time. Further by way of example only and not by way of limitation, a sound that has plateaued or otherwise a sound that has an additional occurrence that a sharp and maintains the loudness thereof for at least 5 or 7 hundred ms or longer, will be processed by a conventional sound processing system so that the output loudness is perceived at the same level for that whole time. Conversely, the processing in accordance with some of the embodiments detailed herein utilizing the product of machine learning will process that sound so that the perceived loudness is reduced over time as would be the case with a person withnormal hearing for example. This is an example of the use of global features as opposed to the mere local features that conventional sound processing utilizes.
[0217] In an exemplary embodiment, reference to local features can be considered features that have a timeframe less than and / or equal to 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 ms or any value or range of values therebetween in 0.1 ms increments. Global features could be features that have a timeframe greater than and / or equal to any of those values or greater, such as 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 10, 175, 200, 225, 250, 275, him 300, 325, 350, 375, 400, 450, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000, 2500, 3000, 3500 ms or more or any value or range of values therebetween in 1 ms increments. In an embodiment, the processing relies on both global and local features. This is contrasted to, for example, the conventional cochlear implant system processing that only relies on local features (at least typically). In an exemplary embodiment, by way of example, the product of machine learning takes into account history of the sound and / or light or otherwise the history of the processing of the sound and / or light going back any one or more or all of the just detailed time periods (for example, the DNN can develop an electrode activation sequence for time T = X based on something that happened or based on a feature that existed or otherwise the processing or determination made by the DNN 700 ms (one of the specified time periods just noted) before time T = X. In some embodiments, this can potentially be a way to correct a previous improper or less than utilitarian output. For example, if the output of the DNN that occurred 100 ms before time T = X, and the DNA and determines that that was not the exact output or otherwise was not the best output that could have been provided based on the audio environment at that time or before or even after for that matter, the DNA and could then adjust the output at time T = X and times thereafter for that matter to attempt to compensate for that prior output. Accordingly, embodiments include the utilization of the DNN in that self corrects based on a self-evaluation of prior input in view of a holistic evaluation of the ultimate goal with respect to the hearing percept that is to be achieved based on the sound input into the prostheses.
[0218] And note that the use of global features is something that happens all the time, or more accurately, the ability to use global features is something that happens all the time if such features are available. It is a part of the processing. This as opposed to, for example, an algorithm that is hardcoded that perhaps utilizes some form of historical basis to adapt the processing, but only does so to a limited extent and otherwise based on a very limited set of features.
[0219] Accordingly, embodiments can be trained to utilize patterns and / or rhythms to influence processing in real time. Accordingly, in an exemplary embodiment, the processing detailed herein can be based on a recognized, albeit by a machine learning or otherwise by a product of machine learning, pattern or rhythm or trend in the ambient sound and / or the ambient light.
[0220] D
[0221] Embodiments can include a system and / or simply an embodiment that includes a non- transitory computer readable medium having recorded thereon, a computer program for executing at least a portion of a method, the computer program including code for executing any one or more of the method actions and / or functionalities detailed herein. Thus, any disclosure herein of a method action or functionality corresponds to a disclosure of a non- transitory computer readable medium having programed thereon code to execute one or more of those actions and also a product to execute one or more of those actions.
[0222] Embodiments include any functionality disclosed herein and / or method action disclosed herein being executed by a computer chip, a processor, software, logic circuitry and / or electronics, and all are not mutually exclusive. Any circuit that can enable the teachings herein can be used providing that the art enables such. Thus, in the interests of textual economy, and disclosure herein of a functionality of an article of manufacture corresponds to any one or more of the aforementioned structures being configured to execute such and otherwise for such, and the same is for any method action disclosed herein, where any such action corresponds to a disclosure of any one or more of the aforementioned structures being configured to execute such and otherwise for such.
[0223] In some embodiments, a neural network, such as a DNN, is used to directly interface to the input into the systems / devices detailed above, and process this input via its neural net, and determine the information detailed above. The network can be, in some embodiments, either a standard pre-trained network where weights have been previously determined (e.g., optimized) and loaded onto the network, or alternatively, the network can be initially a standard network, but is then trained to improve specific recipient results based on outcome oriented reinforcement learning techniques.
[0224] Any disclosure herein of a processor corresponds to a disclosure in an embodiment of a non-processor device or a combined processor-non-processor device where the nonprocessor is a result of machine learning. Embodiments can include a link from the cloud toa clinic to pass information back and forth, enabling the remote processing noted above and / or enabling the obtaining of additional data for retraining purposes. Information can be uploaded to the cloud to the clinic, where the information can be analyzed. Another exemplary system includes a smart device, such as a smart phone or tablet, etc., that is running a purpose built application to implement some of the teachings detailed herein. Any disclosure herein of a processor corresponds to a disclosure of a non-processing device, or includes non-processing devices, such as a chip or the like that is a result of a machine learning algorithm or machine learning system, etc.
[0225] Reference herein is frequently made to the recipient of a hearing prosthesis. It is noted that in at least some exemplary embodiments, the teachings detailed herein can be applicable to a person who is not the recipient of a hearing prosthesis. Accordingly, for purposes of shorthand, at least some exemplary embodiments include embodiments where the disclosures herein directed to a recipient correspond to a disclosure directed towards a person who is not a recipient but instead is only hard of hearing or otherwise has a hearing ailment and is contemplating obtaining a hearing assistance device.
[0226] Any method action and / or functionality disclosed herein where the art enables such corresponds to a disclosure of a code from a machine learning algorithm and / or a code of a machine learning algorithm and / or a product of machine learning for execution of such. Still as noted above, in an exemplary embodiment, the code need not necessarily be from a machine learning algorithm, and in some embodiments, the code is not from a machine learning algorithm or the like. That is, in some embodiments, the code results from traditional programming. Still, in this regard, the code can correspond to a trained neural network. In an embodiment, the trained neural network can be utilized to provide (or extract therefrom) an algorithm that can be utilized separately from the trainable neural network. In one embodiment, there is a path of training that constitutes a machine learning algorithm starting off untrained, and then the machine learning algorithm is trained and “graduates,” or matures into a usable code - code of trained machine learning algorithm. With respect to another path, the code from a trained machine learning algorithm is the “offspring” of the trained machine learning algorithm (or some variant thereof, or predecessor thereof), which could be considered a mutant offspring or a clone thereof. That is, with respect to this second path, in at least some exemplary embodiments, the features of the machine learning algorithm that enabled the machine learning algorithm to learn may not be utilized in the practice someof the method actions, and thus are not present the ultimate system. Instead, only the resulting product of the learning is used.
[0227] And to be clear, in an exemplary embodiment, there are products of machine learning algorithms (e.g., the code from the trained machine learning algorithm) that are included in any one or more of the systems / subsystems detailed herein, that can be utilized to analyze any of the data obtained or otherwise available disclosed above that can be utilized or otherwise is utilized to evaluate the data obtained herein. This can be embodied in software code and / or in computer chip(s) that are included in the system(s).
[0228] An exemplary system includes an exemplary device / devices that can enable the teachings detailed herein, which in at least some embodiments can utilize automation. That is, an exemplary embodiment includes executing one or more or all of the methods and / or functionalities detailed herein and variations thereof, at least in part, in an automated or semiautomated manner using any of the teachings herein. Conversely, embodiments include devices and / or systems and / or methods where automation is specifically prohibited, either by lack of enablement of an automated feature or the complete absence of such capability in the first instance.
[0229] At least some of the actions herein can be practiced and otherwise represented by an algorithm where circuitry receives the input (embodied in an analogue or a digital signal), where the input suite converts the “physical” input into electronic signals using analog to digital converters for example, or in the case of the input suite corresponding to an Internet server, receives the digital signal from a remote location, and the digital data is stored in a memory and / or received by the electronics. The electronics, which is a result of the machine learning, takes the digital signal and deconstructs the digital signal to evaluate properties, and then, using its “knowledge” from its training, provides an output corresponding to the prediction. By analogy, the operation is analogous to how a human being “predicts” how he or she will function if he or she foregoes a meal for example, or stays up all night, or if he or she drinks 5 cups of coffee in one hour. Past experience informs the future results, the prediction
[0230] It is further noted that any disclosure of a device and / or system detailed herein also corresponds to a disclosure of otherwise providing that device and / or system and / or utilizing that device and / or system.
[0231] It is also noted that any disclosure herein of any process of manufacturing or providing a device corresponds to a disclosure of a device and / or system that results therefrom. Is also noted that any disclosure herein of any device and / or system corresponds to a disclosure of a method of producing or otherwise providing or otherwise making such.
[0232] An exemplary system includes an exemplary device / devices that can enable the teachings detailed herein, which in at least some embodiments can utilize automation, as will now be described in the context of an automated system. That is, an exemplary embodiment includes executing one or more or all of the methods detailed herein and variations thereof, at least in part, in an automated or semiautomated manner using any of the teachings herein.
[0233] Any embodiment or any feature disclosed herein can be combined with any one or more or other embodiments and / or other features disclosed herein, unless explicitly indicated and / or unless the art does not enable such. Any embodiment or any feature disclosed herein can be explicitly excluded from use with any one or more other embodiments and / or other features disclosed herein, unless explicitly indicated that such is combined and / or unless the art does not enable such exclusion.
[0234] Any function or method action detailed herein corresponds to a disclosure of doing so an automated or semi-automated manner.
[0235] While various embodiments of the present invention have been described above, it should be understood that they have been presented by way of example only, and not limitation. It will be apparent to persons skilled in the relevant art that various changes in form and detail can be made therein without departing from the spirit and scope of the invention.
Claims
CLAIMSWhat is claimed is:
1. A system, comprising: an input subsystem configured to receive input; and a convolutional based sub-system in signal communication with the input subsystem, wherein the convolutional based sub-system outputs electrode activation sequence based data.
2. The system of claim 1, wherein the system is an ambient environment stimulus processor system.
3. The system of claims 1 or 2, wherein the system is a sound processor system.
4. The system of claims 1, 2 or 3, wherein: the input subsystem is configured to receive data based on sound captured by a sound capture device; and the convolutional based subsystem outputs: electrode activation sequences; or data sets convertible to electrode activation sequences.
5. The system of claims 1, 2, 3 or 4, wherein: bottleneck layers of the convolutional based subsystem develop current levels for electrode channels based on data received at the input subsystem.
6. A cochlear implant, comprising: a sound capture device; the system of claims 1, 2, 3, 4 or 5; and an electrode array, wherein the cochlear implant provides data based on sound captured by the sound capture device to the input end, wherein the input end provides data based on the provided data to the convolutional based sub-system, andthe convolutional based sub-system provides output based on the data provided to the convolutional based subsystem, wherein the cochlear implant is configured to directly control stimulation by the electrode array to stimulate tissue to evoke a hearing percept based on the data output by the convolutional based subsystem.
7. A cochlear implant system, comprising: one or a plurality of sound capture devices; the system of claims 1, 2, 3, 4 or 5; and an electrode array, wherein the cochlear implant system provides data based on sound captured by the plurality of sound capture devices to the input subsystem, wherein the input subsystem provides data based on the provided data to the convolutional based subsystem, and the cochlear implant system controls the electrode array to evoke a hearing percept in response to the sound captured by the plurality of sound capture devices while maintaining a coupling between magnitude and phase thereof.
8. The system of claims 1, 2, 3, 4 or 5, wherein: the system includes a sound coding path for a cochlear implant; and the path is entirely made up of a convolutional encoder-decoder network.
9. The system of claims 1, 2, 3, 4, 5 or 8, wherein: the convolutional based sub-system is a trained sub-system trained at least in part with a loss function that minimizes difference between data based on a signal from an external microphone and data based on a signal from an implanted microphone.
10. A cochlear implant, comprising: a sound capture device; the system of claims 1, 2, 3, 4, 5, 8 or 9; and an electrode array, wherein the cochlear implant provides data based on sound captured by the sound capture device to the input subsystem, wherein the input subsystem provides data based on the provided data to the convolutional based subsystem, andthe cochlear implant is configured to stimulate tissue to evoke a hearing percept based on the electrode activation sequence based data outputted by the convolutional based subassembly, wherein there is no dedicated circuitry solely dedicated to electrode activation sequencing.
11. The system of claims 1, 2, 3, 4, 5, 8 or 9, wherein the system is a sound processor and wherein the convolutional based subsystem is trained to reduce nose in the received input, the training being based on a loss function of minimizing a difference between electrode activation patterns that respectively results from input having a desired base content and input having the desired base content plus noise content.
12. The system of claims 1, 2, 3, 4, 5, 8, 9 or 11, wherein: the input subsystem is configured to receive data based on sound captured by a plurality of separate sound capture devices; the convolutional based subsystem executes noise reduction based on comparing respective data based on sound captured by the respective sound capture devices.
13. The system of claims 1, 2, 3, 4, 5, 8, 9, 11 or 12, wherein: the input subsystem is configured to receive data based on sound captured by a plurality of separate sound capture devices; the convolutional based subsystem executes speech enhancement and spatial filtering based on comparing respective data based on sound captured by the respective sound capture devices.
14. The system of 13, wherein at least one of: the received data based on sound captured by a plurality of separate sound capture devices is contained in a single signal or a plurality of signals; the received data includes normalized cross-correlation data for spatially distinct sound capture devices; or the convolutional based subsystem normalizes cross-correlation of spatially distinct sound capture devices.
15. A device comprising: a product of machine learning, whereinthe product is configured to receive an input based on a captured phenomenon in the visual and / or audio environment, and the product is configured to process the input and automatically develop electrode activation control data.
16. The device of claim 15, wherein: the device is a sensory prosthesis that includes electrodes; and the device is configured to activate the electrodes based on the developed electrode activation control data to evoke a sensory percept.
17. The device of claims 15 or 16, wherein: the device is a cochlear implant; and there is no processing in the signal processing pipe of the device after the product.
18. The device of claims 15, 16 or 17, wherein: the device is a cochlear implant with at least 10 electrode channels; and the electrode activation control data is dimensionalized relative to the number of channels of the at least 10 electrode channels.
19. The device of claims 15, 16, 17 or 18, wherein: the device is a hearing prosthesis; and the electrode control data is devoid of audio content.
20. The device of claims 15, 16, 17, 18 or 19, wherein: the device is a hearing prosthesis; and the developed electrode control data has reduced noise content.
21. The device of claims 15, 16, 17, 18, 19 or 20, wherein: the device is a hearing prosthesis; and the developed electrode control data has phase content.
22. The device of claims 15, 16, 17, 18, 19, 20 or 21, wherein: at least some input side processing and all output side processing is executed by the product of machine learning.
23. The device of claims 15, 16, 17, 18, 19, 20, 21 or 22, wherein: the product receives raw audio and outputs electrode stimulation sequences.
24. The device of claims 15, 16, 17, 18, 19, 20, 21, 22 or 23, wherein: the product receives data based on raw audio and outputs electrode stimulation sequences.
25. The device of claims 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24, wherein: the product is configured to process the input and automatically reduce noise content in the input in the development of the electrode activation control data.
26. The device of claims 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25, wherein: the product automatically reduces noise content at the same time that it automatically develops the electrode activation control data.
27. The device of claims 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or 26, wherein: the product is configured to receive an input based on a captured phenomenon in the audio environment; captured phenomenon includes audio content including speech subcontent from at least two different microphones of a hearing prosthesis, which microphones capture the captured phenomenon; the product includes a separator configured to calculate a mask to reduce noise from the audio content in a latent space of the product.
28. The device of claims 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 or 27, wherein: the product is configured to receive an input based on a captured phenomenon in the audio environment; captured phenomenon includes audio content including speech subcontent from at least two different microphones of a hearing prosthesis, which microphones capture the captured phenomenon; the product includes a separator configured to enhance the speech subcontent.
29. A method, comprising:receiving data based on at least one of ambient sound and / or ambient light in an ambient environment of a device or electromagnetic signals received by the device, the data enabling direct usage thereof by a conventional speaker apparatus and / or a conventional video monitor apparatus to produce a recognizable corresponding sound and / or image without further processing; and generating electrode activation sequences directly from the received data.
30. The method of claim 29, further comprising: generating electrode activation sequences indirectly from the received data.
31. The method of claim 29 or 30, wherein: the action of generating electrode activation sequencies is executed in real time vis-a- vis the action of receiving data; and the method further comprises evoking a hearing percept based on the generated electrode activation sequences.
32. The method of claims 29, 30 or 31, wherein: the actions of receiving data and generating electrode activation sequences are executed by a body -worn cochlear implant without converting data based on the received data to a Fourier domain.
33. The method of claims 29, 30, 31 or 32, wherein: the action of generating electrode activation sequences is executed by a product of a trained neural network.
34. The method of claims 29, 30, 31, 32 or 33, wherein: all backend processing is executed by data-driven developed algorithms.
35. The method of claims 29, 30, 31, 32, 33 or 34, wherein: the actions of receiving data and generating electrode activation sequences are executed by a body -worn cochlear implant; and the action of generating electrode activation sequences is based on both global and local features.
36. The method of claims 29, 30, 31, 32, 33, 34 or 35, wherein: the action of generating electrode activation sequences is based on learned feature representation.
37. The method of claims 29, 30, 31, 32, 33, 34, 35 or 36, wherein: the received data is based on ambient sound with speech and distracting noise to a user of a hearing prosthesis that generates the electrode activation sequences; and the method includes automatically removing noise content from the received data when generating the electrode activation sequences.
38. The method of claims 29, 30, 31, 32, 33, 34, 35, 36 or 37, wherein: the received data is based on ambient sound with speech and distracting noise to a user of a hearing prosthesis that generates the electrode activation sequences; and the method includes automatically removing noise content from the received data when generating the electrode activation sequences.
39. The method of claims 29, 30, 31, 32, 33, 34, 35, 36, 37 or 38, wherein: there is no transitory and no non-transitory signal between the action of receiving data and the action of generating the electrode activation sequencies that has more noise content than elsewhere between the action of receiving data and the action of generating the electrode activation sequencies.
40. The method of claims 29, 30, 31, 32, 33, 34, 35, 36, 37, 38 or 39, wherein: the received data includes audio content including speech subcontent from at least two different microphones of a hearing prosthesis, which microphones capture the ambient sound; the action of generating electrode activation sequences includes enhancing the speech subcontent.
41. A method, compri sing : obtaining data based on audio and / or visual data; processing the obtained data using one or more algorithms based on machine learning; and evoking a sensory percept based on the processed data, wherein the action of processing the data using one or more algorithms one or more of:executes output side processing of a cochlear implant processing pipeline; or executes one or more of: amplitude compression; channel envelope sampling; maxima selection; speech enhancement; binaural processing; or microphone equalization.
42. The method of claim 41, wherein: the processing generates electrode activation sequences directly from the data based on audio and / or visual data.
43. The method of claim 41, wherein: the processing generates electrode activation sequences directly from the based on audio and / or visual data using a convolutional encoder-decoder.
44. The method of claim 41, wherein: the action of processing the data using one or more algorithms executes output side processing of a cochlear implant processing pipeline.
45. The method of claim 41, wherein: the action of processing the data using one or more algorithms executes one or more of: amplitude compression; channel envelope sampling; selects maxima selection; speech enhancement; binaural processing; or microphone equalization.
46. The method of claim 41, wherein: the obtained data is based on audio data; the processing is processing of a sound coding path for a cochlear implant; andthe path is entirely made up of a convolutional encoder-decoder network.
47. The method of claim 41, wherein: the action of processing the data using one or more algorithms executes two or more of: amplitude compression; channel envelope sampling; selects maxima selection; speech enhancement; binaural processing; or microphone equalization.
48. A method, comprising: obtaining output of signal processing of an input signal; and providing data based on the obtained output to a convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided data, wherein the output is electrical stimulation control data for a sensory prosthesis.
49. The method of claim 48, wherein: the signal processing is conventional audio and / or visual signal processing of the input signal.
50. The method of claim 48, wherein: the processing is audio processing; and the output is electrical stimulation control data for a cochlear implant.
51. The method of claim 50, wherein: the convolutional encoder-decoder trains itself on all aspects of a cochlear implant stimulation strategy during the training.
52. The method of claim 49, wherein: the processing is audio processing; the output is electrical stimulation control data for a cochlear implant; and the convolutional encoder-decoder is trained, simultaneously, in two or more of:microphone equalization; automatic gain control; envelope extraction; loudness growth functions; noise reduction; speech enhancement; maxima selection; envelope sampling; acoustic-to-electric mapping; and pulse sequencing.
53. The method of claim 50, wherein: the training includes penalizing for simulated energy quantities used by the convolutional encoder-decoder.
54. The method of claim 50, wherein: the training includes applying a loss function that minimizes differences between electrode activation patterns of the output and an output of the convolutional encoderdecoder.
55. The method of claim 50, further comprising: providing the output to a neural activation model, wherein the output is a first output; obtaining first data based on first neural activation model output based on the first output provided to the neural activation model; obtaining second output from the convolutional encoder-decoder; providing the second output to the neural activation model; and obtaining second data based on second neural activation module output based on the second output.
56. The method of claim 78, wherein: the training includes minimizing difference between neural activation patterns in the first data and the second data.
57. The method of claim 56, wherein:the training further includes penalizing for simulated number of electrical pulses used by the convolutional encoder-decoder.
58. The method of claim 50, wherein: the signal processing is a complex signal processing not usable in real time on a Nucleus 8 as of October 13, 2023.
59. The method of claim 48, wherein: the signal processing is signal processing executed by a product of machine learning.
60. The method of claim 48, wherein: the signal processing is signal processing executed by convolutional encoder-decoder.
61. The method of claim 50, wherein: the input signal is based on audio captured by an external microphone exposed to an air environment; the method further comprises providing a second input signal to the convolutional encoder-decoder, wherein the second input signal is based on audio captured by an implantable microphone in at least a simulated implanted environment, and obtaining a second output from the convolutional encoder-decoder based on the provided second input; the training includes comparing the obtained output to the second output, wherein the provided data is based on the comparison.
62. The method of claim 48, wherein: the input signal is based on audio captured by an external microphone exposed to an air environment; the method further comprises providing a second input signal to the convolutional encoder-decoder, wherein the second input signal is based on audio captured by an implantable microphone in at least a simulated implanted environment, and obtaining a second output from the convolutional encoder-decoder based on the provided second input; the training includes comparing the obtained output to the second output, wherein the provided data is based on the comparison.
63. The method of claim 49, wherein:the input signal is an audio input without noise; the obtained output is first output; the method includes obtaining second output of signal processing of a second input signal, wherein the second input signal is an audio input with noise, wherein the second output is electrical stimulation control data for the sensory prosthesis; and providing second data based on the obtained second output to the convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided second data.
64. A device comprising: a microphone; and a stimulation output assembly, wherein the device is a hearing prosthesis, the device includes a sound processing path that receives input from the implanted microphone, the device controls the tissue stimulation output assembly based on output from the sound processing path to evoke a hearing percept based on the input from the implanted microphone, and at least one of the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; the sound processing path includes a product of machine learning trained to minimize feedback; or a processing path that includes at least a portion of the sound processing path and extends to the tissue stimulation output assembly is made entirely of a product of machine learning.
65. The device of claim 61, wherein: the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output.
66. The device of claim 61, wherein:the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; and the sound processing path includes a sound processing path of a conventional implantable hearing prosthesis downstream of the product of machine learning.
67. The device of claim 61, wherein: the device is a totally implantable cochlear implant; the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; and the device includes a processing path that includes at least a portion of the sound processing path and that extends to the tissue stimulation assembly includes a conventional cochlear implant sound processor.
68. The device of claim 61, wherein: the sound processing path includes a product of machine learning trained to minimize feedback.
69. The device of claim 61, wherein: the device includes a processing path that includes at least a portion of the sound processing path and that extends to the tissue stimulation output assembly is made entirely of a product of machine learning.
70. The device of claim 69, wherein: the tissue stimulation assembly is configured to receive output from the processing path and controls electrical output of the tissue stimulation assembly based directly on the received output.
71. The device of claim 69, wherein: the processing path outputs electrode stimulation sequences to the tissue stimulation assembly.
72. A device, comprising:a signal input configured to receive data based on audio data and / or visual data; and circuitry configured to simultaneously functionally execute at least some input side processing and at least some output side processing of the data based on audio data and / or visual data.
73. The device of claim 72, wherein: the circuitry is configured to simultaneously functionally execute some input side processing and at all output side processing of the data based on audio data and / or visual data.
74. The device of claims 72 or 73, wherein: output side processing executed by the circuitry includes determination of timing and pattern of stimulation on one or more electrodes in signal communication with the device.
75. The device of claims 72, 73 or 74, wherein: output side processing executed by the circuitry includes maxima selection.
76. The device of claims 72, 73, 74 or 75, wherein: input side processing executed by the circuitry includes channel envelope development and / or frequency bin division.
77. The device of claims 72, 73, 74, 75 or 76, wherein: the circuitry mimics hard-coded filter banks in the input side processing.
78. A cochlear implant, comprising: a microphone; an electrode array; and the device of claims 72, 73, 74, 75, 76 or 77, wherein the cochlear implant controls stimulation of tissue with the electrode array based on results of the front end processing and back end processing.
79. The cochlear implant of claim 78, wherein: the output of the device is electrode activation sequence data.
80. The device of claims 72, 73, 74, 75, 76, 77 or 79, wherein: there is no dedicated circuitry dedicated to input side processing outside of front-end processing; and there is no dedicated circuitry dedicated to output side processing.
81. A cochlear implant, comprising: a microphone; an electrode array; and the device of claims 72, 73, 74, 75, 76, 77 or 79, wherein there is no circuitry solely dedicated to electrode activation sequencing.
82. A cochlear implant, comprising: one or more microphones; a product of machine learning in signal communication with the one or more microphones; and an electrode driver in signal communication with the product of machine learning; and an electrode array in electrical communication with the electrode driver, wherein the product is configured to receive an input from the microphones, and the product is configured to process the input and automatically develop electrode activation control data and provide such to the electrode driver.
83. A device and / or system and / or method, wherein: the device and / or system includes an input subsystem configured to receive input; the device and / or system includes a convolutional based sub-system in signal communication with the input subsystem; the convolutional based sub-system outputs electrode activation sequence based data; the system is an ambient environment stimulus processor system; the system is a sound processor system; the input subsystem is configured to receive data based on sound captured by a sound capture device; the convolutional based subsystem outputs: electrode activation sequences; ordata sets convertible to electrode activation sequences; bottleneck layers of the convolutional based subsystem develop current levels for electrode channels based on data received at the input subsystem; the device and / or system is a cochlear implant comprising a sound capture device and an electrode array; the cochlear implant provides data based on sound captured by the sound capture device to the input end, wherein the input end provides data based on the provided data to the convolutional based sub-system; the convolutional based sub-system provides output based on the data provided to the convolutional based subsystem, wherein the cochlear implant is configured to directly control stimulation by the electrode array to stimulate tissue to evoke a hearing percept based on the data output by the convolutional based subsystem; cochlear implant includes one or a plurality of sound capture devices; the cochlear implant system provides data based on sound captured by the plurality of sound capture devices to the input subsystem, wherein the input subsystem provides data based on the provided data to the convolutional based subsystem; the cochlear implant system controls the electrode array to evoke a hearing percept in response to the sound captured by the plurality of sound capture devices while maintaining a coupling between magnitude and phase thereof; the system includes a sound coding path for a cochlear implant; the path is entirely made up of a convolutional encoder-decoder network; the convolutional based sub-system is a trained sub-system trained at least in part with a loss function that minimizes difference between data based on a signal from an external microphone and data based on a signal from an implanted microphone; the cochlear implant provides data based on sound captured by the sound capture device to the input subsystem, wherein the input subsystem provides data based on the provided data to the convolutional based subsystem; the cochlear implant is configured to stimulate tissue to evoke a hearing percept based on the electrode activation sequence based data outputted by the convolutional based subassembly; there is no dedicated circuitry solely dedicated to electrode activation sequencing; the system is a sound processor and wherein the convolutional based subsystem is trained to reduce nose in the received input, the training being based on a loss function of minimizing a difference between electrode activation patterns that respectively results frominput having a desired base content and input having the desired base content plus noise content; the input subsystem is configured to receive data based on sound captured by a plurality of separate sound capture devices; the convolutional based subsystem executes noise reduction based on comparing respective data based on sound captured by the respective sound capture devices; the input subsystem is configured to receive data based on sound captured by a plurality of separate sound capture devices; the convolutional based subsystem executes speech enhancement and spatial filtering based on comparing respective data based on sound captured by the respective sound capture devices; the received data based on sound captured by a plurality of separate sound capture devices is contained in a single signal or a plurality of signals; the received data includes normalized cross-correlation data for spatially distinct sound capture devices; the convolutional based subsystem normalizes cross-correlation of spatially distinct sound capture devices; the device and / or system includes a product of machine learning; the product is configured to receive an input based on a captured phenomenon in the visual and / or audio environment; the product is configured to process the input and automatically develop electrode activation control data;; the device is a sensory prosthesis that includes electrodes; the device is configured to activate the electrodes based on the developed electrode activation control data to evoke a sensory percept; the device is a cochlear implant; there is no processing in the signal processing pipe of the device after the product; the device is a cochlear implant with at least 10 electrode channels; the electrode activation control data is dimensionalized relative to the number of channels of the at least 10 electrode channels; the device is a hearing prosthesis; the electrode control data is devoid of audio content; the device is a hearing prosthesis; the developed electrode control data has reduced noise content;the device is a hearing prosthesis; the developed electrode control data has phase content; at least some input side processing and all output side processing is executed by the product of machine learning; the product receives raw audio and outputs electrode stimulation sequences; the product receives data based on raw audio and outputs electrode stimulation sequences; the product is configured to process the input and automatically reduce noise content in the input in the development of the electrode activation control data; the product automatically reduces noise content at the same time that it automatically develops the electrode activation control data; the product is configured to receive an input based on a captured phenomenon in the audio environment; captured phenomenon includes audio content including speech subcontent from at least two different microphones of a hearing prosthesis, which microphones capture the captured phenomenon; the product includes a separator configured to calculate a mask to reduce noise from the audio content in a latent space of the product; the product is configured to receive an input based on a captured phenomenon in the audio environment; captured phenomenon includes audio content including speech subcontent from at least two different microphones of a hearing prosthesis, which microphones capture the captured phenomenon; the product includes a separator configured to enhance the speech subcontent; the method includes receiving data based on at least one of ambient sound and / or ambient light in an ambient environment of a device or electromagnetic signals received by the device, the data enabling direct usage thereof by a conventional speaker apparatus and / or a conventional video monitor apparatus to produce a recognizable corresponding sound and / or image without further processing; the method includes generating electrode activation sequences directly from the received data; the method includes generating electrode activation sequences indirectly from the received data;the action of generating electrode activation sequencies is executed in real time vis-a- vis the action of receiving data; the method further comprises evoking a hearing percept based on the generated electrode activation sequences; the actions of receiving data and generating electrode activation sequences are executed by a body -worn cochlear implant without converting data based on the received data to a Fourier domain; the action of generating electrode activation sequences is executed by a product of a trained neural network; all backend processing is executed by data-driven developed algorithms; the actions of receiving data and generating electrode activation sequences are executed by a body -worn cochlear implant; the action of generating electrode activation sequences is based on both global and local features; the action of generating electrode activation sequences is based on learned feature representation; the received data is based on ambient sound with speech and distracting noise to a user of a hearing prosthesis that generates the electrode activation sequences; the method includes automatically removing noise content from the received data when generating the electrode activation sequences; the received data is based on ambient sound with speech and distracting noise to a user of a hearing prosthesis that generates the electrode activation sequences; the method includes automatically removing noise content from the received data when generating the electrode activation sequences; there is no transitory and no non-transitory signal between the action of receiving data and the action of generating the electrode activation sequencies that has more noise content than elsewhere between the action of receiving data and the action of generating the electrode activation sequencies; the received data includes audio content including speech subcontent from at least two different microphones of a hearing prosthesis, which microphones capture the ambient sound; the action of generating electrode activation sequences includes enhancing the speech sub content; the method includes obtaining data based on audio and / or visual data;the method includes processing the obtained data using one or more algorithms based on machine learning; the method includes evoking a sensory percept based on the processed data; the action of processing the data using one or more algorithms one or more of: executes output side processing of a cochlear implant processing pipeline; or executes one or more of: amplitude compression; channel envelope sampling; maxima selection; speech enhancement; binaural processing; or microphone equalization; the processing generates electrode activation sequences directly from the data based on audio and / or visual data; the processing generates electrode activation sequences directly from the based on audio and / or visual data using a convolutional encoder-decoder; the action of processing the data using one or more algorithms executes output side processing of a cochlear implant processing pipeline; the action of processing the data using one or more algorithms executes one or more of: amplitude compression; channel envelope sampling; selects maxima selection; speech enhancement; binaural processing; or microphone equalization; the obtained data is based on audio data; the processing is processing of a sound coding path for a cochlear implant; the path is entirely made up of a convolutional encoder-decoder network; the action of processing the data using one or more algorithms executes two or more of: amplitude compression; channel envelope sampling; selects maxima selection;speech enhancement; binaural processing; or microphone equalization; the method includes obtaining output of signal processing of an input signal; providing data based on the obtained output to a convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided data; the signal processing is conventional audio and / or visual signal processing of the input signal; the processing is audio processing; the output is electrical stimulation control data for a cochlear implant; the convolutional encoder-decoder trains itself on all aspects of a cochlear implant stimulation strategy during the training; the processing is audio processing; the output is electrical stimulation control data for a cochlear implant; the convolutional encoder-decoder is trained, simultaneously, in two or more of: microphone equalization; automatic gain control; envelope extraction; loudness growth functions; noise reduction; speech enhancement; maxima selection; envelope sampling; acoustic-to-electric mapping; pulse sequencing; the training includes penalizing for simulated energy quantities used by the convolutional encoder-decoder; the training includes applying a loss function that minimizes differences between electrode activation patterns of the output and an output of the convolutional encoderdecoder; the method includes providing the output to a neural activation model, wherein the output is a first output; the method includes obtaining first data based on first neural activation model output based on the first output provided to the neural activation model;the method includes obtaining second output from the convolutional encoder-decoder; the method includes providing the second output to the neural activation model; the method includes obtaining second data based on second neural activation module output based on the second output; the training includes minimizing difference between neural activation patterns in the first data and the second data; the training further includes penalizing for simulated number of electrical pulses used by the convolutional encoder-decoder; the signal processing is a complex signal processing not usable in real time on a Nucleus 8 as of October 13, 2023; the signal processing is signal processing executed by a product of machine learning; the signal processing is signal processing executed by convolutional encoder-decoder; the input signal is based on audio captured by an external microphone exposed to an air environment; the method further comprises providing a second input signal to the convolutional encoder-decoder, wherein the second input signal is based on audio captured by an implantable microphone in at least a simulated implanted environment, and obtaining a second output from the convolutional encoder-decoder based on the provided second input; the training includes comparing the obtained output to the second output, wherein the provided data is based on the comparison; the input signal is based on audio captured by an external microphone exposed to an air environment; the method further comprises providing a second input signal to the convolutional encoder-decoder, wherein the second input signal is based on audio captured by an implantable microphone in at least a simulated implanted environment, and obtaining a second output from the convolutional encoder-decoder based on the provided second input; the training includes comparing the obtained output to the second output, wherein the provided data is based on the comparison; the input signal is an audio input without noise; the obtained output is first output; the method includes obtaining second output of signal processing of a second input signal, wherein the second input signal is an audio input with noise, wherein the second output is electrical stimulation control data for the sensory prosthesis;providing second data based on the obtained second output to the convolutional encoder-decoder and training the convolutional encoder-decoder based on the provided second data; the device and / or system includes a microphone; a stimulation output assembly; the device is a hearing prosthesis; the device includes a sound processing path that receives input from the implanted microphone; the device controls the tissue stimulation output assembly based on output from the sound processing path to evoke a hearing percept based on the input from the implanted microphone; at least one of: the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; the sound processing path includes a product of machine learning trained to minimize feedback; a processing path that includes at least a portion of the sound processing path and extends to the tissue stimulation output assembly is made entirely of a product of machine learning; the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output; the sound processing path includes a sound processing path of a conventional implantable hearing prosthesis downstream of the product of machine learning; the device is a totally implantable cochlear implant; the sound processing path includes a product of machine learning trained to minimize difference between data based on implanted microphone output and data based on external microphone output;the device includes a processing path that includes at least a portion of the sound processing path and that extends to the tissue stimulation assembly includes a conventional cochlear implant sound processor; the sound processing path includes a product of machine learning trained to minimize feedback; the device includes a processing path that includes at least a portion of the sound processing path and that extends to the tissue stimulation output assembly is made entirely of a product of machine learning; the tissue stimulation assembly is configured to receive output from the processing path and controls electrical output of the tissue stimulation assembly based directly on the received output; the processing path outputs electrode stimulation sequences to the tissue stimulation assembly; the device includes a signal input configured to receive data based on audio data and / or visual data; circuitry configured to simultaneously functionally execute at least some input side processing and at least some output side processing of the data based on audio data and / or visual data; the circuitry is configured to simultaneously functionally execute some input side processing and at all output side processing of the data based on audio data and / or visual data; output side processing executed by the circuitry includes determination of timing and pattern of stimulation on one or more electrodes in signal communication with the device; output side processing executed by the circuitry includes maxima selection; input side processing executed by the circuitry includes channel envelope development and / or frequency bin division; the circuitry mimics hard-coded filter banks in the input side processing; the cochlear implant controls stimulation of tissue with the electrode array based on results of the front end processing and back end processing; the output of the device is electrode activation sequence data; there is no dedicated circuitry dedicated to input side processing outside of front-end processing; there is no dedicated circuitry dedicated to output side processing;an encoder-decoder (ED) structure that has a decoder that instead of returning the latent space representation back to the input domain, transforms the latent space representation into a domain different from the input domain; in the case of audio signal processing, input audio is transformed into the latent space by the convolutional encoder, bottleneck layers process the latent space, and then the convolutional decoder transforms the latent space not back to audio, but instead into electrical pulse data / electrode activation sequence data (this can be done at the decoder); the device configured to convert audio input / audio data to electrode activation sequences (or data directly related thereto) at the decoder; the device is a convolutional encoder-decoder neural networks specifically designed for sensory prosthesis applications, such as CI and retinal implant / bionic eye (BE) applications; the audio domain is transformed to the latent space using a convolutional encoder, bottleneck layers process the latent space, and then the latent space is transformed to the electrode space (e.g., CI or BE space); the dimension of the output corresponds to the number of active implanted electrodes(N), which can be by way of example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19,20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44,45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 65, 70, 75, 80, 90, 100, 125, 150,175, 200, 250, 300, 350, 400 or 450 or more or any value or range of values therebetween in 1 increment; the device and / or system is or includes and / or the method uses a trained or in training DNN, a trained or in training CED, a trained or in training product of machine learning; an output of the device and / or system is electrode activation sequences; the method includes electrode activation sequences are generated directly from an audio input using a convolutional encoder-decoder; one or more or all of the amplitude compression, channel envelope sampling, maxima selection, acoustic-to-electric mapping, and further signal processing such as noise reduction, speech enhancement, binaural processing, and microphone equalization are encapsulated within one deep learning framework (such as a product of machine learning / a product of artificial intelligence); the electrode activation sequence based data has data that corresponds to or is directly translatable to length of time of current application to be applied to an electrode;the electrode activation sequence based data has data that corresponds to or is directly translatable to a relative activation time relative to one or more other electrodes; the electrode activation sequence based data (which can be electrode activation sequences) can include active electrode designation (electrode number), return electrode designation (electrode number), pulse amplitude, phase duration of activation (uniform our could be per channel), interphase gap of activation (uniform or could be per channel), pulse shape (uniform our could be per channel) and / or interpulse interval (uniform our could be per channel); the sequences that are output by the CED can be combined with the fixed phase duration and interphase gap to comprise a complete electrode activation sequence (or not, the driver could be “hard controlled” to have the fixed data); the CED outputs the entire electrode activation sequence; the output of the CED is output that has been adjusted or otherwise is output that takes into account threshold and / or comfort levels of the recipient; there is a downstream component downstream from the CED that takes the output and adjusts for threshold and / or comfort levels of the recipient; the downstream component is the downstream component that is controlled based on fitting data sets that are loaded into the hearing prostheses as a result of a so-called fitting procedure where the prostheses is fitted to the recipient; the output of the CED is normalized on a scale of 0 to 1, or more accurately, the amplitude portion of the electrode activation sequence data is normalized on a scale of 0 to 1, the component downstream of the CED then takes the normalized amplitude for one or more or all channels and adjusts such to accommodate the threshold levels and the comfort levels for those channels for that particular recipient; any filtering is not based on FFT; irrespective of whether or not the data output by the CED is adjusted to take into account fitting, that output is output that is readily usable by the electrode driver or by a controller that controls the electrode driver apparatus; the output from the CED can be utilized to drive the electrode driver; bottleneck layers of the convolutional based subsystem develop current levels and / or current timing for electrode channels based on data received by the input subsystem, or otherwise develop components of the electrode activation sequence, including active electrode, phase duration, interphase gap, return electrode and / or interpulse interval;the device includes a convolutional decoder which can include a ID or 2D transposed convolutional layer; the output of the CED is output that is in a dimension different from that of the input dimension; with respect to a processing pipeline, the output of the CED is never returned to the input dimension(s); the device and / or system includes an encoder-decoder (ED) structure that has a decoder that instead of returning the latent space representation back to the input domain, transforms the latent space representation into a domain different from the input domain the device and / or system includes a network that efficiently capture local and global features, enabling hierarchical feature extraction across various time scales; the device and / or system down-samples input-signals (e.g., with the CED) to create a learned latent representation, expanding temporal context and reducing complexity in bottleneck layers; the device and / or system is configured to have autonomous feature discovery; the device is configured to enhance speech subcotent by at least and / or equal to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 dB or more or any value or range of values therebetween in 0.1 dB increments; the device and / or system executes any one or more of the method actions above; or the method includes executing any one or more of the functionalities and / or uses any one or more of the components above.
Citation Information
Patent Citations
New sound processing techniques
US20210260377A1
Systems and methods for training a machine learning model for use by a processing unit in a cochlear implant system
US20220249844A1
Signal processing in a hearing device
US20230127309A1
Automatic measurement of an evoked neural response concurrent with an indication of a psychophysics reaction
US8190268B2
System and method for real-time cochlear implant localization
WO2020215000A1
Cited By
Fully-implanted artificial cochlea system and noise reduction method thereof
CN120789479A