Auditory cortical implant and system and method for generating stimulation patterns for the auditory cortical implant

A two-dimensional flexible cortical implant with a trained encoder neural network addresses the limitations of cochlear implants by providing high resolution sound information to the auditory cortex, improving hearing perception and discrimination.

WO2026052735A1PCT designated stage Publication Date: 2026-03-12INST PASTEUR +4
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Current cochlear implants suffer from limitations such as inability to stimulate patients without auditory nerves, poor spectral accuracy, and limited perception in noisy environments, due to their low spatial resolution and one-dimensional encoding schemes, while cortical implants face issues with irreversible brain damage and low resolution.

Method used

A two-dimensional flexible surface stimulator with high electrode density and a trained encoder neural network to generate stimulation patterns, using an autoencoder architecture that respects spatial organization and continuity of sound representations, ensuring high resolution sound information transmission to the auditory cortex.

Benefits of technology

The solution provides improved hearing experiences by achieving more than one thousand independent information channels with better spatial resolution, enabling perception of tonal languages and music, and enhancing sound discrimination in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025075227_12032026_PF_FP_ABST
    Figure EP2025075227_12032026_PF_FP_ABST
Patent Text Reader

Abstract

System (100), method and computer program product for training an encoder neural network (110) to generate stimulation patterns (192) for an auditory cortical implant having a predefined number of electrodes arranged in a two‐dimensional electrode array (220) adapted for use in mammals (201) with impaired hearing. The encoder processes soundwave training data records (210i) and outputs for each corresponding soundwave a one‐ dimensional encoded sound vector (191) which is mapped to the electrode array (220) to obtain corresponding two‐dimensional stimulation patterns (192) for the auditory cortical implant. A decoder neural network (130) reconstructs the respective training data record from the encoder output wherein the parameters of the decoder are determined in parallel to the encoder during training. A cost function with a first term (CFr) to measure reconstruction performance and a second term (CFc) to measure continuity of the stimulation patterns is used when training the encoder (110) and decoder (130 while aiming at minimizing a weighted sum of the terms (CFr, CFc) of the cost function.
Need to check novelty before this filing date? Find Prior Art

Description

Auditory Cortical Implant and System and Method for Generating Stimulation Patterns for the Auditory Cortical ImplantTechnical Field

[0001] The present invention generally relates to the field of hearing restoration for mammals with impaired hearing, and in particular relates to a high resolution auditory cortical implant and training an encoder neural network to generate stimulation patterns for the auditory cortical implant.Background

[0002] With almost half a billion impacted people worldwide, hearing loss is a leading cause of disease burden. Disabling hearing loss, defined as an elevation of hearing thresholds of at least 30 dB, has different degrees, which imply different strategies to mitigate the handicap. Particularly challenging is functional deafness, in which the patient is not able to perceive by his / her own means any useful sound, and can only compensate, when possible, with other senses. Fortunately, in many cases hearing can be partially restored by the use of cochlear implants, an electrical device that bypasses the sensory organ to directly stimulate the auditory nerve and provide auditory inputs. Currently, more than a million patients worldwide have received cochlear implants, and most of them have recovered auditory abilities sufficient for some speech perception and production.

[0003] Despite their large success, cochlear implants unfortunately suffer from three major limitations. Firstly, they cannot be applied to a non-negligible minority of patients who have lost their auditory nerve. Some of these patients have been tentatively treated with auditory brainstem implants (ABI) in the recipient structure of the auditory nerve, the cochlear nucleus. In fact, ABIs cannot restore typical hearing, but they can improve sound awareness and identification as well as lip-reading ability. Moreover, results vary widely and serious, even if rare, side effects exist, including the generation of non-auditory perceptions or involuntary movements through stimulation. Secondly, cochlear implants are well-known for their poor spectral accuracy, leaving implanted patients barely able to judge the elevation and fine frequency structure of sounds, information that contributes to the meaning of words in tonal languages (e.g. Mandarin), and that conveys important contextual information about the gender, emotional state and intentions of the speaker. For the samereason, cochlear implants do not allow for the perception of music or of many pleasurable elements in sounds. Third, the poor acoustic resolution of cochlear implants severely limits perception in noisy environments, with a strong impact on social life for implanted patients.

[0004] The reason for the low sound fidelity of cochlear implants is the poor spatial resolution of stimulation provided by the 12-22 electrical contacts over ~30mm implants placed along the so-called tonotopic axis of the cochlea, along which sound frequency is first defined. It is considered that a 12-electrode implant with a typical electrode distance of 2.4mm provides maximal precision. Due to spatial spread of electrical currents, higher electrode resolution does not improve the restored perception because neighboring electrode information is not discriminated by the brain.

[0005] It has been envisioned that the central auditory brain, the auditory cortex, could offer an interesting alternative target. The auditory cortex has a surface of >1600mm2, which, compared to the ~40mm length of the cochlea, offers at least 40 times more space for electrical contact. Based on this, the idea of cortical implants has been proposed in the past. However, current methods to stimulate the cortex, used for pilot clinical trials for visual cortex stimulation, either involve penetrating electrodes which leave irreversible damage in the brain or low-resolution surface electrode arrays, which provide only two- to three-fold more electrical contacts than cochlear implants. In addition, the only encoding scheme for transmitting auditory information to the auditory cortex is a 1-dimensional encoding scheme used for cochlear implants.Summary

[0006] No solution exists so far to address the full potential of auditory cortex stimulation for auditory rehabilitation with high density stimulators. There is therefore a need for improving brain stimulation of hearing-impaired mammals (also referred to as patients) to enable an improved hearing experience. This technical problem is solved by the features of the independent claims.

[0007] In one embodiment, a computer-implemented method is provided for training an encoder neural network (also referred to as encoder herein) to generate stimulation patterns for an auditory cortical implant having a predefined number of electrodes arranged in a two-dimensional electrode array adapted for use in mammals with impaired hearing.The central auditory brain of a human, the primary auditory cortex, has a surface of ~1600mm2, which, compared to the ~40mm length of the cochlea, offers 40 times more space for electrical contact. The electrode array allows for high density stimulation of the auditory cortex, and therefore provides a way to transmit high resolution sound information to the patient's brain. However, current methods to stimulate the cortex typically involve penetrating electrodes which may leave irreversible damage in the brain, or low-resolution surface electrode arrays, which provide only two- to three-fold more electrical contacts than cochlear implants, leaving unknown the maximal resolution that can be achieved with cortical implants. To address this issue, a two-dimensional, flexible surface stimulator is suggested which can be positioned on the brain surface avoiding any brain tissue damage. In an example implementation, the stimulator has a contact density of 25 electrodes per square millimeter and an electrode-to-electrode distance of 180pm, allowing the placement of several thousand electrodes over the human auditory cortex. A person skilled in the art will appreciate that the surface stimulator may have a contact density in the range of 1 to 40 electrodes per square millimeter and an electrode-to-electrode distance in the range of 160pm to 1000pm

[0008] Preclinical trial in mice demonstrated that distinct perceptions were obtained when 100 pV voltage pulses (absolute pulse heights) were applied to neighboring electrodes. Therefore, the spatial resolution of cortical stimulation is also more than one order of magnitude better than the cochlear implant. This ensures that cortical implants can generate more than one thousand independent information channels. To exploit this density and the two dimensions of the cortical implant (instead of one for the cochlear implant), artificial neural networks were used to map sound information to the two-dimensional stimulator while preserving the known spatial organization of sound representations. The algorithm to optimize the cortical implant technology is based on a so-called autoencoder architecture which will be described in detail below.

[0009] The computer-implemented method starts with receiving a dataset of soundwaves as training data records for the encoder.

[0010] In an example experiment for mice, the training data included soundwaves including the following sound types:- pure tones with frequencies in the range of 500Hz to 28kHz,- amplitude-modulated tones with a carrier frequency between 6kHz and 16kHz and a modulating frequency between 10Hz and 200Hz,- frequency steps with starting frequencies between 500Hz and 20kHz, wherein the number of steps is between 2 and 60 - frequency step is a tone with one single frequency played for a certain duration,- chirps (sweeps) with starting frequencies between 500Hz and 20kHz and ending frequencies between 500Hz and 20kHz (with the starting frequency being different from the ending frequency) - chirps / sweeps combine many steps of increasing or decreasing frequencies, and- chords with fundamental frequencies between 1kHz and 10kHz.For the human auditory cortex, other frequency ranges can be used, such as for example:- pure tones with frequencies in the range of 20Hz to 20kHz,- amplitude-modulated tones with a carrying frequency between 20Hz and 20kHz and a modulating frequency between 10Hz and 200Hz,- chirps with starting frequencies between 20Hz and 20 kHz and ending frequencies between 20Hz and 20kHz,- frequency steps with starting frequencies between 20Hz and 20kHz, wherein the number of steps is between 2 and 60 - Combining many steps of increasing or decreasing frequencies and- chords with fundamental frequencies between 20Hz and 20kHz.

[0011] A person skilled in the art may also use other appropriate sound types as training data. For example, in particular with regard to the human auditory cortex, the autoencoder model can be trained with a large dataset of natural sound records (e.g., Audioset available at https: / / research.google.com / audioset / , extracted from a large video databank) in which each natural sound record is labelled by human listeners with the object or action they recognize in the sound (e.g. trumpet, cheering, crowd, motor vehicle, duck, etc..).

[0012] In an optional embodiment, each received training data record of the dataset may be transformed into a respective transformed training data record in the frequency domain. In this embodiment the transformed training data records are used as the training data records for all subsequent steps. In one implementation, transforming into the frequency domain may be achieved by using a Fourier Transform algorithm, or a constant Q-transform, toproduce a compact time-frequency representation of the soundwaves which respects the logarithmic scaling of frequency resolution of the mammalian hearing system. In another implementation, the encoder neural network which receives the training data as input, may comprise additional layers to preprocess the received soundwaves through a filter bank whose parameters are learnt during the training of the encoder.

[0013] The encoder processes each training data record to output for each corresponding soundwave a one-dimensional encoded sound vector of scalar values. That is, each encoded sound vector is a one-dimensional compressed representation of the respective original soundwave. The plurality of encoded sound vectors is sometimes also referred to as the latent space. For example, the encoder neural network may be implemented by a plurality of convolutional layers, each convolutional layer being a custom linear filter bank followed by an activation layer.

[0014] The encoded sound vector is then mapped to the electrode array of the auditory cortical implant to obtain a two-dimensional stimulation pattern for controlling the electrical stimulation of the mammal's auditory cortex via the auditory cortical implant. In other words, the one-dimensional encoded sound vector is transformed into a two-dimensional stimulation pattern whose dimensions match the dimensions of the electrode array.

[0015] The encoder forms part of an autoencoder neural network which tries to reconstruct the original encoder inputs and uses a cost function to measure the mathematical distance between the reconstructed data and the original inputs to train the encoder. In more detail, for each encoder output, a reconstruction of the respective training data record is generated by a decoder neural network (also referred to as decoder herein) by processing the respective encoded sound vector or the corresponding stimulation pattern, wherein the parameters of the decoder are determined in parallel to the encoder during training. The reconstruction performance for the training data is measured through a cost function with a first term that computes a first mathematical distance between any received training data record and its corresponding reconstructed training data record.

[0016] For the generated stimulation patterns to be applied to the auditory cortex, so-called continuity is an important characteristic to enable or improve the hearing capability of the patient. Generally spoken, continuity means that for similar consecutive soundwaves also the resulting stimulation patterns need to be similar. In more detail, stimulation patterns forcontinuously changing sound features (e.g., sound frequency) must be continuous. For example, if the frequency of a pure tone (i.e., a single frequency sound) changes little from one soundwave record to the next one, the corresponding stimulation patterns must also remain very similar. This is not the case when sounds are represented in highly resolved spectrograms, for which the mathematical similarity between two pure tones with a 10% frequency difference is identical to the mathematical similarity between two tones with 100% frequency difference (i.e., one octave interval between both tones). The autoencoder will therefore not necessarily encode similar sound pairs into similar patterns. To impose the continuity of representation between similar sounds that is perceptually experienced and observed in the brain, the neural network is forced to produce a target similarity matrix across a set of reference sounds (pure tone frequency and amplitude modulations) that matches an estimation of the similarity of representations observed in brain for these sounds. This is done by adding a term to the cost function which computes a mathematical distance (e.g., Euclidean distance) between a reference similarity matrix and an actual similarity matrix. This distance is minimized during the training process, thus making the encoding patterns sufficiently continuous. In an example implementation, the reference soundwaves are pure tones in a range of audible frequencies, preferably between 100Hz - 10000Hz, and amplitude modulations in a predefined range of frequencies, preferably from 1Hz to 200Hz. In other words, to enforce continuity between pure tones (PT) or between amplitude-modulated (AM) tones, at each training step, the similarity matrix of a collection of PT (as well as AM tones) is computed (continuity is between different pure tones, or between different AM tones). This similarity matrix is then compared to the reference similarity matrix (i.e., the ground truth), and a respective loss is computed.

[0017] In other words, to enforce continuity of the obtained stimulation patterns, the continuity of obtained stimulation patterns is measured by computing a pairwise similarity matrix between stimulation patterns associated with a predefined set of reference soundwaves that are perceptually continuous. Thereby, the predefined set of reference soundwaves is a subset of the dataset comprising the soundwave training data. A second term of the cost function is then used to enforce the continuity of the obtained stimulation patterns. This second term computes a second mathematical distance between the computed similarity matrix and a corresponding predefined reference similarity matrix that accounts for perceptual similarity.

[0018] Then, the encoder and decoder neural networks (i.e., the autoencoder) are trained (simultaneously) while aiming at minimizing a weighted sum of the terms of the cost function.

[0019] Whereas the continuity of the stimulation patterns is an essential requirement for a good hearing experience, the hearing experience can be further improved by also taking into account spatial considerations regarding the stimulation patterns in an optional embodiment implementing a spatial mapping. The centroids of the latent space activations of the encoder for increasing frequencies of single frequency sounds (i.e., pure tones) follows a defined two-dimensional pattern on the two-dimensional space of the electrode array. This reproduces the known organization of pure tone responses observed in regions of the auditory cortex that the implant targets and which are typically organized in the human brain along multiple frequency gradients. This principle with a single frequency gradient is demonstrated further down below, by adding a penalty term to the minimized cost function which corresponds to the distance between actual and target pure tone activity centroids. In fact, the solution follows a generic approach that can be extended to more complex frequency map patterns and to the spatial organization of other sound features. For example, in the human auditory cortex, a map for amplitude modulation frequency is described in: Weisz, N., & Lithari, C. (2017). Amplitude modulation rate dependent topographic organization of the auditory steady-state response in human auditory cortex. Hearing Research, 354, 102-108. The same approach can be used to force the autoencoder model to organize the stimulation patterns which correspond to different amplitude modulation frequencies in accordance with this map. In other words, in this optional embodiment, the cost function further has a third term corresponding to a third mathematical distance between known spatial patterns of sound-driven responses in auditory cortex regions targeted by the implant, to enforce that the obtained stimulation patterns respect multiple pure tone frequency gradients and amplitude frequency gradients observed in the auditory cortex of the human brain, wherein the weighted sum also takes into account the third term of the cost function. It is to be noted, that this embodiment particularly targets human patients.

[0020] In one embodiment, the method further limits the number of simultaneously active channels of each obtained stimulation pattern to a predefined number of most activechannels by keeping only the predefined number of most active channels of the encoder output with other channels being set to zero. The encoder output with the limited number of simultaneously active channels is also being used as input to the decoder for training data reconstruction. In this embodiment, the limiting of the number of simultaneously active channels to the M most active channels reduces energy consumption and channel crosstalk, and further reduces the risk for a clinical application that the stimulator overdrives the brain which might lead to medical disadvantages for the patient (e.g., epilepsy).

[0021] In one embodiment, the active channels of the stimulation pattern can be subject to a blurring step before being processed by the decoder. It is advantageous when the stimulation patterns are discriminable even after a spatial blurring corresponding to the effective resolution of the stimulation itself. The reason for that is that the size of the active area is typically larger than the stimulating electrode. To make sure that the auditory cortical implant works efficiently, this effect can be taken into account to avoid obtaining a code in the stimulation pattern that concentrates the pixels of very different sounds in nearby regions. This can be implemented by blurring (using a two-dimensional spatial filter whose input is the encoder's output pattern and whose output is the decoder's input) the pattern of the active channels in each obtained stimulation pattern with a two-dimensional kernel of a predefined width, to account for the physical resolution of the stimulation. The physical resolution is measured as the spatial spread of neuronal activity generated by stimulation with one electrode of the cortical implant. For example, the two-dimensional kernel can be a Gaussian kernel of width W.

[0022] In one embodiment, a computer program product is provided for training an encoder neural network to generate stimulation patterns for an auditory cortical implant. The program includes computer-readable instructions that, when loaded into a memory of a computing device and executed by one or more processors of the computing device, cause the computing device to perform the herein disclosed computer-implemented method.

[0023] In one embodiment, a computer system is provided for training an encoder neural network to generate stimulation patterns for an auditory cortical implant having a predefined number of electrodes arranged in a two-dimensional electrode array adapted for use in mammals with impaired hearing. The system has a memory storing computer- readable instructions implementing functional units, and one or more processors adaptedfor executing the functional units. The functional units are adapted to perform the herein disclosed computer-implemented method at system runtime (i.e., when the program is executed and the respective functional units are instantiated).

[0024] In one embodiment, an auditory cortical implant is provided which has a predefined number of electrodes arranged in a two-dimensional electrode array adapted for use in mammals with impaired hearing. The implant is adapted for high resolution electrical stimulation wherein the electrodes are arranged as a flexible two-dimensional surface stimulator adapted for positioning on the surface of a mammal's brain. The surface stimulator has a contact density in the range of 1 to 40 electrodes per square millimeter and an electrode-to electrode distance in the range of 160pm to 1000pm. The electrical stimulation of the electrodes is controlled by stimulation patterns received by the implant from an encoder. Each stimulation pattern encodes a particular soundwave into a two- dimensional array with dimensions corresponding to the two-dimensional electrode array such that stimulation patterns for continuously changing sound frequencies are continuous in that, if the frequency of a single frequency sound changes a little, the corresponding stimulation patterns remain very similar.

[0025] Optionally, the received stimulation patterns may further have a minimized mathematical distance between known spatial patterns of sound-driven responses in auditory cortex regions targeted by the implant, thereby respecting multiple pure tone frequency gradients and amplitude frequency gradients observed in the auditory cortex of the human brain.

[0026] Optionally, the number of simultaneously active channels of each received stimulation pattern may be limited to a predefined number of most active channels. Thereby, the pattern of the predefined number of most active channels in each obtained stimulation pattern may be blurred with a two-dimensional kernel of a predefined width, to account for the physical resolution of the stimulation.

[0027] Further aspects of the invention will be realized and attained by means of the elements and combinations particularly depicted in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only, and are not restrictive of the invention asdescribed.Brief Description of the Drawings

[0028] FIG. 1 illustrates a block diagram including for training an encoder neural network to generate stimulation patterns for an auditory cortical implant according to an embodiment;FIG. 2 is a simplified flow chart illustrating a computer-implemented method for such training according to an embodiment;FIG. 3 illustrates enforcement of continuity during the training using a respective loss function term;FIG. 4A illustrates an optional cost function term taking into account the tonotopy of the auditory cortex;FIG. 4B illustrates effectiveness of the tonotopic loss;FIG. 5 illustrates details of a blurring function between the encoder and decoder according to an embodiment;FIG. 6 shows an example of the representation continuity for sample pure tones and its match to the continuity constraint imposed;FIG. 7 shows examples of reconstructed soundwave spectrograms and corresponding originals for one embodiment;FIG. 8 shows the loss over iterations for example training data records and validation data records used for the training for the auditory cortex of mice;FIG. 9 shows a photo of an embodiment of an auditory cortical implant with electrodes arranged in a two-dimensional electrode array adapted for use in mammals with impaired hearing; andFIG. 10 is a diagram that shows an example of a generic computer device and a generic mobile computer device, which may be used in embodiments of the invention.Detailed Description

[0029] FIG. 1 illustrates a block diagram including a system 100 for training an encoder neural network 110 to generate stimulation patterns 192 for an auditory cortical implant.FIG. 2 is a simplified flow chart of a computer implemented method 1000 for such encodertraining according to an embodiment. FIG. 1 is described now in the context of FIG. 2. For this reason, reference numbers of FIG. 1 and FIG. 2 are used in the following description. The auditory cortical implant has a predefined number of electrodes arranged in a two- dimensional electrode array 220 adapted for use in mammals 201 with impaired hearing. In FIG. 1, mammal 201 is a human but it can also be another mammal (e.g., mouse, cat, dog, etc.). The auditory cortical implant is adapted for high resolution electrical stimulation with the electrodes being arranged as a flexible two-dimensional surface stimulator adapted for positioning on the surface of a mammal's brain. Advantageously, the surface stimulator has a contact density in the range of 1 to 40 electrodes per square millimeter and an electrode- to electrode distance in the range of 160pm to 1000pm.

[0030] In the example embodiment, the encoder 110 is part of an autoencoder module 140 (illustrated by a dash-dotted frame), which also includes a decoder 130. The encoder and decoder can be implemented as neural networks. In an example implementation, the autoencoder model 140 is built using the Pytorch 2.0.1 package and dependencies. The encoder 110 was composed of 4 convolutional layers, each followed by an activation layer. The encoder's output is a custom linear layer encompassing active unit filtering. Table 1 illustrates details of the encoder network architecture used by the example implementation.Table 1 - Encoder Network Architecture

[0031] The encoder of the example embodiment was designed for a 10x10 channels stimulator aiming to target the mouse auditory cortex (i.e. 100 stimulation channels) which was then used in mice.

[0032] Table 2 illustrates details of the decoder network architecture used by the example implementation.Table 2 - Decoder Network Architecture

[0033] System 100 has an interface which is adapted to receive 1100 a dataset 210 of soundwaves as training data records for the encoder 110. In the example implementation, the dataset of soundwaves used to train and validate the model for the auditory cortex of mice was generated using a custom Python library Quicksound. This library is available through PyPi. For reproducibility, a random seed for sound generation was fixed.Soundwaves were generated with a 64kHz sample rate at 70 dB SPL. If there is added noise, the noise is set at 60 dB SPL. The duration of a sound is set at 500ms. It is to be noted no tone was generated but all sounds had an absolute scale. This is important because the encoding is nonlinear, so encoding of a 40dB and of a 70dB sound can be very different. In general, dB is a relative scale. The SPL relates it to a reference (i.e., 0 dB SPL is the human threshold of hearing at 1kHz).

[0034] The dataset 210 used in the example implementation can be described as follows:Pure tones: 3000 occurrences from 500Hz to 28kHzAmplitude-modulated tones: 3000 occurrences with a carrying frequency between 6kHz and 16kHz and a modulating frequency between 10Hz and 200HzChirps (sweeps): 3000 occurrences with starting frequencies between 500Hz and 20kHz and ending frequencies between 500Hz and 20kHz.Frequency steps: 3000 occurrences with starting frequencies between 500Hz and 20kHz. Number of steps between 2 and 60.Chords: A total of 4000 occurrences.200 fundamental frequencies between 1kHz and 10kHz. For each fundamental frequency, 20 of the 32 possible arrangements of 5 harmonics are generated. If a harmonic goes beyond 28kHz, it is discarded.This comes to a total of 16,000 unique soundwaves. This sound dataset was augmented by using a combination of white noise and natural noise extracted from the WSJ0 Hipster Ambient Mixtures WHAM! dataset. The WHAM! dataset pairs each two-speaker mixture in the wsj0-2mix dataset with a unique noise background scene. 800 random sound instances were drawn from the WHAM! By associating randomly each sound with 6 noise samples, an augmented dataset of 112000 soundwaves was obtained. Each noise-corrupted sound was associated with its noise-free counterpart for denoised reconstruction.

[0035] In an optional embodiment, system 100 may include a frequency transformation (FT) module 150 (illustrated by dashed frame) which is adapted to transform 1120 each received training data record of the dataset 210 into a respective transformed training data record 210t in the frequency domain. In this embodiment, the transformed training data records 210t are used as the training data records for all subsequent steps in method 1000. The frequency transformation may be implemented by a Fourier Transform algorithm, or a constant Q-transform, to produce a compact time-frequency representation of the soundwaves which respects the logarithmic scaling of frequency resolution of the mammalian hear. For example, each soundwave of a training data record may be converted by using a Q-transform before being given to the encoder neural network. The Q-transform is a modified version of a spectrogram that outputs a logarithmically scaled frequency axis, employed to as to match Weber's law of sound frequency perception. In a constant-Q transform, the higher the frequencies, the wider the frequency grouping becomes. In an alternative implementation, the transformation module may be implemented by adding layers in the encoder neural network 110 to preprocess the soundwaves 210i through a filter bank whose parameters are learnt during the training of the encoder.

[0036] The encoder 110 processes 1200 each training data record 210i to output for each corresponding soundwave a one-dimensional encoded sound vector 191 of scalar values. As mentioned above the encoder output in the example implementation is a custom linear layer encompassing active unit filtering.

[0037] The encoded sound vector 191 is then mapped 1300 to the electrode array 220 to obtain a two-dimensional stimulation pattern 192 for the auditory cortical implant electrode array 220. In the example implementation a 10x10 stimulation pattern is generated from the encoded sound vector.

[0038] As in an autoencoder, the encoder input is basically the ground truth for a respective training data record, the decoder 130 now reconstructs 1400 the respective training data record for each encoder output by processing the respective encoded sound vector 191 or the corresponding stimulation pattern 192 with the decoder neural network 130. The information in the encoded sound vector 191 is equivalent to the information in the respectively mapped simulation pattern 192. Therefore, it does not matter which one is used as the decoder input. Thereby, the parameters of the decoder are determined in parallel to the encoder during training. The training is performed using a cost function CF which has at least two terms to evaluate reconstruction performance and continuity of the generated / obtained stimulation patterns 192.

[0039] In other words, the encoder part of the autoencoder network compresses the highdimensional input data into a low-dimensional latent space and the decoder part uses the compressed latent representation to reconstruct the original data. The encoder and decoder part of the autoencoder are constructed by a training process during which the connections between the processing units of the network are adjusted by an optimization algorithm based on many training samples until a high reconstruction performance is obtained. The reconstruction performance is measured through a cost function that computes the distance between original and reconstructed sound information. Training the neurons aims to minimize the cost function. The stimulation output of the autoencoder is the trained latent space that encodes a maximum of information in a given number of information channels in the form of the stimulation patterns. The decoder part of the network is only used during the training and evaluation of the network compression to ensure that the relevant information from the input data is included in the latent representation. This allows togenerate stimulation patterns for the auditory cortical implant electrode array by using a customized autoencoder network to compress sound information into a latent space that has as many channels as the electrode array. Further, the composition of the cost function allows to respect a crucial biological constraint (continuity) as demonstrated in the following.

[0040] The reconstruction performance for training data is measured 1500 through a first term CFr of the cost function CF by computing a first mathematical distance between any received training data record 210i and its corresponding reconstructed training data record 210r.

[0041] Continuity of the obtained stimulation patterns a, b, c is measured 1600 by computing a pairwise similarity matrix SMI between stimulation patterns a, b, c associated with a predefined set of reference soundwaves 210ref that are perceptually continuous. Thereby, the predefined set of reference soundwaves 210ref is a subset of the dataset 210. Advantageously, part of the reference soundwaves are pure tones in a range of audible frequencies, preferably between 100Hz -10000Hz, and amplitude modulations in a predefined range of frequencies, preferably from 1Hz to 200Hz. A second term CFc of the cost function is then used 1700 to enforce the continuity of the obtained stimulation patterns 192. The second term CFc computes a second mathematical distance between the computed similarity matrix SMI and a corresponding predefined reference similarity matrix SM2 that accounts for perceptual similarity. Continuity is a biological constraint which is crucial for an appropriate hearing experience of the patient. It means that the stimulation patterns for continuously changing sound features (e.g. sound frequency) must be continuous as already explained above. For example, if the frequency of two single frequency sounds changes by a value df, the similarity between the two corresponding stimulation patterns remain smaller or equal to g(df), where g(.) is a strictly decreasing function such that g(0) is equal to the similarity of identical vectors.

[0042] As mentioned earlier, the above-mentioned natural sound records can be used as additional training data for the generation of stimulation patterns addressing the human auditory cortex. To generate a corresponding reference similarity matrix for natural soundwaves, the system computes the similarity between two sound records as the fraction of sound labels which are shared by the two records. For example, if one natural sound record has 4 labels and the other has 2 labels, and one label is common to the two sounds,then the similarity is 1 / 5 (number of common labels / total number of labels). During training, after each batch of M sound records (typically M is in the order of 1000 sound records), the corresponding MxM similarity matrix is computed based on the sound encoding vectors, and the mean square error MSE with the reference similarity matrix for these M sound records is computed. The MSE is added to the loss to be minimized by the autoencoder model.

[0043] Alternatively, the decoder may be extended with a multilayer deep neural network that receives the sound encoding vectors from the encoder and produces a prediction of the sound labels as an output. The difference between the predicted and actual labels (MSE) becomes the continuity loss which is updated on each training step and minimized with the reconstruction loss.

[0044] In both cases, the encoder produces sound encoding vectors that reflect a perceptual continuity between sounds.

[0045] Turning briefly to FIG. 3, an example illustrates in more detail how to enforce continuity between pure tones (PT) or between amplitude-modulated (AM) tones. The example training data include a pure tone collection 310 with pure tone training data records, and may also include amplitude-modulated training data records, for which the encoder predicts the corresponding simulation patterns 320. Typically, the training is performed for batches of training data records. For each batch, the system computes a similarity matrix 330 of the predicted stimulation patterns for the subset of PT soundwaves (and AM soundwaves) in the respective batch. This similarity matrix is then compared to a corresponding reference similarity matrix which serves as ground truth. Then, the MSE is computed and added to the loss resulting in a respective loss value that allows to enforce continuity.

[0046] Finally, the encoder 110 and decoder 130 neural networks are trained 1800 while aiming at minimizing a weighted sum of the terms CFr, CFc of the cost function CF.

[0047] In optional embodiments, system 100 can be further improved with regard to the quality of the generated stimulation patterns by taking into account further constraints. FIG. 4A illustrates an embodiment, where a third term CFt of the cost function CF takes into account the tonotopy of the auditory cortex. Tonotopy is the spatial arrangement of wheresounds of different frequency are processed in the brain. Tones close to each other in terms of frequency are represented in topologically neighboring regions in the brain. Tonotopic maps are a particular case of topographic organization. For healthy mammals, tonotopy in the auditory system begins at the cochlea that sends information about sound to the brain. Different regions of the basilar membrane in the organ of Corti, the sound-sensitive portion of the cochlea, vibrate at different sinusoidal frequencies due to variations in thickness and width along the length of the membrane. Nerves that transmit information from different regions of the basilar membrane therefore encode frequency tonotopically. This tonotopy then projects through the vestibulocochlear nerve and associated midbrain structures to the primary auditory cortex via the auditory radiation pathway. Throughout this radiation, organization is linear with relation to placement on the organ of Corti, in accordance to the best frequency response (that is, the frequency at which that neuron is most sensitive) of each neuron. However, binaural fusion in the superior olivary complex onward adds significant amounts of information encoded in the signal strength of each ganglion. Thus, the number of tonotopic maps varies between species and the degree of binaural synthesis and separation of sound intensities. For example, in humans, six tonotopic maps have been identified in the primary auditory cortex.

[0048] The third term CFt of the cost function computes a third mathematical distance (tonotopic loss) between known spatial patterns of sound-driven responses 410 to pure tones in auditory cortex regions targeted by the implant and the respective predicted (mapped) stimulation patterns, and enforces 1720 that the obtained stimulation patterns 420 respect multiple pure tone frequency gradients and amplitude frequency gradients observed in the auditory cortex of the human brain. In this embodiment, the weighted sum of cost function terms to be minimized also takes into account the third term of the cost function when training the autoencoder. FIG. 4B demonstrates the effectiveness of the tonotopic loss by forcing the encoder model to create stimulation patterns for pure tones of different frequencies (0.5kHz, 2kHz, ..., 28kHz) that are aligned on a single axis such that frequency changes continuously when moving from one side of the axis to the other side. This is shown in diagram 440 which illustrates center of masses of 96 pure tone stimulation patterns (latent representations) after training highlighting the enforced one-axis tonotopy of the latent representations. The color scale 441 represents the tone frequencies between 0.5kHz and 28kHz. The dark bullets on the left represent the respective centers of mass ofsimulation patterns for lower tone frequencies and the bright bullets on the right for higher tone frequencies. It can be seen that the centers of mass are lined up along an axis from left to right. Pure tone responses 430 were modeled aligned along on one axis. When constraining the model to reproduce such spatial organization for sample pure tones, a spatial arrangement of pure tone responses was obtained along one axis as shown by plotting the centroids of the responses to many pure tones, color-coded according to their frequency (diagram 440).

[0049] In one embodiment, the autoencoder 140 includes a spatial block (optional channel reduction module (CRM) 120). CRM 120 is placed between the encoder output and the decoder input layers. CRM 120 is adapted to limit 1740 the number of simultaneously active channels of each obtained stimulation pattern 192 to a predefined number of most active channels by keeping only the predefined number of most active channels of the encoder output with other channels being set to zero. That is, the electrodes of the electrode array 220 which are with a zero channel will not send stimulation signals to the auditory cortex of the patient. In this embodiment, the encoder output with the limited number of simultaneously active channels is also being used as input to the decoder for training data reconstruction. It is to be noted that in FIG. 1, the arrow from optional CRM 120 to the stimulation pattern 192 is to interpreted such that, in case that CRM 120 is present, the obtained stimulation pattern only includes the corresponding reduced number of active channels, whereas in the basic embodiment of system 100 (without CRM 120), the entire encoded output vector 191 is mapped to the stimulation pattern without any limitation of the active channels. The active channels in a stimulation pattern are indicated as grey colored squares where different shades of grey represent different activity strength of the respective channels - the darker, the more active. As already mentioned, the embodiment with CRM 120 reduces energy consumption and channel crosstalk which can be seen as advantageous effects regarding the technical implementation of the mapping of encoded sound vector to a corresponding stimulation pattern. It further reduces the risk for a clinical application that the stimulator overdrives the brain and leads to a deterioration of the patient's health state (e.g., by triggering epileptic seizure), thus also achieving an advantageous effect on the patient's medical state. Further, it has been shown that only a few active channels are necessary for an accurate reconstruction of the soundwaves.

[0050] The CRM 120 embodiment can be further improved by adding a blurring function to CRM 120 to blur 1760 the pattern of the active channels in each obtained stimulation pattern with a two-dimensional kernel 121 of a predefined width, to account for the physical resolution of the stimulation. Typically, the active area of the auditory cortex is larger than the respective stimulating electrode. The blurring 1760 ensures that the auditory cortex implant can work efficiently, by avoiding that an obtained stimulation pattern concentrates the pixels of very different sounds in nearby regions. This is achieved by blurring the pattern of the active channels with the 2D kernel 121 (e.g., a Gaussian kernel) of width W. In the embodiment with the M most active channels, the pattern with these most active channels is blurred. 2-photon recordings in the auditory cortex show (cf. FIG. 5) that the activity evoked by pure tone is correlated at 50% when the sounds are separated by 0.33 octaves. This measure is a reference for computing the spatial blurring kernel (o = 1.32).

[0051] FIG. 5 illustrates details of the blurring function 121 between the encoder and decoder by way of an example. A stimulation pattern which is obtained by the mapping of the encoded sound vector may be reduced to the reduced stimulation pattern 192r which has 10 active channels with an activity in the range from 0 and 1 (as indicated by the color scale 125). The reduced stimulation pattern is then convoluted with the 2D-kernel 121. The result is the blurred stimulation pattern 192b.

[0052] In addition, a person skilled in the field of deep neural networks can take advantage of impressive optimization capabilities of deep neural networks to make the encoder resistant to moderate noise background by training the network with input soundwaves superposed with different types of noise while the respective reconstruction target (target output) is noise-free.

[0053] In an example implementation, the autoencoder network 140 (with all optional components disclosed herein) was trained through 7500 iterations while monitoring training and validation loss. The corresponding model comprises four convolutional layers for both the encoder and decoder parts. A reconstruction error is computed between the noisy input spectrogram and the denoised output. Latent representations are manipulated using a continuity cost function term and a tonotopic cost function term. The latent representation is spatially filtered with a Gaussian kernel to mimic the spread of neuronal stimulation in thebrain. Thanks to this filtering step, the optimized encoder generates stimulation patterns whose precision is less affected by the spread of neuronal stimulation.

[0054] The loss over iterations is shown in FIG. 8 for training data records (black curve 80-1) and validation data records (grey curve 82-2) used for training the autoencoder neural network for the auditory cortex of mice. It shows that the network can still converge and remain stable for thousands of iterations. In the example implementation, the encoder was designed for a 10x10 channels stimulator aiming to target the mouse auditory cortex (i.e., 100 stimulation channels in the obtained stimulation patterns). The stimulator was then used in mice. The target similarity matrix for the continuity constraint was built based on measurements of sound representations in mice (cf., A spatial code for temporal cues is necessary for sensory learning, Bagur et al., bioRxiv 2022.12.14.520391; doi: https: / / doi.org / 10.1101 / 2022.12.14.520391) (Bagur et al. 2023).

[0055] The spatial mapping was simulated as a single linear gradient frequency following one dimension of the 2D array (cf., FIG. 9). A sparsity constraint was set as a maximum of 10 co-activated channels, and the effective resolution constraint was applied with a 3x3 channel 2D Gaussian blurring kernel which a standard deviation of 1.32 (cf. FIG. 4).

[0056] The training and test sets were constituted of 100,000 laboratory-generated soundwaves (as described earlier as the example training dataset), each lasting 500ms, including pure tones, amplitude-modulated tones, chords, chirps, and frequency steps. Each soundwave was produced to be at 70dB SPL. The encoder was trained to generate one activity pattern for each soundwave, meaning that the temporal resolution of the latent representation becomes 500ms, which accounts for the fact that in the auditory cortex temporal details are encoded by specific spatial patterns (cf., Bagur et al. 2023). This results in a significant compression of the spectro-temporal information over a long timescale. Resistance to moderate noise was performed by adding different types of background noises to every sound at 60dB SPL while the target reconstruction sound was noise-free. The type of the noise used was either white noise or background sounds extracted from an environmental natural sound dataset. Every soundwave was converted using a Q-transform before being given to the autoencoder network.

[0057] The loss function (cost function for optimization) was :LOSS=0.85*MSE(X, Y)+0.05*CROSSPT+0.05*CROSSAM+TONOLOSS where MSE is the mean square error between the reconstructed Q-transform spectrogram and the target Q-Transform spectrogram. CROSSPT and CROSSAM are the mean squared errors between target and observed similarity matrices for pure tones, and amplitude modulation representations, respectively. Similarity was measured as the Pearson correlation coefficient correlation between latent space activations of the two sounds. TONOLOSS refers to the mean squared error between the predicted position of the centroid of the stimulation pattern on the 10x10 array and the centroid position of the encoder's output pattern.

[0058] At the end of training, the loss was below 0.01 (cf. FIG. 8) meaning that the reconstruction error was below 1%, indicating the excellent performance of the autoencoder network and demonstrating that the multi-optimization process implemented during training was successful. This was also verified for the two continuity and tonotopic constraints imposed to the neural network through respective terms added to the cost (loss) function. FIG 6 represents the similarity 610 between the encoding vectors of sample pure tones observed after training the model and its match with the target similarity 620 for the same tones derived from Bagur et al. 2023 in an embodiment which uses a cost function with all three terms CFr, CFc and CFt, as well as channel reduction and blurring function. As mentioned above, FIG. 4B demonstrates the efficiency of the tonotopic loss in the same embodiment.

[0059] FIG. 7 shows three examples of reconstructed soundwave spectrograms 72 and the corresponding originals 71 for an embodiment which uses a cost function with all three terms CFr, CFc and CFt, as well as channel reduction and blurring function. The autoencoder model can reconstruct an almost perfect denoised version of the original input data.

[0060] FIG. 9 shows a photo 12 of an embodiment of an auditory cortical implant which has a predefined number of electrodes arranged in a two-dimensional electrode array 12-1 adapted for use in mammals with impaired hearing. The implant is adapted for high resolution electrical stimulation wherein the electrodes are arranged as a flexible two- dimensional surface stimulator adapted for positioning on the surface of a mammal's brain. The high resolution of the surface stimulator is achieved through a contact density in therange of 1 to 40 electrodes per square millimeter and an electrode-to electrode distance in the range of 160pm to 1000pm.

[0061] The electrical stimulation of the electrodes is controlled 12c by stimulation patterns 12-3 which are received by the implant via an interface 12-2 from an encoder (e.g. encoder 110 of FIG. 1). Each received stimulation pattern encodes a particular soundwave into a two- dimensional array with dimensions corresponding to the two-dimensional electrode array such that stimulation patterns for continuously changing sound frequencies are continuous in that, if the frequency of two single frequency sounds changes by a value df, the similarity between the two corresponding stimulation patterns remain smaller or equal to g(df), where g(.) is a strictly decreasing function such that g(0) is equal to the similarity of identical vectors. In the example, the control stimulation patterns 12-3 are generated for a series of pure tones ranging from 6kHz (top left) to 16kHz (bottom right).

[0062] In one embodiment, the received stimulation patterns 12-3 further have a minimized mathematical distance between known spatial patterns of sound-driven responses in auditory cortex regions targeted by the implant, thereby respecting multiple pure tone frequency gradients and amplitude frequency gradients observed in the auditory cortex of the human brain.

[0063] In one embodiment, the number of simultaneously active channels of each received stimulation pattern 12-3 is limited to a predefined number of most active channels. In the example, each control pattern has 10 active channels illustrated by grey squares. As describe above, the pattern of the predefined number of most active channels in each obtained stimulation pattern may be blurred with a two-dimensional kernel of a predefined width, to account for the physical resolution of the stimulation.

[0064] In the following, some experimental results are described which were achieved with mice as mammal patients. The two-dimensional electrode array was made of parylene C (PaC), a biocompatible and mechanically flexible polymer of a 4 pm thickness (total probe thickness), which is highly conformable to the brain's mechanical properties and curvature. The electrodes are coated with PEDOT:PSS, a highly conductive and biocompatible polymer, which increases the signal-to-noise-ratio (SNR) by enhancing the conversion of ionic (in biology) to electronic (in electronic systems) current. Hence, PEDOT improves the interfacing of the implant with the brain and substantially increases the sensitivity. The electrode arraycontains 32 electrodes, with a size of 20x20 pm2, distributed in a 700x875 pm2 area. The edge-to-edge spacing between electrodes is 175 pm. The electrode array was successfully implanted over the mouse auditory cortex and showed that an increase of electrical stimulation's strength gives rise to a stronger neural response in the auditory cortex. It was verified that the neurons' health was not affected after the electrical stimulation.

[0065] Implanted mice were then trained to respond with a lick on a reward spout if they detected the cortical stimulation of one electrode. Stimulation used a 1000ms train of electrical pulses (0.1 mA) at 100Hz. Once the mice achieved this training goal (>75% hits), they were trained to perform a discrimination of spatial and temporal stimulation patterns. The spatial discrimination task consisted in learning to lick on the reward spout if one of four electrodes placed in the low frequency region of the auditory cortex was activated and not to lick if one of four electrodes placed in the high frequency region of the auditory cortex was activated. After about 15 days of training, the mice were able to discriminate very well between the rewarded and non-rewarded electrodes. The day after the pure electrical tone task, the mice were presented with pure tone sounds. Interestingly, they were also able to discriminate between different sound frequencies without any previous sound discrimination learning. This indicates that the perception generated by the implant is close to the perception of the sounds that usually activate the stimulated areas. These experiments showed that the spatial resolution of the implant is as good as 175pm (the distance between two neighboring electrodes), which was used as the minimal distance between rewarded and non-rewarded electrodes. These preclinical experiments demonstrate that cortical stimulation is perceived by the herein disclosed approach, and is similar to sounds.

[0066] The mice were also tested to discriminate temporal information. In this task, the animals learnt to discriminate between two groups of pulse amplitude modulation frequencies (20 to 80Hz rewarded, 100 to 200 Hz non rewarded) for the stimulation of a single electrode. They achieved this also with a good efficiency indicating that the electrical implant has an excellent temporal resolution of at least 100Hz. Once the mice have learnt this task, the background activity was superimposed to the amplitude modulation to simulate the masking of the temporal signal resulting from background noise. These resultsindicated that the resulting perception was resistant to a certain degree of noise, up to a simulated 60dB noise level for a 70dB SPL signal.

[0067] To further demonstrate that cortical stimulation results in high resolution perception, the perceptual resolution was compared with a 4-electrodes cochlear implant (500pm electrode distance) inserted in the cochlea of another set of mice. Stimulation protocols were identical with cortical and cochlear implants. It was observed that spatial and temporal resolution was better with cortical than with cochlear implants. This result is striking considering that the inter-electrode distance in the cochlear implant is 500pm while it is 175pm for the cortical implant.

[0068] FIG. 10 is a diagram that shows an example of a generic computer device 900 and a generic mobile computer device 950, which may be used with the techniques described here. Computing device 900 relates in an exemplary embodiment to the system 100 for training a respective encoder (cf. FIG. 1). Computing device 950 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, and other similar computing devices. In an exemplary embodiment of this disclosure, the computing device 950 may serve as a frontend control device of the system 900 which can be used by the user to interact with the training system. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and / or claimed in this document.

[0069] Computing device 900 includes a processor 902, memory 904, a storage device 906, a high-speed interface 908 connecting to memory 904 and high-speed expansion ports 910, and a low-speed interface 912 connecting to low-speed bus 914 and storage device 906. Each of the components 902, 904, 906, 908, 910, and 912, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 902 can process instructions for execution within the computing device 900, including instructions stored in the memory 904 or on the storage device 906 to display graphical information for a GUI on an external input / output device, such as display 916 coupled to high-speed interface 908. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices 900 may be connected, with each deviceproviding portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0070] The memory 904 stores information within the computing device 900. In one implementation, the memory 904 is a volatile memory unit or units. In another implementation, the memory 904 is a non-volatile memory unit or units. The memory 904 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0071] The storage device 906 is capable of providing mass storage for the computing device 900. In one implementation, the storage device 906 may be or contain a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. A computer program product can be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 904, the storage device 906, or memory on processor 902.

[0072] The high-speed controller 908 manages bandwidth-intensive operations for the computing device 900, while the low-speed controller 912 manages lower bandwidthintensive operations. Such allocation of functions is exemplary only. In one implementation, the high-speed controller 908 is coupled to memory 904, display 916 (e.g., through a graphics processor or accelerator), and to high-speed expansion ports 910, which may accept various expansion cards (not shown). In the implementation, low-speed controller 912 is coupled to storage device 906 and low-speed expansion port 914. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, ZigBee, WLAN, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0073] The computing device 900 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 920, or multiple times in a group of such servers. It may also be implemented as part of a rack server system 924. In addition, it may be implemented in a personal computer such as a laptop computer 922. Alternatively, components from computing device 900 may becombined with other components in a mobile device (not shown), such as device 950. Each of such devices may contain one or more of computing device 900, 950, and an entire system may be made up of multiple computing devices 900, 950 communicating with each other.

[0074] Computing device 950 includes a processor 952, memory 964, an input / output device such as a display 954, a communication interface 966, and a transceiver 968, among other components. The device 950 may also be provided with a storage device, such as a microdrive or other device, to provide additional storage. Each of the components 950, 952, 964, 954, 966, and 968, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0075] The processor 952 can execute instructions within the computing device 950, including instructions stored in the memory 964. The processor may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor may provide, for example, for coordination of the other components of the device 950, such as control of user interfaces, applications run by device 950, and wireless communication by device 950.

[0076] Processor 952 may communicate with a user through control interface 958 and display interface 956 coupled to a display 954. The display 954 may be, for example, a TFT LCD (Thin-Film-Transistor Liquid Crystal Display) or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 956 may comprise appropriate circuitry for driving the display 954 to present graphical and other information to a user. The control interface 958 may receive commands from a user and convert them for submission to the processor 952. In addition, an external interface 962 may be provided in communication with processor 952, so as to enable near area communication of device 950 with other devices. External interface 962 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0077] The memory 964 stores information within the computing device 950. The memory964 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. Expansion memory984 may also be provided and connected to device 950 through expansion interface 982, which may include, for example, a SIMM (Single In Line Memory Module) card interface. Such expansion memory 984 may provide extra storage space for device 950, or may also store applications or other information for device 950. Specifically, expansion memory 984 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, expansion memory 984 may act as a security module for device 950, and may be programmed with instructions that permit secure use of device 950. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing the identifying information on the SIMM card in a non-hackable manner.

[0078] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one implementation, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 964, expansion memory 984, or memory on processor 952, that may be received, for example, over transceiver 968 or external interface 962.

[0079] Device 950 may communicate wirelessly through communication interface 966, which may include digital signal processing circuitry where necessary. Communication interface 966 may provide for communications under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communication may occur, for example, through radio-frequency transceiver 968. In addition, short-range communication may occur, such as using a Bluetooth, WiFi, ZigBee or other such transceiver (not shown). In addition, GPS (Global Positioning System) receiver module 980 may provide additional navigation- and location- related wireless data to device 950, which may be used as appropriate by applications running on device 950.

[0080] Device 950 may also communicate audibly using audio codec 960, which may receive spoken information from a user and convert it to usable digital information. Audio codec 960 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of device 950. Such sound may include sound from voice telephone calls, may 1include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on device 950.

[0081] The computing device 950 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 980. It may also be implemented as part of a smart phone 982, personal digital assistant, or other similar mobile device.

[0082] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0083] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" "computer-readable medium" refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine- readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0084] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g.,visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0085] The systems and techniques described here can be implemented in a computing device that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0086] The computing device can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0087] A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the scope of the invention.

[0088] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems.

Claims

Claims1. A computer-implemented method (1000) for training an encoder neural network (110) to generate stimulation patterns (192) for an auditory cortical implant having a predefined number of electrodes arranged in a two-dimensional electrode array (220) adapted for use in mammals (201) with impaired hearing, comprising: receiving (1100) a dataset (210) of soundwaves as training data records for the encoder (110); processing (1200), by the encoder (110), each training data record (210i) to output for each corresponding soundwave a one-dimensional encoded sound vector (191) of scalar values; mapping (1300) the encoded sound vector (191) to the electrode array (220) to obtain a two-dimensional stimulation pattern (192) for the auditory cortical implant; reconstructing (1400) the respective training data record for each encoder output by processing the respective encoded sound vector (191) or the corresponding stimulation pattern (192) with a decoder neural network (130), wherein the parameters of the decoder are determined in parallel to the encoder during training; measuring (1500) reconstruction performance for training data through a cost function with a first term (CFr) that computes a first mathematical distance between any received training data record (210i) and its corresponding reconstructed training data record (210r); measuring (1600) continuity of obtained stimulation patterns (a, b, c) by computing a pairwise similarity matrix (SMI) between stimulation patterns (a, b, c) associated with a predefined set of reference soundwaves (210ref) that are perceptually continuous, wherein the predefined set of reference soundwaves (210ref) is a subset of the dataset (210); using (1700) a second term (CFc) of the cost function to enforce the continuity of the obtained stimulation patterns (192), wherein the second term (CFc) computes a second mathematical distance between the computed similarity matrix (SMI) and acorresponding predefined reference similarity matrix (SM2) that accounts for perceptual similarity; and training (1800) the encoder (110) and decoder (130) neural networks while aiming at minimizing a weighted sum of the terms (CFr, CFc) of the cost function.

2. The method of claim 1, wherein the cost function further has a third term (CFt) corresponding to a third mathematical distance between known spatial patterns (410) of sound-driven responses in auditory cortex regions targeted by the implant and the stimulation patterns obtained from the respective encoded sound vectors, to enforce (1720) that the obtained stimulation patterns respect multiple pure tone frequency gradients and amplitude frequency gradients observed in the auditory cortex of the human brain, wherein the weighted sum also takes into account the third term of the cost function.

3. The method of any of the previous claims, further comprising: limiting (1740) the number of simultaneously active channels of each obtained stimulation pattern to a predefined number of most active channels by keeping only the predefined number of most active channels of the encoder output with other channels being set to zero, wherein the encoder output with the limited number of simultaneously active channels also being used as input to the decoder for training data reconstruction.

4. The method of any of the previous claims, further comprising: blurring (1760), by a spatial block (120) between encoder (110) and decoder (120), a pattern of active channels in each obtained stimulation pattern with a two- dimensional kernel (121) of a predefined width, to account for the physical resolution of the stimulation, wherein the physical resolution is measured as the spatial spread of neuronal activity generated by stimulation with one electrode of the cortical implant.

5. The method of any of the previous claims, further comprising: transforming (1120) each received training data record of the dataset (210) into a respective transformed training data record in the frequency domain and using thetransformed training data records (210t) as the training data records for all subsequent steps from processing (1200) to training (1800).

6. The method of claim 5, wherein transforming into the frequency domain is achieved by: using a Fourier Transform algorithm, or a constant Q-transform, to produce a compact timefrequency representation of the soundwaves which respects the logarithmic scaling of frequency resolution of a respective mammal's hearing system, or by adding layers in the encoder neural network to preprocess the soundwaves through a filter bank whose parameters are learnt during the training of the encoder.

7. The method of any of the previous claims, wherein the encoder neural network is implemented by a plurality of convolutional layers, each convolutional layer being a custom linear filter bank followed by an activation layer.

8. The method of any of the previous claims, wherein the reference soundwaves are pure tones in a range of audible frequencies in the range of 16Hz -20000Hz, and amplitude modulations in a predefined range of frequencies, preferably from lOHz to 200Hz.

9. The method of any of the previous claims, wherein the training data records comprise a dataset of natural soundwave records in which each natural soundwave record is labelled with an object or action recognized in the soundwave by a human, further comprising: generating a reference similarity matrix for the natural soundwaves by computing the similarity between pairs of soundwave records as the fraction of sound labels which are shared by the two records of each respective pair; during training, computing after each batch of M natural soundwave records, a corresponding MxM similarity matrix based on the respective sound encoding vectors or their transformation by an additional neural network whose weights are optimized with the encoder and decoder neural networks, and computing the mean square error with the reference similarity matrix for the M soundwave records; and adding the mean square error to the cost function to be minimized by the encoder and decoder neural networks.

10. The method of any of the previous claims, wherein the two-dimensional stimulation patterns match the dimensions of the auditory cortical implant which is adapted for highresolution electrical stimulation with the electrodes being arranged as a flexible two- dimensional surface stimulator adapted for positioning on the surface of a mammal's brain, wherein the surface stimulator has a contact density in the range of 1 to 40 electrodes per square millimeter and an electrode-to electrode distance in the range of 160pm to 1000pm.

11. A computer program product for training an encoder neural network to generate stimulation patterns for an auditory cortical implant, comprising computer-readable instructions that, when loaded into a memory of a computing device and executed by one or more processors of the computing device, cause the computing device to perform the computer-implemented method according to any of the previous claims.

12. A computer system for training an encoder neural network to generate stimulation patterns for an auditory cortical implant having a predefined number of electrodes arranged in a two-dimensional electrode array adapted for use in mammals with impaired hearing, the system comprising a memory storing computer-readable instructions implementing functional units, and one or more processors adapted for executing the functional units, wherein the functional units are adapted to perform the computer-implemented method according to any of the claims 1 to 10 when being executed.

13. An auditory cortical implant having a predefined number of electrodes arranged in a two-dimensional electrode array (12-1) adapted for use in mammals with impaired hearing, the implant adapted for electrical stimulation wherein the electrodes are arranged as a flexible two-dimensional surface stimulator adapted for positioning on the surface of a mammal's brain, wherein the surface stimulator has a contact density in the range of 1 to 40 electrodes per square millimeter and an electrode-to-electrode distance in the range of 160pm to 1000pm, and wherein the electrical stimulation is controlled (12c) by stimulation patterns (12-3), the stimulation patterns received by the implant from an encoder, wherein each stimulation pattern encodes a particular soundwave into a two-dimensional array with dimensions corresponding to the two-dimensional electrode array such that stimulation patterns for continuously changing sound frequencies are continuous in that, if the frequency of two single frequency sounds changes by a value df, the similarity between the two corresponding stimulation patterns remain smaller or equal to g(df), where g(.) is a strictly decreasing function such that g(0) is equal to the similarity of identical vectors.

14. The auditory cortical implant of claim 13, wherein the received stimulation patterns (12- 3) further have a minimized mathematical distance between known spatial patterns of sound-driven responses in auditory cortex regions targeted by the implant, thereby respecting multiple pure tone frequency gradients and amplitude frequency gradients observed in the auditory cortex of the human brain.

15. The auditory cortical implant of claim 13 or 14, wherein the number of simultaneously active channels of each received stimulation pattern (12-3) is limited to a predefined number of most active channels, and wherein the pattern of the predefined number of most active channels in each obtained stimulation pattern is blurred with a two-dimensional kernel of a predefined width, to account for the physical resolution of the stimulation.

16. A hearing restoration system for mammals with impaired hearing, comprising: an encoder neural network (110) adapted to generate, in response to sensed soundwaves, corresponding stimulation pattern (192) for an auditory cortical implant adapted for use in a mammal (201) with impaired hearing, wherein each stimulation pattern encodes a particular soundwave into a two-dimensional array with dimensions corresponding to the two-dimensional electrode array such that stimulation patterns for continuously changing sound frequencies are continuous in that, if the frequency of two single frequency sounds changes by a value df, the similarity between the two corresponding stimulation patterns remain smaller or equal to g(df), where g(.) is a strictly decreasing function such that g(0) is equal to the similarity of identical vectors; and the auditory cortical implant having a predefined number of electrodes arranged in a two-dimensional electrode array (12-1), the implant adapted for electrical stimulation of the mammal's auditory cortex wherein the electrodes are arranged as a flexible two-dimensional surface stimulator adapted for positioning on the surface of the mammal's brain, wherein the surface stimulator has a contact density in the range of 1 to 40 electrodes per square millimeter and an electrode-to-electrode distance in the range of 160pm to 1000pm, and wherein the electrical stimulation is controlled (12c) by the stimulation patterns (12-3) received by the implant from the encoder.

Citation Information

Patent Citations

  • Neural Network Audio Scene Classifier for Hearing Implants

    US20210174824A1