Method based on artificial neural network and its uses, auditory device, and computer program product
Through the individual auditory response model based on artificial neural network, the hearing aid algorithm is solved, and the problem of hearing impairment such as synapses is difficult to compensate for hearing impairments, which is achieved to improve speech clarity and effective treatment of auditory periphery, and is suitable for cochlear implants and wearable hearing aids.
Patent Information
- Application Number
- CN202180026269.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-01
- Filing Date
- 2021-04-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-04-01
AI Technical Summary
Existing hearing aid algorithms are difficult to accurately compensate for different types of hearing impairments, especially sensory neurological hearing loss caused by synapses, and lack individualized and effective auditory signal processing methods, resulting in insufficient recovery of speech clarity.
Using an artificial neural network-based method, by generating a personalized auditory response model, using a differentiable auditory response poor to train a neural network model, and developing an individualized auditory signal processing model to match the expected auditory response, compensate for hearing loss and improve speech clarity.
It achieves individual compensation for different hearing impairments, improves speech clarity, especially hearing loss caused by synapses, enhances the processing ability of the auditory peripherals, and is suitable for hearing equipment such as cochlear implants and wearable hearing aids.
Smart Images

Figure CN115362689B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of auditory devices. In particular, the present invention provides a method for converting an auditory stimulus into a processed auditory output. The present invention also relates to the use of the method, an auditory device configured to perform the method, and a computer program configured to perform the method to convert an auditory stimulus into a processed auditory output. Background Art
[0002] Over the past decade, the number of people suffering from hearing loss has been steadily increasing, and society has been continuously exposed to increasingly noisy environments and lifestyles. However, despite extensive research on the compensation of cochlear gain loss in the past few years, the correct diagnosis and treatment of hearing impairments remain unclear. To address this issue, computational models of the human auditory periphery can be used as tools for developing effective auditory signal processing algorithms aimed at restoring degraded speech auditory representations caused, for example, by outer hair cell loss. At the same time, these computational models can benefit the new "enhanced hearing" field, in which sound signals are transformed in such a way as to enhance the listener's hearing experience. Audio signal processing manipulations inspired by the models can improve sound perception or sound quality, or incorporate noise reduction or other manipulations. However, it is still not straightforward to design such processing methods that can accurately compensate for different types of hearing impairments or create effective enhanced hearing algorithms for processing complex stimuli such as speech.
[0003] Provide an example of audio signal processing in hearing aids: Hearing aid algorithms are typically optimized to compensate for frequency-specific damage to outer hair cells in the inner ear (or cochlea), such as the NAL-NL or DSL schemes. Therefore, the signal processing algorithms do not incorporate important aspects of sensorineural hearing loss related to damaged synapses (synaptopathy) between inner hair cells and auditory nerves in the cochlea. At the same time, currently few include metrics from biophysical signals such as otoacoustic emissions (OAE), middle ear muscle reflex (MEMR) responses, or auditory evoked potentials (AEP) to individualize the processing of hearing aid algorithms.
[0004] Several attempts have been made to automate and predict the auditory performance of basic human sound perception tasks. Experiments of this type are time-consuming to conduct, so it is beneficial to use listener models instead. These systems typically use (individualized) auditory models (front-ends) as input to a task simulation system (back-end), which is usually an automatic speech recognition (ASR) system that can be used to train and predict the task performance (i.e., psychoacoustics) of sound perception tasks. Psychoacoustic tasks are used to objectively quantify an individual's sound perception ability, and a typical task is to measure speech intelligibility in noise, i.e., to determine the SNR threshold at which a listener can correctly identify 50% of the words in a sentence. However, it remains a challenge to develop systems that can predict the results of different experiments and generalize well to listeners, taking into account various aspects such as their hearing impairment or language. Summary of the Invention
[0005] The present invention overcomes one or more of these problems. Preferred embodiments of the present invention overcome one or more of these problems.
[0006] Advantages of embodiments of the present invention are that they explain how synaptopathy affects supra-threshold speech coding and help individuals who do not fully recover speech intelligibility based solely on gain prescriptions.
[0007] Advantages of embodiments of the present invention are that the individual-based recovery algorithm for synaptopathy provides a means to help improve the speech intelligibility of self-reporting listeners with normal audiograms who are currently untreated.
[0008] Advantages of embodiments of the present invention are that the model-based processing algorithm takes into account the individual degree of synaptopathy and other aspects of sensorineural hearing loss.
[0009] Advantages of embodiments of the present invention are that they can include both OAE and AEP metrics to construct an individualized hearing loss model that will serve as the basis for the processing algorithm.
[0010] Advantages of embodiments of the present invention are that they include an NN-based auditory model that can provide a differentiable auditory response.
[0011] Advantages of embodiments of the present invention are that they include an NN-based auditory model that can accurately describe the processing (auditory processing) of the auditory periphery in a biophysically inspired manner.
[0012] Advantages of embodiments of the present invention are that they include an NN-based auditory model that can capture the characteristics of the auditory periphery, up to the level of inner hair cell and auditory nerve processing and the resulting population responses.
[0013] Advantages of embodiments of the present invention are that they include NN-based auditory models that can include combinations of outer hair cell damage, inner hair cell damage, cochlear synaptopathy, or even hearing loss at all different stages in the auditory periphery.
[0014] Advantages of embodiments of the present invention are that they include NN-based auditory models that can simulate auditory brainstem responses and thus provide the ability to generate a generator that restores auditory evoked potentials.
[0015] Advantages of embodiments of the present invention are that they use accurate NN-based auditory models as input to an NN-based automatic speech recognition (ASR) system to simulate and compensate for the degraded performance of hearing-impaired listeners in speech intelligibility tasks.
[0016] Advantages of embodiments of the present invention are that they use a closed-loop method based on the aforementioned NN-based auditory model to generate an NN-based processing model that can minimize a suitably designed metric that reflects the degraded hearing ability and perception of a human listener.
[0017] The present invention relates to a method based on an artificial neural network for obtaining an individualized auditory signal processing model suitable for converting an auditory stimulus into a processed auditory output. The method preferably comprises the following steps:
[0018] a. obtaining, preferably generating, an NN-based personalized auditory response model that represents the expected auditory response of a subject with an auditory profile to an auditory stimulus;
[0019] b. comparing the output of the personalized auditory response model with the output of an NN-based desired auditory response model to determine an auditory response difference; whereby the auditory response difference is differentiable, i.e., it can be used to train / develop an NN model that can backpropagate to a solution; and,
[0020] c. using the determined differentiable auditory response difference to develop an NN-based individualized auditory signal processing model for the subject, wherein the individualized auditory signal processing model is configured to minimize the determined auditory response difference.
[0021] The method can thus obtain an individualized auditory signal processing model that can process an auditory stimulus to generate a processed auditory output that, when given as input to the individualized auditory response model or to the subject, matches the desired auditory response.
[0022] The present invention also relates to an artificial neural network-based method for converting an auditory stimulus into a processed auditory output. The method preferably includes the steps of obtaining an individualized auditory signal processing model or an implementation thereof as described herein; and,
[0023] d. Applying the individualized neural network-based auditory signal processing model to the auditory stimulus to generate a processed auditory output that preferably matches the desired auditory response when given as an input to the personalized auditory response model or the subject.
[0024] The present invention also relates to an artificial neural network-based method for obtaining an individualized auditory signal processing model suitable for converting an auditory stimulus into a processed auditory output, the method comprising the steps of:
[0025] a. Generating a neural network-based personalized auditory response model based at least on the integrity of the subject's auditory nerve fibers (ANF) and / or auditory nerve synapses (ANS), preferably also based on the integrity of the subject's inner hair cell (IHC) damage and / or outer hair cell (OHC) damage; the personalized auditory response model represents the expected auditory response of the subject with an auditory distribution to an auditory stimulus;
[0026] b. Comparing the output of the personalized auditory response model with the output of a neural network-based desired auditory response model to determine an auditory response difference; wherein the neural network-based model consists of non-linear operations that make the auditory response difference differentiable;
[0027] c. Using the determined differentiable auditory response difference to develop a neural network-based individualized auditory signal processing model for the subject, wherein the individualized auditory signal processing model is configured to minimize the determined auditory response difference; and,
[0028] d. Applying the individualized neural network-based auditory signal processing model to the auditory stimulus to generate a processed auditory output that matches the desired auditory response when given as an input to the personalized auditory response model or the subject.
[0029] In some preferred embodiments, the personalized auditory response model in step a. is determined by deriving and including an auditory distribution specific to the subject.
[0030] In some preferred embodiments, the auditory distribution specific to the subject is an auditory damage distribution specific to the subject; preferably based on the integrity of the auditory nerve fibers (ANF) and / or auditory nerve synapses (ANS), and / or based on the outer hair cell (OHC) damage of the subject.
[0031] In some preferred embodiments, the desired auditory response is a response from a normally hearing subject or a response with enhanced characteristics.
[0032] In some preferred embodiments, the desired auditory response model and the personalized auditory response model include models of different stages of the auditory periphery.
[0033] In some preferred embodiments, a reference neural network describing the normally hearing auditory periphery is used as the desired auditory response model; a corresponding neural network for hearing-impaired is used as the personalized auditory response model; and the individualized auditory signal processing model is the following signal processing neural network model, which is trained to process auditory inputs and compensate for the degraded output of the hearing-impaired model when connected to the input of the hearing-impaired model or the subject.
[0034] In some preferred embodiments, a reference neural network simulating the enhanced hearing perception and / or ability of a normally hearing listener is used as the desired auditory response model; a corresponding neural network for normally hearing or hearing-impaired is used as the personalized auditory response model; and the individualized auditory signal processing model is a signal processing neural network model trained to process auditory inputs and provide an enhanced auditory response.
[0035] In some preferred embodiments, the individualized auditory signal processing model is trained to minimize a specific auditory response difference metric, such as the absolute difference or the squared difference between two auditory response models at several or all tonal frequencies.
[0036] In some preferred embodiments, the processed auditory output is selected from:
[0037] (i) a modified auditory stimulus designed to compensate for a hearing impairment or to generate enhanced hearing; or,
[0038] (ii) a modified auditory response corresponding to a specific processing stage along the auditory pathway, which can be used, for example, to stimulate an auditory prosthesis such as a cochlear implant or a deep brain implant.
[0039] In some preferred embodiments, the difference between the auditory nerve outputs of the normally hearing and hearing-impaired peripheries is minimized; or the difference between the simulated auditory brainstem and / or cortical responses expressed in the time domain or frequency domain is minimized.
[0040] In some preferred embodiments, a task-optimized speech "backend" simulating the performance of the listener in different tasks is connected to the output of the auditory response model, which is also referred to as the "frontend"; and the output of the backend is used to determine and minimize the auditory response difference.
[0041] In some preferred embodiments, the method is used to configure an auditory device, where the auditory device is a cochlear implant or a wearable hearing aid.
[0042] The invention also relates to the use of the method or its embodiments as described herein in hearing aid applications.
[0043] The invention also relates to a processing device, such as a processing unit of an auditory device, which is configured to perform the method and / or any of its embodiments as described herein. Preferably, the processing unit is configured to:
[0044] a. Generate a neural network-based personalized auditory response model at least based on the integrity of the subject's auditory nerve fibers (ANF) and / or synapses (ANS), and preferably also based on the integrity of the subject's inner hair cell (IHC) damage and / or outer hair cell (OHC) damage; the personalized auditory response model represents the expected auditory response of the subject with an auditory distribution to an auditory stimulus;
[0045] b. Compare the output of the personalized auditory response model with the output of a neural network-based desired auditory response model to determine an auditory response difference; wherein the neural network-based model consists of non-linear operations that make the auditory response difference differentiable;
[0046] c. Use the determined differentiable auditory response difference to develop a neural network-based individualized auditory signal processing model for the subject, wherein the individualized auditory signal processing model is configured to minimize the determined auditory response difference; and
[0047] d. Apply the individualized neural network-based auditory signal processing model to the auditory stimulus to generate a processed auditory output that, when given as an input to the personalized auditory response model or the subject, matches the desired auditory response.
[0048] The invention also relates to an auditory device, preferably a cochlear implant or a wearable hearing aid, which includes a processing device configured to perform the method and / or any of its embodiments as described herein.
[0049] In some preferred embodiments, the auditory device includes:
[0050] - An input device, which is configured to pick up input sound waves from the environment and convert the input sound waves into an auditory stimulus;
[0051] - A processing unit, which is configured to perform the method and / or any of its embodiments as described herein; and,
[0052] - An output device, which is configured to generate a processed auditory output from the processor.
[0053] In some preferred embodiments, the auditory device includes:
[0054] - An input device disposed on the auditory device, the input device being configured to pick up input sound waves from the environment and convert the input sound waves into auditory stimuli;
[0055] - A processing unit configured to perform the methods and / or any of its embodiments as described herein; and,
[0056] - An output device disposed on the auditory device, the output device being configured to generate a processed auditory output from the processor.
[0057] The present invention also relates to a computer program configured to perform the methods or its embodiments as described herein, or a computer program product directly loadable into the internal memory of a computer, or a computer program product stored on a computer-readable medium, or a combination of such computer programs or computer program products. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The following description of the figures of the present invention is given by way of example only and is not intended to limit this description, its application or use. In the figures, the same reference numerals denote the same or similar parts and features.
[0059] Figure 1 A flowchart presenting preferred steps for determining the distribution of auditory nerve fibers and synapses and optionally using reference data to determine a subject-specific auditory distribution. Such a distribution can be used in the methods according to some embodiments of the present invention.
[0060] Figure 2 A flowchart presenting preferred steps for determining the ANS / ANF and OHC distributions and optionally using reference data to determine a subject-specific auditory distribution. Such a distribution can be used in the methods according to some embodiments of the present invention.
[0061] Figure 3 A flowchart presenting preferred steps for determining a desired auditory response. The determined auditory response can be used to configure an auditory device, such as a cochlear implant or a hearing aid. Such an auditory response can be used in the methods according to some embodiments of the present invention.
[0062] Figure 4 A method is shown for extracting, approximating, training, and evaluating the outputs of different stages of an auditory periphery model that can be used in the methods according to some embodiments of the present invention.
[0063] Figure 5Shows a closed-loop method for designing a compensation strategy for hearing impairment according to some embodiments of the present invention. In this example, the simulation results from a normal-hearing model and a hearing-impaired model are compared to stimulate a signal processing algorithm that makes the hearing-impaired response closer to the normal-hearing response.
[0064] Figure 6 Shows a closed-loop method for designing a simulator for hearing impairment according to an embodiment of the present invention. In this example, the simulation results from a normal-hearing model and a hearing-impaired model are compared to stimulate a signal processing algorithm that provides signals that can emulate the hearing perception of a listener with such a periphery.
[0065] Figure 7 Shows using a personalized auditory response model and a reference auditory response model to generate a difference signal based on the difference in their outputs. The auditory response model can be a model of the auditory periphery or an ASR system or any NN-based auditory model. The individualized auditory model can be adapted to an individual subject using different sensors and measurement data, including experimental data of OAEs, AEPs, or performance in psychoacoustic tasks such as speech reception threshold (SRT). By using an NN-based auditory model, the difference signal can be differentiated and thus backpropagated through these models.
[0066] Figure 8 Shows using the above difference signal as a loss function to train an individualized NN-based auditory signal processing model. During training, the output of the processing model will be given as the input to the personalized auditory response model, and its parameters are adjusted to minimize the difference signal. After successful training, the NN-based auditory processing model can be directly used to process auditory stimuli and generate a processed output that adapts to the individualized response model or a human listener.
[0067] Figure 9 Shows the real-time optimization of a pre-trained individualized auditory signal processing model adapted to a specific subject. In this schematic diagram, the AEP response of the subject to the processed stimulus is collected by a sensor and compared with the simulated AEP response of the reference auditory model output for the unprocessed stimulus. The weights of the processing model are adjusted online such that the measured AEP response is optimized to better match the reference AEP response.
[0068] Figure 10 Shows the use of an NN-based ASR model for an auditory response model. The individualized ASR model can be a hearing-impaired ASR model or a combination of a simple ASR backend and a hearing-impaired front end.
[0069] Figure 11Shows an implementation of a preferred neural network-based model called "CoNNear", which is a fully convolutional encoder-decoder neural network with strided convolutions and skip connections to map an audio input to 201 basilar membrane vibration outputs in different cochlear sections (N CF ). The CoNNear architectures with (a) context and (b) without context are shown. The final CoNNear model has four encoder and decoder layers, uses context, and includes a tanh activation function between the CNN layers. (c) Provides an overview of the model training and evaluation procedures. Although the CoNNear parameters are trained using a reference analysis TL model simulation of a speech corpus, the model is evaluated using simple acoustic stimuli commonly used in cochlear mechanics research.
[0070] Figure 12 Shows training an audio signal processing DNN model using the ConNear output. (a) The audio signal processing DNN model is trained to minimize the difference in the outputs of two CoNNear IHC-ANF models (orange pathways). (b) When processed by the trained DNN model, the input stimulus results in an excitation rate output of the second model that closely matches that of the first model. Detailed Description
[0071] As used herein below, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" include both the singular and the plural.
[0072] The terms "comprise", "comprises" used hereinafter are synonymous with "including", "include" or "contain", "contains", and are inclusive or open-ended and do not exclude additional unmentioned parts, elements or method steps. When this description refers to a product or process that "comprises" a particular feature, component or step, this means that there is a possibility that other features, components or steps may also be present, but it may also refer to an embodiment that only contains the listed features, components or steps.
[0073] Enumerating numerical values by a numerical range includes all values and fractions within these ranges, as well as the recited endpoints.
[0074] The term "about" when referring to a measurable value such as a parameter, quantity, period of time, etc. is intended to include variations of + / - 10% or less, preferably + / - 5% or less, more preferably + / - 1% or less, and still more preferably + / - 0.1% or less, starting from the specified value, provided that these variations are applicable to the present invention disclosed herein. It should be understood that the values themselves referred to by the term "about" are also disclosed.
[0075] All references cited in this specification are hereby incorporated by reference in their entirety.
[0076] The percentages used herein may also be expressed as dimensionless fractions, or vice versa. For example, a value of 50% may also be written as 0.5 or 1 / 2.
[0077] Unless otherwise defined, all terms disclosed in this invention, including technical and scientific terms, shall have the meanings commonly ascribed to them by those of ordinary skill in the art. For further guidance, definitions are included to further explain the terms used in the description of this invention.
[0078] The present invention relates to a method based on an artificial neural network for obtaining an individualized auditory signal processing model suitable for converting an auditory stimulus into a processed auditory output. The method preferably comprises the following steps:
[0079] a. Obtaining, preferably generating, a neural network-based personalized auditory response model that represents the expected auditory response of a subject with an auditory profile to an auditory stimulus;
[0080] b. Comparing the output of the personalized auditory response model with the output of a neural network-based desired auditory response model to determine an auditory response difference; whereby, the auditory response difference is differentiable, i.e., it can be used to train / develop a neural network model that can be backpropagated to a solution; and,
[0081] c. Using the determined differentiable auditory response difference to develop a neural network-based individualized auditory signal processing model for the subject, wherein the individualized auditory signal processing model is configured to minimize the determined auditory response difference.
[0082] The method can thus obtain an individualized auditory signal processing model that can process the auditory stimulus to generate a processed auditory output that, when given as an input to the personalized auditory response model or to the subject, matches the desired auditory response.
[0083] The present invention also relates to a method based on an artificial neural network for converting an auditory stimulus into a processed auditory output. The method preferably comprises the step of obtaining an individualized auditory signal processing model or an implementation thereof as described herein; and,
[0084] d. Applying the individualized neural network-based auditory signal processing model to the auditory stimulus to generate a processed auditory output that, when given as an input to the personalized auditory response model or to the subject, preferably matches the desired auditory response.
[0085] In some embodiments, the method is a computer-implemented method.
[0086] In some preferred embodiments, the subject is a human or animal subject, preferably a human subject. In some embodiments, the human subject has a hearing impairment. In some embodiments, the human subject has a synaptopathy. In some embodiments, the human subject has outer hair cell (OHC) loss. In some embodiments, the human subject has inner hair cell (IHC) damage. In some embodiments, the human subject has demyelination. In some embodiments, the human subject has presbycusis or brainstem / midbrain inhibitory changes. In some embodiments, the human subject has a hearing impairment of the above types at different stages of the auditory periphery. In some embodiments, the human subject particularly has both synaptopathy and outer hair cell (OHC) loss, for example, due to aging or noise exposure.
[0087] The method can be applied to most people of all ages and various sensorineural hearing impairments and can be applied in different situations: watching movies, sleeping, subconsciously, non-verbal (e.g., newborns). In addition, people undergoing cancer treatment can be considered.
[0088] The method according to the invention preferably comprises the following steps:
[0089] a. Obtaining, preferably generating, a personalized auditory response model based on a neural network, the personalized auditory response model representing the expected auditory response of a subject with an auditory profile to an auditory stimulus.
[0090] The personalized auditory response model can be pre-determined or determined using the measured responses of the subject to sensitive stimuli (e.g., AEP, OAE) or using the performance results of psychoacoustic tasks such as speech intelligibility or amplitude modulation (AM) detection tasks. As used herein, the term "auditory evoked potential" (AEP) refers to a type of EEG signal emitted from the brain scalp by acoustic stimulation. As used herein, the term "otoacoustic emission" (OAE) refers to a sound generated within the inner ear, which is typically recorded using a sensitive microphone and is commonly used as a measure of inner ear health.
[0091] As used herein, an artificial neural network (ANN or NN) is preferably a deep neural network (DNN), and the deep neural network preferably has at least 2 layers between the input layer and the output layer. The neural network can be a convolutional neural network (CNN).
[0092] The neural network-based models in this disclosure can consist of non-linear operations that make the auditory response differentiable. As understood in the art regarding neural networks, the term "differentiable" refers to a mathematical model that has a computable gradient and is capable of reiterating at least one component through gradient optimization using a mathematical optimization algorithm. Thus, providing a neural network-based model that is differentiable enables the use of gradient-based optimization of the parameters in the model, such as gradient descent, to accurately solve problems. Therefore, differentiability is an inherent property of this neural network-based model, which enables the training of the model to backpropagate to solutions that would otherwise not be achievable through, for example, gradient-free optimization, without resorting to mathematical simplifications that sacrifice model accuracy to solve the problem. Those skilled in the art know which mathematical expressions are differentiable, and since most neural networks only consist of differentiable components, those skilled in the art have no difficulty in selecting a differentiable NN-based model.
[0093] In some embodiments, the NN-based model typically consists of highly non-linear but parallel operations. This has the advantage of further significantly accelerating the computation when implemented on a dedicated chip, compared to the computation of complex mathematical feed-forward expressions. At the same time, these operations are differentiable, which means that the neural network can be trained to backpropagate to solutions that would otherwise not be achievable. Therefore, this method is preferably used in the closed-loop compensation method.
[0094] Using the NN-based auditory model, the above-mentioned difference signal is differentiable and reflects a specific degraded hearing ability.
[0095] An additional benefit of connecting the field of individualized neural network (NN)-based auditory signal processing models and NN-based audio signal processing is that this combination can improve the performance of state-of-the-art speech recognition, noise suppression, sound quality, and robotic systems to operate under more adverse conditions such as negative signal-to-noise ratio (SNR). The NN-based auditory signal processing model, classifier, or recognition system can help utilize the extraordinary frequency selectivity and noise reduction ability of the human cochlea, which helps with the perception of speech in noise at negative SNR (< -6dB), while spectro-temporal traditional audio signal processing applications start to fail when the SNR is below 0dB.
[0096] In the context of the present invention, an auditory stimulus can be of various types and refers to an acoustic signal (e.g., a pressure wave) to which the human or animal hearing is sensitive, such as a signal for the human auditory system that includes and transmits acoustic energy in the range from approximately 20 Hz to approximately 20 kHz, depending on age and health status. Clearly, for non-human animals, different frequency ranges apply. As used herein, the term "auditory processing" refers to the processing of sound by the auditory periphery and includes cochlear and neural processing of sound across the various stages in the ascending auditory pathway. Thus, as used herein, the term "auditory processing" can refer to the processing of the auditory periphery or pathway, which includes cochlear processing as well as brainstem and midbrain neuronal processing and the processing of neuronal populations at any of the foregoing steps. Thus, the term "cochlear processing" refers to the processing that occurs in the middle ear, on the basilar membrane (BM), in outer and inner hair cells (OHC and IHC), at the synapses of auditory nerve fibers (ANF), and within neurons.
[0097] As used herein, the term "individualized auditory response model" is preferably defined as an NN-based model of the biophysical sound processing stages along the auditory pathway. The NN-based model can include stages corresponding to the ear canal, middle ear, cochlear basilar membrane filtering, and responses from cochlear nerve elements such as inner and outer hair cells (IHC and OHC), auditory nerve fibers (ANF), brainstem / midbrain neurons, and their synapses. In addition, population responses from several of these elements can form the outcome of the individualized model: for example, otoacoustic emissions (OAE), which are population basilar membrane and OHC responses; and auditory evoked potentials (AEP), which are neuronal population responses generated at the ANF and / or brainstem / midbrain neuron levels. The individualized auditory response model can individualize one or more frequency-related parameters associated with hearing impairments of the above-described structures. The model can be a single NN model that encompasses all aspects of hearing impairment and auditory processing, or it can consist of modules, each of which covers a specific aspect of auditory processing and / or hearing impairment.
[0098] As used herein, the term "individualized auditory signal processing model" is preferably defined as an NN-based auditory signal processing algorithm that has an auditory stimulus as input and has, as the processed auditory output, for example, (i) a modified auditory stimulus designed to compensate for a hearing impairment or produce enhanced hearing or (ii) a modified auditory response corresponding to a specific processing stage along the auditory pathway that can be used, for example, to stimulate an auditory prosthesis such as a cochlear implant or a deep brain implant.
[0099] Thus, in some preferred embodiments, the processed auditory output is selected from:
[0100] (i) a modified auditory stimulus designed to compensate for a hearing impairment or produce enhanced hearing; or,
[0101] (ii) A modified auditory response corresponding to a specific processing stage along the auditory pathway that can be used, for example, to stimulate an auditory prosthesis such as a cochlear implant or a deep brain implant.
[0102] As used herein, the terms "enhanced hearing" and "enhanced auditory response" preferably relate to the purpose of an individualized auditory signal processing algorithm. In addition to compensating for individual forms of hearing impairment, the algorithm can also be designed to improve hearing (even for normal-hearing listeners), with the aim of improving the perception or quality of hearing or enhancing the auditory response (e.g., AEP, OAE). This can be achieved by aiming to perform noise reduction or by enhancing certain neural response characteristics through means such as audio signal onset or modulation enhancement.
[0103] In some preferred embodiments, the individualized auditory response model of step a. is determined by deriving and including a subject-specific auditory profile.
[0104] This step is preferably pre-executed by using a sensitivity metric to measure the biological response of the subject to a specific sound stimulus (e.g., OAE, AEP) or by using an additional sensor that detects human biological signals. These data are compared with model simulations to determine the best-matching auditory profile.
[0105] In some preferred embodiments, the subject-specific auditory profile is a subject-specific auditory impairment profile; preferably based on the integrity of the auditory nerve fibers (ANF) and / or synapses (ANS), and / or based on the damage to the outer hair cells (OHC) of the subject.
[0106] As is known to those skilled in the art, hearing loss can be attributed to several measurable factors at different stages of the auditory periphery, including but not limited to:
[0107] - Damage / loss of outer hair cells (OHC);
[0108] - Dysfunction or loss of the auditory nerve (AN);
[0109] - Damage / loss of inner hair cells (IHC);
[0110] - Demyelination;
[0111] - Presbycusis; and,
[0112] - Alternation of neural inhibition intensity.
[0113] Once the exact auditory profile (auditory impairment profile) of the hearing loss has been estimated for an individual, an individualized signal processing auditory response model can be developed. For example, the individualized signal processing auditory response model can accurately compensate for a specific hearing impairment. In some embodiments, the method includes the step of developing an individualized hearing aid signal processing model as described herein. The auditory impairment profile can include outer hair cell damage, inner hair cell damage, cochlear synaptopathy, brainstem inhibition changes, or even a combination of hearing losses at all different stages in the auditory periphery, such as those described above. Using sensitive metrics based on otoacoustic emissions (OAEs) and auditory evoked potentials (AEPs), an individualized model can be constructed that accounts for an individual's synaptopathy and hair cell damage.
[0114] In some embodiments, using sensitive metrics based on otoacoustic emissions (OAEs) and auditory evoked potentials (AEPs), a personalized auditory response model has been established that can account for synaptopathy and outer hair cell damage. Thus, preferably, the personalized auditory response model includes both synaptopathy and outer hair cell damage.
[0115] In some embodiments, the subject-specific auditory impairment profile includes an auditory nerve fiber and / or synapse damage profile; that is, the auditory impairment profile is based on the integrity of the auditory nerve fibers (ANFs) and / or synapses (ANSs).
[0116] In some embodiments, the subject-specific auditory impairment profile includes an outer hair cell damage profile; that is, the auditory impairment profile is based on the integrity of the outer hair cells (OHCs).
[0117] In some embodiments, the subject-specific auditory impairment profile includes an inner hair cell damage profile; that is, the auditory impairment profile is based on the integrity of the inner hair cells (IHCs).
[0118] In some embodiments, the subject-specific auditory impairment profile includes a demyelination damage profile.
[0119] In some embodiments, the subject-specific auditory impairment profile includes a presbycusis damage profile.
[0120] In some embodiments, the subject-specific auditory impairment profile includes a brainstem / midbrain inhibition change profile.
[0121] In some embodiments, the subject-specific auditory impairment profile includes an auditory nerve fiber and / or synapse damage and outer hair cell damage profile; that is, the auditory impairment profile is based on the integrity of the auditory nerve fibers (ANFs) and / or synapses (ANSs), and on the integrity of the subject's outer hair cell (OHC) damage.
[0122] In some embodiments, the subject-specific auditory damage profile includes brainstem / midbrain damage, auditory nerve fiber and / or synapse damage, and hair cell damage profiles; that is, the auditory damage profile is based on the integrity of the brainstem / midbrain, the integrity of the auditory nerve fibers (ANF) and / or synapses (ANS), and the hair cell damage of the subject.
[0123] The developed neural network model of the auditory periphery can also assist in this step by providing a faster way to cluster experimental data into simulated outputs such that an individualized hearing loss profile can be constructed with better accuracy. To this end, a pre-configured personalized auditory response model of hearing impairment can be used, which includes various aspects of hearing loss at various degrees.
[0124] Thus, the term "integrity" can relate to one or both of the function or loss of elements in the auditory periphery, such as inner hair cell loss, outer hair cell loss, or other types of hearing damage as described herein. For example, ANF integrity can refer to one or both of the function of the remaining ANF, and the innervation of the afferent cochlear synapses (ANS) to them. The term "integrity" can also relate to the quantification of the number and / or type of damaged elements in the auditory periphery, such as inner hair cell loss, outer hair cell loss, or damaged ANF and / or ANS. As used herein, the terms "measuring integrity" or "determining integrity" can interchangeably designate qualitative or quantitative measurements. By incorporating at least ANF and / or ANS integrity into a network-based personalized auditory response model, a biophysically accurate model can be generated that can be personalized to fit subgroups of individuals and / or fit a single individual.
[0125] Auditory damage can be evaluated by any means known to those skilled in the art. For example, it has been found that ANF exhibits a strong response to specific auditory stimuli (audio stimuli or stimuli), that is, the auditory stimuli are capable of inducing highly synchronous ANF responses in populations of ANF and ANS along the cochlea. The ANF response can be recorded by measuring the electrical activity of the brain. This activity is mapped by invasive recording electrodes (in animals) or by electroencephalogram (EEG, in humans or animals), preferably AEP. For EEG, a number of electrodes are attached to the scalp of the subject, and these electrodes record all brain activity as a fluctuating graph. The EEG data can be processed to determine the integrity of ANF and / or ANS in the subject. The integrity can be determined for the entire ANF population or a subset of the ANF population.
[0126] Other functional neuroimaging techniques can be used in the present invention. For example, the brain activity of a subject can also be mapped by magnetoencephalography (MEG) or electrocochleography (EcochG). Those skilled in the art understand that EcochG / MEG data can be processed in a manner equivalent to the embodiments described for EEG data, and the application of this auditory stimulation is not limited to any particular neuroimaging technique. Data from different neuroimaging and / or auditory tests can also be combined to obtain more accurate or alternative results, such as determining damage to other auditory components, such as outer hair cell (OHC) damage. In some embodiments, the damage distribution specific to a subject can be extended to also include, for example, a simulated and / or experimentally frequency-specific OHC damage distribution. The OHC damage component can be determined based on experimental data, i.e., an estimate of frequency-specific OHC damage (e.g., derived from audiogram tests, otoacoustic emissions). Alternatively, the OHC damage distribution can be kept variable so that the matching algorithm can be optimized for both ANF and OHC distributions simultaneously.
[0127] In some embodiments, the auditory damage distribution is obtained from brain activity data, such as via AEP. In some embodiments, the brain activity data is obtained from a signal, preferably, the signal is an EEG (electroencephalogram) or MEG (magnetoencephalogram) signal, preferably an EEG signal, preferably an AEP signal. These EEG and MEG methods can provide a non-invasive approach with high temporal precision for hearing screening. As used herein, the term "EEG" also includes EcochG (electrocochleogram) because this setup is essentially an EEG recording from a tiptrode in the ear canal or a transtympanic electrode through the tympanic membrane (requiring a clinical setup).
[0128] The method according to the present invention preferably comprises the following steps:
[0129] b. Comparing the output of the personalized auditory response model with the output of the desired auditory response model based on a neural network to determine an auditory response difference; whereby the auditory response difference is differentiable, i.e., it can be used to train / develop a neural network model that can be backpropagated to the solution.
[0130] In some embodiments, the desired auditory response is automatically determined based on an auditory response model of a subject without hearing loss. In some embodiments, the desired auditory response is determined based on sensor input or data obtained from the subject. In some embodiments, the desired auditory response is experimental or simulated.
[0131] In some embodiments, the desired auditory response is an enhanced response. In some preferred embodiments, the desired auditory response is a response from a subject with normal hearing or a response with enhanced characteristics.
[0132] A normal-hearing auditory periphery can simulate the hearing perception / ability of a normal-hearing listener. Examples of enhanced features include, but are not limited to, improved sound perception or sound quality, combined noise reduction, or other manipulations.
[0133] In some embodiments, the desired auditory response is from a hearing-impaired subject. This can provide a processed audio stimulus that, when played back to a normal-hearing listener, will simulate the hearing degradation experienced by a hearing-impaired listener.
[0134] In some embodiments, the desired auditory response model and the personalized auditory response model include task-oriented neural network auditory models, such as automatic speech recognition (ASR) / word recognition systems, speech enhancement models (noise suppression, dereverberation), or audio / speech quality models.
[0135] In some embodiments, the desired auditory response model and the personalized auditory response model include psychoacoustic neural network models, such as loudness models.
[0136] In some embodiments, the desired auditory response model and the personalized auditory response model include different combinations of neural network models, such as an auditory model (front end) and an ASR system (back end); or combinations of more models, such as a noise suppression model as an intermediate step between the front end and the back end.
[0137] In some preferred embodiments, the desired auditory response model and the personalized auditory response model include models of different stages of the auditory periphery, as described herein.
[0138] The method according to the present invention preferably comprises the following steps:
[0139] c. Using the determined differentiable auditory response difference to develop a neural network-based individualized auditory signal processing model for the subject, wherein the individualized auditory signal processing model is configured to minimize the determined auditory response difference.
[0140] The neural network-based individualized auditory signal processing model can be used for various applications, depending on the selected personalized auditory response model and the desired auditory response model. Examples of such specific applications are shown below.
[0141] In some preferred embodiments, a reference neural network describing a normal-hearing auditory periphery is used as the desired auditory response model; a corresponding hearing-impaired neural network is used as the personalized auditory response model; and the individualized auditory signal processing model is a signal processing neural network model that is trained to process auditory inputs and compensate for the degraded output of the hearing-impaired model when connected to the input of the hearing-impaired model or the subject.
[0142] In some preferred embodiments, a reference hearing-impaired neural network is used as the desired auditory response model; a corresponding neural network describing the normal-hearing auditory periphery is used as the personalized auditory response model; and the individualized auditory signal processing model is the following signal processing neural network model, which is trained to process auditory inputs and simulate the degraded output of the hearing-impaired model when connected to the input of the normal-hearing model.
[0143] In some preferred embodiments, a reference neural network simulating the enhanced hearing perception and / or ability of a normal-hearing listener is used as the desired auditory response model; a corresponding normal-hearing neural network or hearing-impaired neural network is used as the personalized auditory response model; and the individualized auditory signal processing model is a signal processing neural network model trained to process auditory inputs and provide an enhanced auditory response.
[0144] In some embodiments, the method includes calibrating an individual hearing impairment model of a subject through OAE / AEP experiments. The experimentally recorded OAEs and audiometric thresholds can be used to determine the personalized OHC distribution. AEPs can be simulated for a range of synaptopathy distributions, i.e., for different degrees of ANF damage. Depending on the type of AEP, such as auditory brainstem response (ABR) or envelope following response (EFR), a feature set containing temporal peaks and latencies, spectral amplitudes, and related metrics can be constructed for each simulated cochlear synaptopathy distribution. Using clustering techniques, the CS distribution that best matches the feature set extracted from the measurements can be determined, and the corresponding OHC and ANF damage parameters can be used to set the parameters of the NN-based individual auditory response model.
[0145] The above procedure can be further optimized by involving both OHC loss and synaptopathy parameters to determine the best-matching distribution. The procedure includes more degrees of freedom, and instead of pre-determining the OHC parameters before iteratively determining the ANF distribution, all OHC- and ANF-related model parameters can now be iteratively run to minimize the difference between the experimental and simulated feature sets. In this way, the OHC and ANF damage parameters of the NN-based auditory response model can be optimized simultaneously.
[0146] In some embodiments, the subject auditory response model can be individualized based on recorded biophysical data (e.g., individual parameters of ANS, ANF, OHC, and / or IHC damage) from the subject to simulate the auditory periphery of an individual listener. Those skilled in the art will thus understand that the individualized model as used herein is different from the personalized model. A personalized model will fit a subgroup of individuals, while an individualized model is for a single individual.
[0147] In particular, an individualized auditory response model refers to an NN-based model, such as obtained from a single measurement (e.g., an audiogram that determines OHC damage) and / or by aggregating data into a single model (based on combined hearing damage, such as OHC and / or IHC damage); while an individualized auditory response model refers to the individualization of all included NN-based models (e.g., the individual contributions of ANS, ANF, OHC, and / or IHC).
[0148] The above individualized subject auditory response model can provide the ability to design an individualized hearing aid model using a closed-loop system, which optimally compensates for the specific sensorineural hearing loss of an individual listener, regardless of the perceptual constraints (e.g., the perceived loudness of the gain solution method) currently used in state-of-the-art hearing aid algorithms.
[0149] After determining the individual auditory profile of the listener, the corresponding parameters can be used to train a personalized NN-based auditory response model, which can capture the hearing damage at each different stage of the listener's periphery up to the level of the auditory nerve or brainstem / midbrain processing. Then, the individual auditory model is used in a closed-loop pathway and its output is compared with the output of a "reference" normal-hearing auditory model.
[0150] The neural network-based models in this disclosure can consist of non-linear operations that make the auditory response differentiable. In some embodiments, the NN-based model can consist of highly non-linear but parallel operations. Since their operations are differentiable, this can enable the use of gradient-based parameter optimization such as gradient descent in the model to accurately solve problems. Thus, differentiability is an inherent property of this neural network-based model, enabling it to be trained to backpropagate to solutions that would otherwise be unattainable. For example, a non-differentiable auditory model might have to resort to mathematical simplifications to reach a solution through, for example, gradient-free optimization, thereby reducing the accuracy of the solution.
[0151] Therefore, by providing a neural network-based model consisting of non-linear operations that make the auditory response differentiable, the above two auditory models can be used to design a closed-loop compensation pathway, where the "hearing aid" neural network model is trained to process auditory inputs and compensate for the degraded output of the individual hearing-impaired model (as Figure 5 shown).
[0152] The closed-loop method is made possible due to the differentiable nature of the auditory models used. The outputs of these two models can provide a difference metric that can be used as a penalty / loss term to train the hearing aid model. This metric is used for backpropagation through the NN-based auditory model and accordingly modify the weights of the hearing aid model such that it can be trained to minimize a particular metric in the best possible way. The hearing aid model is trained to process auditory stimuli such that when the auditory stimulus is given as input to the hearing-impaired model, it can produce an output that can match (or partially match) the output of a "reference" normal-hearing model.
[0153] The present invention also relates to an artificial neural network-based method for converting an auditory stimulus into a processed auditory output. The method preferably includes the steps of obtaining an individualized auditory signal processing model or an implementation thereof as described herein; and,
[0154] d. applying the individualized neural network-based auditory signal processing model to the auditory stimulus to produce a processed auditory output that preferably matches the desired auditory response when given as input to the individualized auditory response model or the subject.
[0155] In some preferred embodiments, the individualized auditory signal processing model is trained to minimize a particular auditory response difference metric, such as the absolute difference or squared difference between two auditory response models at several or all tonal frequencies.
[0156] In some embodiments, the absolute difference between two models is used to minimize the difference between the desired auditory response and the auditory response. In some embodiments, the squared difference between two models is used to minimize the difference between the desired auditory response and the auditory response.
[0157] In some embodiments, the responses of two models expressed in the frequency domain are used to minimize the difference between the desired auditory response and the auditory response. In some embodiments, the responses of two models expressed in different frequency representations such as power or amplitude spectrograms are used to minimize the difference between the desired auditory response and the auditory response.
[0158] In some preferred embodiments, the difference in the total auditory response across a series of simulated frequencies is minimized. This allows for the optimal recovery of the generators of auditory evoked potentials when used as input to the brainstem and cortical processing models.
[0159] In some preferred embodiments, the difference in the auditory nerve outputs of the normal-hearing and hearing-impaired peripheries is minimized; or the difference between simulated auditory brainstem and / or cortical responses expressed in the time domain or frequency domain is minimized.
[0160] The choice of optimization metric has an impact on closed-loop compensation. Given the complexity of these representations, minimizing the difference between the outputs of a normal-hearing model and a hearing-impaired model, as used in some embodiments, may not always be desirable or even possible. In some embodiments, a personalized auditory signal processing model (a hearing aid model in this example) can be selected to compensate for a single aspect of hearing impairment (e.g., outer hair cell damage or synaptopathy) at several or all tonal frequencies. In some other embodiments, the simulated cochlear response is used as an input to a brainstem and cortical processing model such that additional auditory evoked potential features can be simulated and used to determine the parameters of the hearing aid model. In some other embodiments, the hearing aid model can be trained to optimally restore the generators of auditory evoked potentials, in which case the total cochlear response across a range of simulated frequencies is used as an input to the brainstem and cortical processing model to determine the parameters of the hearing aid model.
[0161] In some other embodiments, the hearing aid model is trained to process auditory signals such that for a perception task such as speech intelligibility, the “reference” performance of a normal-hearing subject can be achieved. In this case, a task-optimized speech “backend” is connected to the outputs of normal-hearing and hearing-impaired cochlear models (i.e., the “frontends”), which will simulate the performance of the listener in different tasks. The output of the backend can then be used to train the hearing aid model, which minimizes the difference between the hearing-impaired and normal-hearing performances. The frontend can be a cochlear model or a cochlear model connected to an auditory brainstem / cortical processing. The task-optimized backend can be a neural network (NN)-based automatic speech recognition (ASR) system. In some embodiments, as a next step, noise or reverberation is introduced into the auditory signals to generalize the performance of these models in more realistic scenarios. In this case, an NN-based noise / reverberation suppression model can also be added as an intermediate step between the frontend and the backend.
[0162] In some preferred embodiments, a task-optimized speech “backend” that simulates the performance of the listener in different tasks is connected to the output of an auditory response model, also referred to as the “frontend”; and the output of the backend is used to determine and minimize the auditory response difference.
[0163] In some embodiments, the auditory response model is trained to process auditory signals such that for a perception task such as speech intelligibility, the “reference” performance of a normal-hearing subject can be achieved.
[0164] In some embodiments, a task-optimized speech “backend” is connected to the outputs of a desired auditory response and a simulated auditory response “frontend”, which simulates the performance of the listener in different tasks.
[0165] In some embodiments, the output of the backend is used to minimize the difference between the desired auditory response and the simulated auditory response.
[0166] In some embodiments, the front end is a cochlear model or a model of the entire auditory periphery.
[0167] In some embodiments, the task-optimized backend is a neural network (NN)-based automatic speech recognition (ASR) system.
[0168] In some embodiments, as a next step, noise or reverberation is introduced into the auditory signal to generalize the performance of these models in more realistic scenarios. In some embodiments, an NN-based noise / reverberation suppression model is added as an intermediate step between the front end and the backend.
[0169] In some embodiments, step d. includes the following steps:
[0170] - Suppressing the auditory stimulus when the amplitude of the input sound wave exceeds the generated maximum threshold.
[0171] In some embodiments, step d. includes the following steps:
[0172] - Enhancing the auditory stimulus when the amplitude of the input sound wave is before the generated minimum threshold.
[0173] Once the neural network for individualized signal processing (e.g., hearing aid) is trained by the closed-loop method, it can be used alone to process auditory signals and compensate for specific hearing losses. The neural network can be implemented on a dedicated chip for parallel computing, which is integrated in the hearing aid or may be on a portable low-resource platform (e.g., Raspberry Pi). The signal processing model will preferably run in real time, thus receiving the input through a sensor (e.g., microphone) and providing the processed output to an output device (e.g., headphones, in-ear inserts) with a specific delay.
[0174] The neural network for individualized signal processing (e.g., hearing aid) is preferably trained to adjust the auditory signal in an optimal way, which can vary according to the task, auditory distribution, or application. Preferably, an autoencoder architecture based on convolutional filters is used, so the auditory signal is processed in the time domain, thus providing a processed output with the same representation.
[0175] Preferably, the neural network architecture for individualized signal processing (e.g., hearing aid) includes a mirrored version of the encoder as the decoder. As described above, this architecture will provide an output representation that is the same as the input representation. However, different architectures can be used instead of the autoencoder to provide the input for the hearing-impaired model.
[0176] In some embodiments, step d. includes the following steps:
[0177] - Including additional signal processing algorithms to adjust the audio stimulus.
[0178] In some embodiments, additional signal processing algorithms include filtering, onset sharpening, compression, noise reduction, and / or expanding the audio stimulus.
[0179] In some embodiments, additional signal processing models may include a noise / reverberation suppression stage, a word recognition stage, a frequency analysis or synthesis stage to generalize different acoustic scenarios and tasks.
[0180] In some embodiments, the individualized signal processing model provides an output representation different from the input representation, such as a cochleogram, a neural map, or different auditory feature maps, depending on the desired input of the auditory response model.
[0181] In some embodiments, the individualized signal processing model provides an output representation that simulates the listener's performance in different tasks, such as speech intelligibility / recognition prediction or speech quality assessment, depending on the desired input of the auditory response model.
[0182] In some preferred embodiments, the (individualized and / or simulated) auditory response to sound (e.g., auditory EEG response such as AEP, sound perception, cochlea, ANF, and brainstem processing) is used to adjust specific aspects of the sound stimulus in the time domain or frequency domain, preferably adjusting the intensity and / or the shape of the time envelope (e.g., onset sharpening / envelope depth enhancement). The desired auditory response to sound can be simulated or recorded (e.g., normal hearing or enhanced auditory feature response). The difference between the desired auditory response and the auditory response corresponding to the subject's AN fiber and synaptic integrity and / or OHC damage distribution can then form a feedback loop to the processing unit of the auditory device. For example, the feedback loop can be used to optimize the signal processing algorithm to adjust the sound excitation in these devices.
[0183] After developing an individualized hearing aid NN model for a specific listener, the output of the model to a specific stimulus can be simulated, and these processed stimuli can be used to measure the listener's auditory response (e.g., EGG response such as AEP). By comparing the measured response to the processed stimulus with the measured response to the original stimulus, the improvement of the signal processing algorithm can be evaluated. If necessary, the difference between the measured responses can be used to further optimize the signal processing algorithm.
[0184] In some embodiments, the efficiency of the trained individualized signal processing model can be evaluated on an individual, e.g., by AEP measurement, psychoacoustic tasks (e.g., speech intelligibility, AM detection), or listening tests (e.g., MUSHRA). Compared with the results of the unprocessed stimulus, the results of these tasks can demonstrate the improvement of the processed stimulus and can also be used to further optimize the signal processing model.
[0185] In some preferred embodiments, the method is used to configure an auditory device, where the auditory device is a cochlear implant or a wearable hearing aid.
[0186] The present invention also relates to the use of the method or its embodiments as described herein in hearing aid applications. Examples thereof are described herein.
[0187] In some embodiments, the method is used in a reversible cochlear filter bank. The reversible cochlear filter bank allows for the analysis of a single input sequence into N output sequences, and then the resynthesis of these output sequences (by summing or combining in a more refined manner) to recreate the single input sequence again. Such a filter bank also provides the ability to process the N output sequences in a more detailed, frequency-dependent manner in order to receive the processed input sequence. This is useful for hearing aid applications, such as for compensating for outer hair cell and / or auditory nerve damage.
[0188] Thus, in some embodiments, the method comprises the following steps:
[0189] - Analyzing a single input sequence into N output sequences and then resynthesizing these output sequences, for example by summing, to recreate the single input sequence again; and / or,
[0190] - Synthesizing N output sequences of a time-frequency representation such as an auditory feature map, for example by summing, to recreate a single time-domain input sequence.
[0191] The present invention also relates to an auditory device, preferably a cochlear implant or a wearable hearing aid, which is configured to perform the method and its embodiments as described herein.
[0192] The present invention also relates to an auditory device, preferably a cochlear implant or a wearable hearing aid. The auditory device preferably comprises:
[0193] - An input device provided on the auditory device, which is configured to pick up input sound waves from the environment and convert the input sound waves into auditory stimuli;
[0194] - A processing unit, which is configured to perform the method and its embodiments as described herein; and,
[0195] - An output device provided on the auditory device, which is configured to generate a processed auditory output from the processor.
[0196] In some embodiments, the processed auditory output includes sound waves. In some embodiments, the processed auditory output includes electrical signals. In some embodiments, the processed auditory output includes deep brain stimulation.
[0197] In some embodiments, the input device includes a microphone.
[0198] In some embodiments, the processing unit is a processor, and a dedicated processor for parallel computing (e.g., GPU, VPU, AI-accelerator) is the best choice because it can calculate the output of the NN-based model faster than a CPU.
[0199] The processing device can be a specially designed processing unit such as an ASIC, or can be a dedicated energy-saving machine learning hardware module, such as a convolutional accelerator chip, suitable for portable and embedded applications, such as battery-powered applications.
[0200] In some embodiments, the output device includes at least one transducer.
[0201] In some embodiments, the output device is configured to provide an audible time-varying pressure signal, basilar membrane vibration, or corresponding auditory nerve stimulation associated with at least one auditory stimulus. For example, the transducer can be configured to convert the output sequence generated by the neural network into an audible time-varying pressure signal, basilar membrane vibration, or corresponding auditory nerve stimulation associated with at least one auditory stimulus.
[0202] The present invention also relates to a computer program configured to execute the method or its embodiments as described herein, or a computer program product that can be directly loaded into the internal memory of a computer, or a computer program product stored on a computer-readable medium, or a combination of such computer programs or computer program products.
[0203] A preferred neural network-based model, hereinafter referred to as the CoNNear model, is described below and will be used as one or more models in the method described herein. Due to the differentiable nature of neural networks, any NN-based auditory model can be used in this closed-loop schematic, including the developed CoNNear model described below. However, no other NN-based model can describe the characteristics of the auditory periphery in such detail, up to the level of inner hair cells and auditory nerves.
[0204] In some embodiments, the method includes the following steps:
[0205] - Providing a multi-layer convolutional encoder-decoder neural network, the multi-layer convolutional encoder-decoder neural network including
[0206] - Encoders and decoders, where the encoder and decoder together include at least a plurality of consecutive convolutional layers, for example each including at least one convolutional layer, preferably each including at least a plurality of consecutive convolutional layers. The consecutive convolutional layers of the encoder have a stride relative to the input to the neural network, such as a decreasing, constant, and / or increasing stride, preferably a constant and / or increasing stride, to sequentially compress the above input. The consecutive convolutional layers of the decoder have a stride relative to the compressed input from the encoder, such as a decreasing, constant, and / or increasing stride, preferably a constant and / or increasing stride, to sequentially decompress the compressed input. Each convolutional layer in the convolutional layers includes a plurality of convolutional filters for convolving with the input to the convolutional layer to generate a corresponding plurality of activation maps as output.
[0207] - At least one non-linear unit for applying a non-linear transformation to the activation maps generated by at least one convolutional layer of the neural network, where the non-linear transformation mimics the level-dependent cochlear filter tuning associated with cochlear processing, such as cochlear mechanics, basilar membrane vibration, outer hair cell processing, inner hair cell processing, or auditory nerve processing and combinations thereof, such as cochlear mechanics and outer hair cells.
[0208] - One or more skip connections between the encoder and the decoder, preferably a plurality of skip connections, for directly forwarding the input to the convolutional layers of the encoder to at least one convolutional layer of the decoder.
[0209] - An input layer for receiving the input to the neural network, and
[0210] - An output layer for generating N output sequences of cochlear response parameters for each input to the neural network. The N output sequences correspond to N simulated cochlear filters associated with N different center frequencies across the cochlear tonotopic frequency map. The cochlear response parameters of each output sequence indicate cochlear processing, such as cochlear mechanics, such as cochlear basilar membrane vibration and / or inner hair cells and / or outer hair cells and / or auditory nerve responses, such as position-dependent time-varying cochlear basilar membrane vibration and / or inner hair cell receptor potential and / or outer hair cell response and / or auditory nerve fiber firing pattern, such as the position-dependent time-varying vibration of the cochlear basilar membrane.
[0211] - Provide at least one input sequence indicating the predetermined length of a time-sampled auditory stimulus, and apply the at least one input sequence to the input layer of the neural network to obtain N output sequences of cochlear response parameters, and
[0212] - Optionally, sum or combine the obtained N output sequences, preferably sum, to generate a single output sequence of cochlear response parameters.
[0213] In some embodiments, the non - linear unit applies a non - linear transformation as an element - wise non - linear transformation, preferably hyperbolic tangent.
[0214] In some embodiments, the number of convolutional layers of the encoder is equal to the number of convolutional layers of the decoder.
[0215] In some embodiments, the neural network includes skip connections between each convolutional layer of the encoder and a corresponding convolutional layer of the decoder.
[0216] In some embodiments, the neural network includes a skip connection between the first consecutive convolutional layer in the consecutive convolutional layers of the encoder and the last consecutive convolutional layer in the consecutive convolutional layers of the decoder.
[0217] In some embodiments, the stride of the consecutive convolutional layers of the encoder relative to the input to the neural network is equal to the stride of the consecutive convolutional layers of the decoder relative to the compressed input, so as to match each convolutional layer of the encoder with a corresponding convolutional layer of the decoder to transpose the convolutional operation of the convolutional layers of the encoder.
[0218] In some embodiments, the number of samples of at least one input sequence is equal to the number of cochlear response parameters in each output sequence.
[0219] In some embodiments, the neural network includes a plurality of non - linear units for applying a non - linear transformation to the activation maps generated by each convolutional layer of the neural network.
[0220] In some embodiments, at least one input sequence includes a pre - context portion and / or a post - context portion before and / or after a plurality of input samples indicating an auditory stimulus respectively, and wherein the method further includes cropping each output sequence in the generated output sequences so that the number of cochlear response parameters included is equal to the number of input samples indicating the auditory stimulus in the plurality of input samples.
[0221] In some embodiments, the method includes:
[0222] - providing a training data set including a plurality of training input sequences, each of the plurality of training input sequences including a plurality of input samples indicating a time - sampled auditory stimulus,
[0223] - Provide a biophysically accurate verification model for cochlear processing, preferably a cochlear transmission line model, the accuracy degree of which is evaluated relative to cochlear response parameters measured experimentally indicating cochlear processing, such as cochlear mechanics, such as cochlear basilar membrane vibration and / or inner hair cells and / or outer hair cells and / or auditory nerve responses, such as position-dependent time-varying cochlear basilar membrane vibration and / or inner hair cell receptor potential and / or outer hair cell responses and / or auditory nerve fiber discharge patterns, such as position-dependent time-varying basilar membrane vibration according to the cochlear tonotopic frequency map.
[0224] - Generate N training output sequences for each training input sequence, each of the N training output sequences being associated with a different center frequency of the cochlear tonotopic map.
[0225] - Execute a simulation method using the training input sequence to generate a simulation sequence of corresponding cochlear response parameters for a neural network for the same cochlear tonotopic map, and evaluate the deviation between the simulation sequence and the training output sequences arranged as training pairs, the simulation sequence and the training output sequences of each training pair being associated with the same training sequence.
[0226] - Use the error backpropagation method to update the neural network weight parameters, the neural network weight parameters including the weight parameters associated with each convolutional filter.
[0227] - Optionally, retrain the neural network weight parameters for different sets of neural network hyperparameters to further reduce the deviation, the different sets of neural network hyperparameters including one or more of the following: different non-linear transformations applied by at least one non-linear unit, different numbers of convolutional layers in the encoder and / or decoder, different numbers of convolutional filters in any convolutional layer of the neural network, lengths different from a predetermined length of the input sequence, different skip connection configurations, or optionally different sizes of convolutional filters in any convolutional layer of the neural network.
[0228] In some embodiments, the method further comprises the steps of: providing a modified verification model reflecting cochlear processing of a person with a hearing impairment, and retraining the neural network weight parameters for the modified verification model or for a combination of the verification model and the modified verification model.
[0229] In some embodiments, the auditory device comprises:
[0230] - Pressure detection means for detecting a time-varying pressure signal indicative of at least one auditory stimulus; and / or a sensor for detecting a biological signal of a person, such as an EEG sensor, or a pressure sensor such as an ear canal pressure sensor.
[0231] - A sampling device for sampling detected auditory stimuli to obtain an input sequence including a plurality of input samples, and
[0232] - At least one transducer for converting an output sequence generated by a neural network into an audible time-varying pressure signal, a cochlear response; for example, basilar membrane vibration, inner hair cell response, outer hair cell response, auditory nerve response or corresponding auditory nerve response and combinations thereof, such as basilar membrane vibration; or a corresponding auditory nerve stimulus associated with at least one auditory stimulus.
[0233] Embodiment
[0234] Example 1: Method for determining the integrity of AN fibers and synapses in a subject
[0235] Reference Figure 1 Possible models for determining the integrity of the auditory nerve fibers and synapses of a subject according to preferred embodiments of the present invention are discussed, Figure 1 A flowchart presenting preferred steps for determining the ANF integrity distribution and optionally using reference data to determine the subject-specific auditory distribution is presented. This record is compared with a standard data set of "normal" people with normal ANF. By comparing the reference for the subject, a subject-specific auditory distribution can be obtained.
[0236] 100 is an auditory stimulator (such as a sound), which induces an auditory response in a population of AN fibers and synapses along the cochlea. The stimulus can be used for AEP recording to diagnose ANF damage. The stimulus characteristics can be designed for a limited or broad hearing frequency range. In a preferred embodiment, the auditory stimulus can be a carrier signal c(t) (such as broadband noise or a pure tone), which is amplitude-modulated by a periodic modulator having a non-sinusoidal (rectangular) waveform m(t).
[0237] 200 is a biophysical model of the signal processing of the auditory periphery (which preferably includes a numerical description of cochlear mechanics, outer and inner hair cell functions and represents the firing rate of AN synapses and discharges). The model can include data from, for example, simulated and / or experimental frequency and / or type-specific ANF damage distributions 210. The ANF damage distribution 210 can be determined based on experimental data (such as AEP recordings). The ANF data can be subdivided according to a subset of the ANF population; this can include high spontaneous rate fibers (HSR), medium spontaneous rate fibers (HSR) and low spontaneous rate fibers (LSR) and / or these fiber subtypes within a selected hearing frequency range.
[0238] The responses of all or a subset of the ANF population can be simulated to obtain predicted auditory responses (300) to auditory stimuli. The auditory responses can be simulated auditory EEG responses such as AEP, simulated auditory sound perception, and / or simulated cochlear, ANF, and brainstem processing). Calculating the response amplitudes of the EEG responses to current or different stimuli (from the simulation) can allow the creation of various auditory responses corresponding to different ANF distributions or other input parameters. The auditory responses can be further segmented using category-based parameters, based on, for example, age, gender, etc., or other parameters. The calculated auditory responses and the corresponding ANF damage distributions can be stored in or made available through a database.
[0239] The EEG responses of a subject to a current auditory stimulus 100 can be experimentally measured (400) using EEG settings. Processing of the EEG data allows calculation of the specific EEG response amplitudes of the subject to the stimulus.
[0240] Predicted simulation data can be used to interpret the processed EEG response data of the subject to assign the subject to an auditory distribution (500). The assignment can be performed automatically by a matching algorithm. The assigned distribution is preferably based on the best possible match between the simulated and recorded EEG response amplitudes. Based on the assigned auditory distribution, the integrity of the AN fibers and synapses of the subject can be determined. For example, in this figure, the subject is assigned an ANF distribution characterized by a damaged distribution of 54% HSR, 0% MSR, and 0% LSR. Since the best-matched ANF distribution does not return 100% ANF types in all ANF categories, the subject has a certain degree of cochlear synaptopathy.
[0241] Example 2: Method for determining outer hair cell (OHC) damage in a subject
[0242] Regarding the above Example 1 further, the possible methods for determining the integrity of the AN fibers and synapses of a subject can be extended to also determine the outer hair cell (OHC) damage of the subject. The method is described with reference to Figure 2 and Figure 2 presents a flowchart of the preferred steps for determining the individual ANF and OHC damage distributions and optionally using subject data to determine a subject-specific auditory distribution.
[0243] In particular, the biophysical model 200 of the auditory periphery can be extended to also include, for example, a simulated and / or experimental frequency-specific OHC damage distribution 220. The OHC damage distribution 220 can be determined based on experimental data of frequency-specific hearing loss (e.g., from audiogram tests, otoacoustic emissions). Alternatively, the OHC damage distribution 220 can remain variable such that a matching algorithm for finding the best subject match can be optimized for both the AN and OHC distributions. For example, in this figure, an OHC distribution characterized by 50% OHC damage was assigned to the subject based on the subject's experimental AEP recordings and the best match to a specific simulated auditory response to the same stimulus within a database of simulated auditory responses to many auditory distributions (including ANF and OHC damage). The subject in the figure was determined to have a certain degree of OHC-related hearing loss.
[0244] Example 3: Method for modifying the auditory response of a subject to an expected sound
[0245] Further to the above embodiments, according to an embodiment of the present invention, a method for determining the integrity of the ANF / ANS and / or OHC damage of a subject can be used to modify the desired auditory response of the subject to sound. The method is described with reference to Figure 3 and Figure 3 presents a flowchart of the preferred steps for determining a signal processing algorithm for modifying an auditory stimulus that produces a desired auditory response. The determined signal processing algorithm can be used to configure an auditory device, such as a cochlear implant or a hearing aid.
[0246] The captured (personalized) auditory response to sound (e.g., auditory EEG response, such as AEP, sound perception, cochlea, ANF, and brainstem processing) can be used to determine a subject-specific ANF and OHC damage auditory distribution (500). This auditory distribution can be included in the auditory periphery model to simulate the auditory response to any acoustic stimulus (600). The individually simulated auditory responses can be compared with the desired auditory response (700). The desired response can be experimental or simulated and can be, for example, a response from a normally hearing subject or a response with enhanced features. Subsequently, a signal processing algorithm 800 is included to adjust the sound stimulus such that the simulated auditory response matches the desired auditory response. For example, the matching algorithm 800 can ultimately filter, onset sharpen, compress, and / or expand the audio stimulus 100.
[0247] Example 4: Training a hearing aid neural network
[0248] Figure 5An example of an embodiment of the present invention is shown. In this example, using a "reference" neural network that can describe the normal-hearing auditory periphery and a corresponding hearing-impaired neural network, a "hearing aid" neural network model can be trained to process auditory inputs and compensate for the degraded output of the hearing-impaired model.
[0249] The "hearing aid" model for such an individual will produce a signal that can match (or partially match) the output of a specific hearing-impaired cochlea to the output of a "reference" normal-hearing cochlea. In this example, the hearing aid model is trained to minimize a specific metric, such as the absolute difference or squared difference between two other models, or a more complex metric indicating the degraded hearing ability. Once the exact auditory profile of the hearing loss has been estimated for an individual, an individualized hearing aid model can be developed to accurately compensate for the specific hearing impairment.
[0250] In different embodiments, the hearing-impaired neural network can be used as the "reference" model, and its auditory input can instead be processed by a "hearing disorder" neural network that will be trained to "degrade" the output of the normal-hearing model to match the "reference" hearing-impaired model. This will provide a processed audio stimulus that, when played back to a normal-hearing listener, will simulate the hearing degradation experienced by a corresponding hearing-impaired listener with the corresponding periphery, as Figure 6 shown.
[0251] Example 5: Adjusting the output at different stages of an auditory periphery model
[0252] Figure 4 A method for extracting, approximating, training, and evaluating the outputs of different stages of an auditory periphery model according to an embodiment of the present invention is shown. The top dashed box shows all the elements included in the model of the auditory periphery, which includes an analytical description of the middle ear, cochlear BM vibration, inner hair cells, auditory nerve, and cochlear nucleus, inferior colliculus processing. The simulated outputs of the above processing stages (for all simulated CFs or as the sum of multiple CFs) can be used to train different processing stages of the CoNNear model. An example is shown here where the TL model BM vibration output to the speech corpus is used to train the BM vibration CoNNear model. During training, the L1 loss between the simulated CoNNear output and the TL model output is used to determine the CoNNear parameters. After training, basic acoustic stimuli are used to evaluate the performance of the resulting CoNNear model, which were not presented during training and are typically used in auditory neuroscience and hearing research.
[0253] Example 6: Generating a difference signal and training a signal processing model
[0254] Figure 7It shows that a personalized auditory response model and a reference auditory response model are used to generate a difference signal based on the difference in their outputs. The auditory response model can be a model of the auditory periphery or an ASR system or anything. The individualized auditory model can be adapted to an individual subject using different sensors and measurement data, including experimental data of OAEs, AEPs, or performance in psychoacoustic tasks such as speech reception threshold (SRT). By using an NN-based auditory model, the difference signal can be differentiated and thus backpropagated through these models.
[0255] Figure 8 It shows that the above difference signal is used as a loss function to train an individualized NN-based auditory signal processing model. During training, the output of the processing model is given as the input to the individualized response model, and its parameters are adjusted to minimize the difference signal. After successful training, the NN-based auditory processing model can be directly used to process auditory stimuli and produce a processed output that adapts to the individualized response model or a human listener.
[0256] Example 7: Training a signal processing model to match an expected performance
[0257] Figure 9 It shows the real-time optimization of a pre-trained individualized auditory signal processing model adapted to a specific subject. In this schematic diagram, the AEP response of the subject to the processed stimulus is collected by a sensor and compared with the simulated AEP response of the reference auditory model output for the unprocessed stimulus. The weights of the processing model are adjusted online such that the measured AEP response is optimized to better match the reference AEP response.
[0258] Figure 10 It shows the use of an NN-based ASR model for an auditory response model. The individualized ASR model can be an ASR model for the hearing impaired, or a combination of a simple ASR backend and a hearing-impaired frontend. The difference in the predicted outputs is calculated, i.e., the difference in the percentage of correct answers predicted by the two models, and this difference is used to train the individualized auditory signal processing NN model. The successfully trained processing model will process auditory stimuli such that the prediction performance of the individualized ASR model can reach the performance of the reference model. If the individualized ASR system can accurately predict the performance of the listener through a simulated auditory periphery, then this will lead to a similar improvement in the listener's performance in the same task.
[0259] Similarly, if a normal-hearing ASR is used as the individualized model and an ASR with enhanced features is used as the reference model (e.g., a model that can correctly identify sentences at a low SNR), then the processing model will be trained to process stimuli so that an increased / enhanced performance can be achieved for the ASR system.
[0260] Example 8: Exemplary implementation of a preferred neural network-based model
[0261] Reference Figure 11 , the implementation of a preferred neural network-based model is discussed. This model is referred to as the CoNNear model in this article.
[0262] The CoNNear model has an autoencoder CNN architecture and uses several CNN layers and dimensionality changes to transform a 20 kHz sampled acoustic waveform (in [Pa]) into an NCF cochlear BM displacement waveform (in [μm]). The first four layers are encoder layers, and stride convolution is used after each CNN layer to halve the time dimension. The next four are decoder layers that map the condensed representation to the LxNCF output using deconvolution operations. L corresponds to the initial size of the audio input, and NCF corresponds to 201 center frequencies (CFs) between 0.1 and 12 kHz of cochlear filters. The CFs employed are spaced according to the Greenwood place-frequency map of the cochlea and span the frequency range to which human hearing is most sensitive. It is important to maintain the temporal alignment (or phase) of the input throughout the architecture because this information is crucial for speech perception.
[0263] For this purpose, U-shaped skip connections are used. Skip connections were earlier adopted in image-to-image translation and speech enhancement applications; they directly transfer temporal information from the encoder layer to the decoder layer ( Figure 11 (a) in ; the dashed arrow). In addition to maintaining phase information, skip connections can also improve the model's ability to learn how to best combine the non-linearities of several CNN layers to mimic the level-dependent characteristics of human cochlear processing.
[0264] Each CNN layer consists of a set of filter banks and non-linear operations, and the CNN filter weights are trained using TL simulated BM displacements from the NCF cochlear channels. Although the training is performed using a speech corpus presented at 70 dB SPL, the model evaluation is based on the ability to reproduce key cochlear mechanical properties using basic acoustic stimuli (such as clicks, pure tones) not seen during training ( Figure 11 (c) in ).
[0265] During training and evaluation, the audio input is segmented into 2048-sample windows (100 ms), after which the corresponding BM displacements are simulated and concatenated over time. Since CoNNear processes each input independently and resets its adaptive properties at the start of each simulation, this concatenation process may lead to discontinuities near the window boundaries. To address this issue, we also evaluated an architecture in which previous and subsequent input samples can be used as context ( Figure 11 (b) in ). Compared to the context-free architecture ( Figure 11Unlike (a) in [reference], a final cropping layer is added to remove the simulated context and produce a final BM displacement waveform of size L.
[0266] Finally, due to its convolutional architecture, training CoNNear with audio inputs of a fixed duration does not prevent it from processing inputs of other durations after training. This flexibility is significantly better than matrix multiplication-based neural network architectures that can only operate on inputs of a fixed duration.
[0267] Example 9: Training a preferred neural network-based neural network model on normal and pathological models
[0268] Reference Figure 12 discusses examples where a deep neural network (DNN) model is trained to minimize the difference between the outputs of the following two IHC-ANF models: a normal model and a pathological model. Each model includes CoNNearIHC and CoNNearANfH modules, and the firing rates of each model are multiplied by 10 and 8 times respectively to simulate the innervation of a normal-hearing human IHC at 4 kHz and a pathological IHC with 20% fiber afferent block due to cochlear synaptopathy.
[0269] The DNN model is trained based on the responses of these two CoNNear models to modify the stimulus so as to restore the output of the pathological model back to the output of the normal-hearing model. Figure 12 (a) in [reference] shows that the training is done using a small input dataset of 4 kHz tones with different levels and modulation depths, normalized to the amplitude range of the IHC input, and the DNN model is trained to minimize the L1 loss between the temporal representation and the frequency representation of the output.
[0270] After training, the DNN model provides the processed input to the 8-fiber model to generate an output that matches the normal-hearing firing rate as closely as possible. The results of the modulated tone stimuli are shown in Figure 12 (b) in [reference], for which the amplitude of the 8-fiber model response is restored to the amplitude of the normal-hearing IHC-ANF. This example demonstrates the backpropagation ability of our CNN models, and their application scope can be extended to more complex datasets such as speech corpora to derive suitable signal processing strategies for speech processing in hearing-impaired cochleas.
Claims
1. A method based on an artificial neural network, the method being used to convert an auditory stimulus into a processed auditory output, the method comprising the following steps: a. Generating a neural network-based personalized auditory response model at least based on the integrity of the subject's auditory nerve fibers (ANF) and / or auditory nerve synapses (ANS), the integrity of the subject's inner hair cell (IHC) damage and outer hair cell (OHC) damage; the personalized auditory response model represents the expected auditory response of the subject with an auditory distribution to an auditory stimulus; wherein, the personalized auditory response model is determined by obtaining a subject-specific auditory distribution at least based on the integrity of the subject's auditory nerve fibers (ANF) and / or auditory nerve synapses (ANS), the integrity of the inner hair cell (IHC) and outer hair cell (OHC) damage and including the subject-specific auditory distribution; b. Comparing the output of the personalized auditory response model with the output of a neural network-based desired auditory response model to determine an auditory response difference; wherein, the neural network-based personalized auditory response model and the neural network-based desired auditory response model include non-linear operations that make the auditory response difference differentiable, such that the neural network-based personalized auditory response model and the neural network-based desired auditory response model can be repeatedly optimized for at least one component by using an optimization algorithm along a computable gradient; wherein, a reference neural network describing a normal-hearing auditory periphery is used as the desired auditory response model, and a corresponding hearing-impaired neural network is used as the personalized auditory response model; c. Using the determined differentiable auditory response difference to develop a neural network-based individualized auditory signal processing model for the subject, wherein the individualized auditory signal processing model is configured to minimize the determined auditory response difference; wherein, the individualized auditory signal processing model is the following signal processing neural network model: the signal processing neural network model is trained to process an auditory input and compensate for the degraded output of the hearing-impaired model when connected to the hearing-impaired model or the input of the subject; and, d. Applying the neural network-based individualized auditory signal processing model to the auditory stimulus to generate a processed auditory output, the processed auditory output matching the desired auditory response when given as the input to the personalized auditory response model or to the subject, wherein, the outer hair cell (OHC) damage distribution remains variable such that the matching can be optimized simultaneously for the auditory nerve fibers (ANF) and the outer hair cell (OHC) damage.
2. The method according to claim 1, wherein The desired auditory response is a response from a subject with normal hearing or a response with enhanced features.
3. The method according to claim 1 or 2, wherein The desired auditory response model and the personalized auditory response model include models of different stages of the auditory periphery.
4. The method according to claim 1 or 2, wherein The personalized auditory response model based on a neural network and the desired auditory response model based on a neural network comprise a convolutional encoder-decoder neural network having stride convolutions and skip connections.
5. The method according to claim 1 or 2, wherein A reference neural network simulating enhanced hearing perception and / or capabilities of a hearing-normal listener is used as the desired auditory response model; wherein, a corresponding hearing-normal neural network or a hearing-impaired neural network is used as the personalized auditory response model; and wherein, the individualized auditory signal processing model is a signal processing neural network model trained to process the auditory input and provide an enhanced auditory response.
6. The method according to claim 1 or 2, wherein The individualized auditory signal processing model is trained to minimize a specific auditory response difference metric.
7. The method according to claim 6, wherein, The specific auditory response difference metric is the absolute difference or the squared difference between two auditory response models at multiple tonal frequencies or all tonal frequencies.
8. The method according to claim 1 or 2, wherein The processed auditory output is a modified auditory stimulus designed to compensate for a hearing impairment or produce enhanced hearing.
9. The method according to claim 1 or 2, wherein The processed auditory output is a modified auditory response corresponding to a specific processing stage along the auditory pathway that can be used to stimulate an auditory prosthesis.
10. The method according to claim 9, wherein, The auditory prosthesis is a cochlear implant or a deep brain implant.
11. The method according to claim 1 or 2, wherein, Minimize the difference between the auditory nerve outputs of a hearing-normal periphery and a hearing-impaired periphery; or wherein, minimize the difference between simulated auditory brainstem and / or cortical responses expressed in the time domain or the frequency domain.
12. The method according to claim 1 or 2, wherein A task-optimized speech "backend" is connected to the output of the auditory response model, also referred to as the "front-end", and the task-optimized speech "backend" simulates the performance of the listener in different tasks; and wherein, the output of the backend is used to determine the auditory response difference and minimize the auditory response difference.
13. The method according to claim 1 or 2, the method being for configuring an auditory device, wherein, The auditory device is a cochlear implant or a wearable hearing aid.
14. Use of the method according to any one of claims 1 to 13 in a hearing aid application.
15. An auditory device, the auditory device comprising: - an input device configured to pick up input sound waves from the environment and convert the input sound waves into an auditory stimulus; - a processing unit configured to perform the method according to any one of claims 1 to 13 to generate a processed auditory output; and, - an output device configured to generate the processed auditory output from the processing unit.
16. The auditory device according to claim 15, the auditory device being a cochlear implant or a wearable hearing aid.
17. A computer program product, the computer program product comprising a computer program which, when executed by a computer, implements the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Binaural adaptive hearing aid
US20050069162A1