Method and device for audio signal generation
A machine learning model aligns neural responses to correct hearing impairments by transforming audio signals, addressing the limitations of current hearing aids and enhancing auditory perception.
Patent Information
- Application Number
- PCT/GB2025/051799
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-15
- Filing Date
- 2025-08-14
- Publication Date
- 2026-02-19
AI Technical Summary
Current hearing aids provide limited benefits for hearing-impaired individuals, particularly in complex audio environments, as they primarily rely on amplification and fail to address the neural distortions caused by hearing loss, leading to inadequate speech-in-noise perception and social isolation.
A machine learning model is trained to generate modified audio signals by aligning neural responses in hearing-impaired and hearing-unimpaired brains, using pre-trained modules to simulate auditory brain responses and minimize differences due to hearing impairment, while preserving individual brain variations.
The model effectively transforms audio signals to restore normal auditory perception by correcting neural responses, improving speech understanding and reducing distortions for both hearing-impaired and unimpaired individuals.
Smart Images

Figure GB2025051799_19022026_PF_FP_ABST
Abstract
Description
[0001] AL Ref: P46575WO1 14 August 2025 Method and Device for Audio Signal Generation Field
[0001] The present techniques generally relate to a method and user device for generating modified audio signals using a machine learning, ML, model. In particular, the present application relates to a computer-implemented method for training a ML model to generate modified audio signals that make it easier for people to interpret, hear or understand the content of the signals. Applications include modifying audio signals for hearing-impaired persons, and modifying distorted audio signals to provide distortion reduction or removal. Background
[0002] Approximately 500 million people globally are affected by hearing loss, making it the fourth leading cause of years lived with disability. The resulting burden imposes enormous personal and societal consequences. By impeding communication, hearing loss leads to social isolation and associated decreases in quality of life and wellbeing. It has also been linked to declines in mental health, with hearing loss having been identified as the leading modifiable risk factor for incident dementia. As the impact of hearing loss continues to grow, the need for improved treatments is becoming increasingly urgent. In most cases, the only treatment available is a hearing aid. But current devices provide only limited benefit in many real-world settings.
[0003] Current state-of-the-art hearing aids are little more than simple amplifiers. Most share the same core processing algorithm known as multi-channel wide dynamic range compression (WDRC). This algorithm provides listeners with frequency-specific amplification based on measured deficits in their hearing sensitivity to improve the audibility of low-intensity sounds. But it ignores many other critical aspects of hearing loss and its utility is limited for complex sounds such as speech in background noise or music. The introduction of directional microphones and algorithms for background noise suppression have helped to improve speech-in-noise perception in certain contexts, but performance remains unsatisfactory overall.
[0004] Therefore, the present applicant has identified the need for an improved method for generating modified audio signals so they can be better understood by hearing-impaired and hearing-unimpaired persons. Summary
[0005] In a first approach of the present techniques, there is provided a computer-implemented method for training a machine learning, ML, model to generate modified audio signals for AL Ref: P46575WO1 14 August 2025 individuals, the method comprising: obtaining a plurality of audio signals; for each audio signal of the plurality of audio signals: generating a first representation of a neural response to the audio signal in a first set of brains; transforming, using a sound transformation module of the ML model, the audio signal to generate a modified audio signal; generating a second representation of a neural response to the modified audio signal in a second set of brains; and generating an aligned second representation of a neural response in the second set of brains by minimising differences in the generated second representation relative to the generated first representation other than differences arising from one or both of: a difference in hearing ability between the first set of brains and the second set of brains, and a difference between the audio signal and the modified audio signal; and training the sound transformation module of the ML model by: minimising a loss function that is based on a difference between the aligned second representation of a neural response to the modified audio signal and the first representation of a neural response to the audio signal.
[0006] The term “sound transformation module” is used herein to mean an algorithm which receives an input audio signal and transforms the input audio signal to generate an output audio signal, wherein the input and output audio signals are, in general, different from each other. For example, the sound transformation module may be an encoder-decoder neural network, NN, architecture which has been trained to perform a sound transformation through an encoding-decoding process. However, it will be appreciated that a sound transformation module may take any appropriate form for generating an output audio signal based on an input audio signal.
[0007] The term “representation of a neural response” is used to broadly mean any way to represent a predicted neural response to an auditory stimulus. As explained in more detail with reference to the examples below, the “representation of a neural response” may be a simulation of neural activity in brains in response to the auditory stimulus, or may be a latent representation of underlying neural dynamics. In one example, an auditory brain response may refer to the activity generated in the inferior colliculus (IC) of the human brain by human speech. However, it will be appreciated that an auditory brain response may refer to activity induced in any part of a brain in response to any type of sound.
[0008] The term “modified audio signal” used herein is used to mean an input audio signal that has been transformed by the sound transformation module. Three example use cases of the present techniques are described below: (i) hearing impairments, (ii) distortion reduction or removal, and (iii) hearing impairments and distortion reduction / removal. In example (i), during training, the sound transformation module transforms an input audio signal to generate a modified audio signal that is more suitable for hearing-impaired individuals. Thus, the input into the sound transformation module is the same audio signal that is used to generate a first AL Ref: P46575WO1 14 August 2025 representation of a neural response. In example (ii), during training, the sound transformation module receives an already altered version of the input audio signal, and uses this to generate a modified audio signal in which distortions are removed or reduced. The already altered version of the input audio signal is therefore further modified by the sound transformation module. The already altered or pre-altered version of the input audio signal has been altered to include a distortion, as described below. In example (iii), during training, the sound transformation module receives an already altered version of the input audio signal, and uses this to generate a modified audio signal in which distortions are removed or reduced and which is more suitable for hearing-impaired individuals. The term “modified audio signal” is referring to the action performed by the sound transformation module to transform the input audio signal, whether that is the originally obtained input audio signal or a pre-altered version of the obtained input audio signal.
[0009] Advantageously, the present techniques provide an effective way of training a ML model to generate modified audio signals that improve what a hearer is able to hear and understand. The hearer may have normal / unimpaired hearing, or may have impaired hearing.
[0010] One focus of the present techniques is the restoration of brain responses in hearing- impaired persons such that the brain responses appear “healthy” or as they would otherwise have been produced without the hearing impairment. To do so, a sound transformation module learns a sound transformation so that, when an input audio signal is transformed to a modified audio signal, the modified audio signal produces a brain response that is equivalent to that which would have been produced by the input audio signal without the hearing impairment. Thus, auditory perception may be corrected in a hearing-impaired brain by virtue of the sound transformation.
[0011] General Features
[0012] In some cases, the step of generating a second representation and the step of generating an aligned second representation may be performed one after the other. Alternatively, the step of generating a second representation and the step of generating an aligned second representation may be performed simultaneously / jointly. Different modules of the ML model may be used to perform each of these steps, or the same module may be used for both steps.
[0013] In some cases, an “alignment module” of the ML model may be used to perform the step of generating an aligned second representation. The term “alignment module” is used herein to mean an algorithm for performing a transformation / modification of a neural response in a way that enhances the similarity between (or aligns) a pair of neural responses. For example, the alignment module may be used to generate an aligned neural response by aligning the AL Ref: P46575WO1 14 August 2025 first and second neural responses. In other words, the similarity between the first and second neural responses is enhanced by the alignment.
[0014] As noted above, current techniques for treating hearing impairments generally operate through a crude amplification and underperform for complex audio signals. Instead, the present techniques operate at the level of the brain response, and the use of an ML model which learns a sound transformation based on brain responses ensures that audio signals can be transformed accurately and without loss of information.
[0015] The present techniques may make use of pre-trained modules which simulate auditory brain responses, based on techniques described in Sabesan, S., et al. These pre-trained modules are part of the ML model being trained, but have already been trained using data obtained using electrodes positioned in real brains, as opposed to simulation data. Thus, a first pre-trained module may simulate auditory brain responses in at least one hearing- unimpaired brain, wherein the first pre-trained module is trained using electrode data from a real hearing-unimpaired brain. Similarly, a second pre-trained module may simulate auditory brain response in at least one hearing-impaired brain, having been trained using electrode data from a real hearing-impaired brain. However, in addition to differences between the first and second neural responses arising from the hearing impairment, there are also individual differences between brains i.e. differences that are not related to the hearing impairment. This is because there is always a degree of natural variation in brain responses arising due to the individual characteristics of brains.
[0016] Thus, in training the ML model to correct for the hearing impairment and / or to correct for differences between distorted or noisy audio signals and undistorted or clean audio signals, it is important that the ML model does not also learn how to correct for these individual differences. This is achieved by the generation of the aligned second representation of a neural response above. For example, generating an aligned second representation of a neural response may comprise: modifying, using the alignment module, the second neural response to generate the aligned second neural response, by: minimising differences between the first and second neural responses arising due to individual differences between brains, and preserving the differences between the first and second neural responses arising only from hearing impairments. In other words, an alignment process is performed, wherein the second neural response is modified in a way that corrects for individual differences between brains, whilst preserving the differences arising only from the hearing impairment. This is so that the ML model learns only to correct for the hearing impairment.
[0017] There are multiple different ways that the aligned second neural response may be generated. In other words, there are multiple possible modifications to the second neural AL Ref: P46575WO1 14 August 2025 response that bring the second neural response into alignment with the first neural response, whilst maintaining the differences due to the hearing impairment.
[0018] In a first example, generating an aligned second neural response may comprise: applying, using the alignment module, a projection matrix to the second neural response, where values of the projection matrix are optimised to align first and second neural responses. In other words, the aligned second neural response may be generated through a linear transformation of the second neural response. In this example, applying a projection matrix to the second neural response may comprise: constraining the projection matrix to be a permutation matrix having a plurality of columns, where each column has a plurality of zero elements and a single non-zero element. That is, the projection matrix may take the form of a permutation matrix, which serves to, for example, re-arrange or permute components of the second neural response, so that the re-arranged components more closely resemble corresponding components of the first neural response.
[0019] In the case where the projection matrix is constrained to be a permutation matrix, training the alignment module may comprise adjusting a position of the non-zero elements in each column of the permutation matrix. In other words, the permutation matrix is recursively adjusted during the training to produce an optimal permutation that has the desired characteristics noted above i.e. to minimise those differences arising due to individual brain differences and to preserve those arising only from the hearing impairment.
[0020] In a second example, generating an aligned second neural response may comprise: decomposing, using the alignment module, a covariance matrix into a first weight matrix and a second weight matrix, wherein the covariance matrix corresponds to a covariance of the first and second neural responses; and applying the first and second weight matrices to the second neural response. That is, a statistical measure of the difference between the first and second neural responses, i.e. the covariance matrix, may be used to generate the aligned second neural response. The covariance matrix may be decomposed into weight matrices which encode these differences, and applied to the second neural response. In this case, training the alignment module may comprise adjusting values of elements in the first and second weight matrices. In other words, values of elements in the first and second weight matrices may be recursively adjusted, in a similar way to the example of the permutation matrix above.
[0021] As noted above, the first and second pre-trained modules are trained on electrode data from real brains. In particular, the first and second pre-trained modules may be trained on data obtained from a plurality of electrode channels. Accordingly, the first pre-trained module may simulate auditory brain responses in electrode channels measuring auditory brain response in real brains to audio inputs. In this case, generating the first neural response AL Ref: P46575WO1 14 August 2025 comprises generating a response for each electrode channel in the first plurality of electrode channels to the audio signal.
[0022] Similarly, the second pre-trained module may simulate auditory brain responses produced in a second plurality of electrode channels measuring auditory brain responses in real brains to audio inputs. In this case, generating the second neural response comprises: generating a response for each electrode channel in the second plurality of electrode channels to the modified audio signal.
[0023] Differences in the positions of electrode channels within the real brains are also a potential source of discrepancy between the first and second neural responses, in addition to the individual characteristics of the brains. Thus, the aligned second neural response may be generated by transforming electrode channels positions of the second neural response, so that these differences are minimised.
[0024] Thus, in a third example, generating an aligned second neural response may comprise: applying, using the alignment module, a transformation matrix to the second neural response, wherein the transformation matrix translates and / or rotates electrode channel positions in the second neural response. In this case, each channel in the second plurality of electrode channels may have a neural position in the hearing-impaired brain, and applying the transformation matrix may comprise transforming, using the translation and / or rotation of the transformation matrix, the neural positions of each electrode channel to generate a plurality of transformed electrode channel positions.
[0025] In this third example, generating an aligned second neural response may further comprise modifying the response in each electrode channel of the second plurality of electrode channels to generate an aligned response in each electrode channel. That is, channel-wise modifications may be made to the second neural response. Modifying the response in each electrode channel may comprise: for each response generated in the second plurality of electrode channels: applying a sampling kernel to the response to generate an electrode channel contribution, wherein the sampling kernel is based on the transformed electrode channel position and a corresponding position of an equivalent electrode channel in the first plurality of electrode channels; and summing the electrode channel contributions to generate the aligned response in the electrode channel.
[0026] Minimising a loss function based on a difference between the aligned second neural response to the modified audio signal and the first neural response to the audio signal may comprise: minimising a difference between the aligned second neural response and the first neural response in corresponding electrode channels by adjusting translation and / or rotation parameters of the transformation matrix. AL Ref: P46575WO1 14 August 2025
[0027] As noted above, the first pre-trained module simulates auditory brain responses in at least one hearing-unimpaired brain. That is, the first neural response may be based on the simulation of a single hearing-unimpaired brain. In other cases, the first neural response may be based on simulating auditory brain responses in multiple brains. In the case of a multi- brain response, the pre-trained module has learned how to generate a single latent representation that is sufficient for generating all of the responses for the multiple brains, but does not need to actually generate the individual responses themselves. The single latent representation is then used to provide the first, single neural response for the multiple brains. This is so that the first neural response is representative of a generic neural response, rather than being specific to a particular hearing-unimpaired brain.
[0028] The step of generating an aligned second representation of a neural response may comprise minimising differences in the generated second representative relative to the generated first representation arising from one or both of: anatomical differences between the first set of brains and the second set of brains that are unrelated to hearing impairments; and differences in position of apparatus for recording neural responses to audio signals. Thus as noted above, in order to generate the modified audio signal, it is necessary to focus on only those differences between the responses to an audio signal that arise from differences in hearing ability (in the case of training the model for hearing impaired people), and / or differences between the audio signals (in the case of training the model to provide distortion reduction or removal). Thus, aligning the representations may comprise minimising differences arising from the natural differences between the brains and in any recording equipment.
[0029] The step of obtaining a plurality of audio signals may comprise obtaining a training dataset comprising a plurality of audio signals. Thus, the training dataset may be a static dataset of audio signals. The audio signals in the dataset may be the same as, or be similar to, the audio signals used in the training of the pre-trained modules described below.
[0030] Alternatively, obtaining a plurality of audio signals may comprise using a generative model to generate a plurality of audio signals. Thus, a separate generative model may be used to generate audio signals having the required characteristics / features for the training. Again, the generated audio signals may be the same as, or be similar to, the audio signals used in the training of the pre-trained modules described below.
[0031] Obtaining the plurality of audio signals may comprise obtaining audio signals that are single-channel and / or multi-channel.
[0032] Obtaining the plurality of audio signals may comprise obtaining a plurality of audio signals comprising speech. In other words, the sound transformation module may be trained for the important use-case of transforming speech audio signals. AL Ref: P46575WO1 14 August 2025
[0033] In the case where obtaining the plurality of audio signals comprises obtaining a plurality of audio signals comprising speech, the method may further comprise: identifying, using an automatic speech recognition, ASR, module, phonemes from the first representation of a neural response and from the aligned second representation of a neural response; wherein minimising a loss function may comprise minimising the loss function using differences between the identified phonemes. That is, a pre-trained automatic speech recognition module may be used to ensure that the sound is transformed accurately by identifying phonemes from the first neural response and the aligned second neural response and by incorporating any differences between the identified phonemes into the loss function for training the ML model.
[0034] More generally, in the case where obtaining the plurality of audio signals comprises obtaining a plurality of audio signals comprising speech, the method may comprise using a perceptual loss for optimisation of the training process. That is, ASR is one non-limiting example technique that may be used to perform the training, and other example perception- based techniques, such as sound quality, intelligibility, listening experience, listening effort, and so on may be used to perform the training (via incorporation into the loss function used for training the ML model). In other words, one or more additional loss terms (in addition to the loss based on a difference between the aligned second representation of a neural response to the modified audio signal and the first representation of a neural response to the audio signal) may be used during the training of the model. The one or more additional loss terms may arise from additional downstream perception tasks, such as ASR, sound quality, etc. Each of these downstream perception tasks may be performed using an additional module or backend.
[0035] In some cases, obtaining the plurality of audio signals may comprise obtaining a plurality of audio signals comprising music. In such cases, the method may comprise using a perceptual loss for optimisation of the training process based on the perception of the modified audio signal, such as music sound quality, intelligibility, listening experience, listening effort, understandability of any lyrics, etc.
[0036] More generally, the method may further comprise personalising the training for a specific hearing-impaired person. That is, to enhance the personalisation of the present techniques such that the sound transformation is tailored to a specific hearing-impaired person, the training of the ML model may be personalised. This may be done through encoding the audio signal into a latent representation, and decoding the latent representation with the incorporation of additional personalisation parameters. These personalisation parameters quantify the severity and full character of the hearing impairment in the brain of the specific person, where the personalisation parameters are in a continuous parameter space. Advantageously, the present techniques may correct any or all of the effects of hearing loss AL Ref: P46575WO1 14 August 2025 and not just the severity. That is, the present techniques may quantify the full effects of a hearing impairment.
[0037] In other words, the method may further comprise: generating a latent representation of the audio signal; and obtaining a personalisation parameter for the specific hearing-impaired person; wherein generating the second representation of a neural response comprises transforming, using the obtained personalisation parameter and a personalisation parameter for each hearing-impaired brain in the second set of brains, the latent representation to generate the second neural response for the specific hearing-impaired person, wherein the personalisation parameter corresponds to a quantification of the full effects of the hearing impairment in the brain of the specific hearing-impaired person.
[0038] Hearing Impairment
[0039] In some cases, the method is for training a machine learning, ML, model to generate modified audio signals for hearing-impaired individuals. In such cases, the first set of brains may be a set of hearing-unimpaired brains, and the second set of brains may be a set of hearing-impaired brains, wherein the hearing-impaired brains are changed due to a hearing impairment. In such cases, the step of transforming the audio signal to generate a modified audio signal may comprise transforming the audio signal for hearing-impaired individuals. Thus, the modified audio signal that is generated by the sound transformation module may be modified so that a neural response to the modified audio signal in hearing-impaired brains is similar to a neural response to the (original) audio signal in hearing-unimpaired brains.
[0040] In a first example, generating the first representation of a neural response to the audio signal may comprises using a first pre-trained forward model of the ML model, and generating the second representation of a neural response to the modified audio signal may comprise using a second pre-trained forward model of the ML model.
[0041] In this example, the first pre-trained forward model may be for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation comprises generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains. Similarly, the second pre-trained forward model may be for simulating neural responses in the second set of hearing-impaired brains, and generating the second representation comprises generating a single second neural response to the modified audio signal that reflects common features of neural response across the second set of hearing-impaired brains.
[0042] The term “neural response” is used herein to mean a simulated auditory brain response generated using ML techniques. For example, the first pre-trained module of the ML model has been trained to simulate auditory brain responses in at least one hearing-unimpaired brain, thereby generating a first neural response. In other words, the first pre-trained module AL Ref: P46575WO1 14 August 2025 simulates neural activity in brain(s) without a hearing impairment. Similarly, the second pre- trained module of the ML model above simulates neural activity in brain(s) with a hearing impairment.
[0043] The pre-trained forward models may be or may comprise encoders and decoders. In this example, generating the single first neural response to the audio signal may comprise: generating, using an encoder of the first pre-trained forward model, a representation of the audio signal, and generating, using a decoder of the first pre-trained forward model and the representation of the audio signal, a single representation of the neural activity in the first set of hearing-unimpaired brains. Similarly, generating the second neural response to the modified audio signal may comprise: generating, using an encoder of the second pre-trained forward model, a representation of the modified audio signal, and generating, using a decoder of the second pre-trained forward model and the representation of the modified audio signal, a representation of the neural activity in each brain in the second set of hearing-impaired brains.
[0044] In a second example, generating the first representation of a neural response to the audio signal may comprise using a first pair of pre-trained encoders of the ML model; and generating the second representation of a neural response to the modified audio signal may comprise using a second pair of pre-trained encoders of the ML model. Thus, instead of using pre-trained forward models comprising encoders and decoders, in this example, the decoders are ‘reversed’ in direction to act as encoders, such that sound and neural activity are mapped into a shared latent space between the pairs of encoders.
[0045] In this example, generating the first representation may comprise: generating, using a first encoder of the first pair of pre-trained encoders, a first latent representation of features of the audio signal; generating, using a second encoder of the first pair of the pre-trained encoders, a second latent representation of features of a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing- unimpaired brains, wherein the first and second latent representations are in the same latent space; and generating a first representation by minimising a difference between the first and second latent representations. Similarly, generating the second representation may comprise: generating, using a first encoder of the second pair of pre-trained encoders, a third latent representation of features of the modified audio signal; and generating, using a second encoder of the second pair of the pre-trained encoders, a fourth latent representation of features of a second neural response to the audio signal for each brain in the second set of hearing-impaired brains, wherein the third and fourth latent representations are in the same latent space; and generating a second representation by minimising a difference between the third and fourth latent representations. AL Ref: P46575WO1 14 August 2025
[0046] Distortion Removal / Reduction
[0047] In some cases, the method is for training a machine learning, ML, model to generate less distorted audio signals. In such cases, the first set of brains may be a set of hearing- unimpaired brains; and the second set of brains may also be a set of hearing-unimpaired brains. The first and second set of brains may be the same or different.
[0048] The first representation may be generated by determining neural responses to the audio signals that are considered “clean” audio signals. That is, these audio signals do not have any distortions or have minimal distortion. The second representation may be generated by determining neural responses to pre-distorted versions of the same clean / undistorted audio signals. That is, the distorted versions contain distortions. Thus, the goal of the training in these cases is for the sound transformation module to transform the distorted audio signals so that the modified distorted audio signals generate a similar response in the second set of brains as the clean / undistorted audio signals generate in the first set of brains. In other words, the sound transformation module receives as input a pre-distorted version of an audio signal, and transforms this into a modified audio signal, which is also referred to herein as a “modified distorted audio signal”.
[0049] In such cases, the method may further comprise: modifying, prior to transforming the audio signal, the audio signal of the plurality of audio signals to include distortions, thereby generating a modified distorted audio signal. That is, the signal that is used to generate a second representation to the modified audio signal in the second set of brains may be additionally modified to include a distortion or distortions. This is in addition to the modifications which take place to cause the responses in the second set of brains to be similar to the responses in the first set of brains. Any sort of distortion may be added to generate the modified distorted audio signal such as, for example, low frequency noise, impulsive noise, intermittent noise, continuous noise, background noise, speech, babble noise, ambient or environmental noise, reverberation, compression, and so on.
[0050] In one example, generating the first representation of a neural response to the audio signal may comprise using a first pre-trained forward model of the ML model for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation may comprise generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains. Similarly, generating the second representation of a neural response to the modified audio signal may comprise using a second pre-trained forward model of the ML model for simulating neural responses in the second set of hearing-unimpaired brains, and generating the second representation may comprise generating a second neural response to the modified noisy audio signal for each brain in the second set of hearing-unimpaired brains. As described above with AL Ref: P46575WO1 14 August 2025 reference to the hearing impairment cases, the pre-trained forward models may comprise encoders and decoders. Thus, the features described above apply similarly here.
[0051] In another example, generating the first representation of a neural response to the audio signal may comprise using a first pair of pre-trained encoders of the ML model; and generating the second representation of a neural response to the modified audio signal comprises using a second pair of pre-trained encoders of the ML model. In this example, generating the first representation may comprise: generating, using a first encoder of the first pair of pre-trained encoders, a first latent representation of features of the audio signal; generating, using a second encoder of the first pair of the pre-trained encoders, a second latent representation of features of a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains, wherein the first and second latent representations are in the same latent space; and generating a first representation by minimising a difference between the first and second latent representations. Similarly, generating the second representation may comprise: generating, using a first encoder of the second pair of pre-trained encoders, a third latent representation of features of the modified audio signal (where the modified audio signal contained distortions prior to the modifying); and generating, using a second encoder of the second pair of the pre-trained encoders, a fourth latent representation of features of a second neural response to the audio signal for each brain in the second set of hearing-impaired brains, wherein the third and fourth latent representations are in the same latent space; and generating a second representation by minimising a difference between the third and fourth latent representations.
[0052] Distortion Cancellation or Reduction for Hearing Impairments
[0053] In some cases, the techniques for distortion cancellation / reduction and the techniques for hearing impairments may be combined. Thus, the method is for training a machine learning, ML, model to generate less noisy audio signals. In such cases, the first set of brains may be a set of hearing-unimpaired brains; and the second set of brains may be a set of hearing-impaired brains.
[0054] In such cases, the method may further comprise: modifying, prior to transforming the audio signal, the audio signal of the plurality of audio signals to include distortion, thereby generating a modified distorted audio signal.
[0055] In such cases, generating the first representation of a neural response to the audio signal may comprise using a first pre-trained forward model of the ML model for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation may comprise generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains. Similarly, generating the second representation of a neural response to the modified audio AL Ref: P46575WO1 14 August 2025 signal may comprise using a second pre-trained forward model of the ML model for simulating neural responses in the second set of hearing-impaired brains, and generating the second representation may comprise generating a second neural response to the modified distorted audio signal for each brain in the second set of hearing-impaired brains.
[0056] More generally, the features described above with respect to the hearing impairment and the distortion reduction examples apply equally to this example where both problems are tackled together.
[0057] In a second approach of the present techniques, there is provided a computer- implemented method for generating, on a user device, modified audio signals for a hearing- impaired person using a trained machine learning, ML, model trained according to the method set out above in relation to the first approach, the method comprising: receiving an input audio signal; generating, using a sound transformation module of the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the hearing-impaired person that is similar to an auditory brain response to the input audio signal in a hearing- unimpaired brain; and outputting the modified audio signal.
[0058] In a third approach of the present techniques, there is provided a computer-implemented method for generating, on a user device, modified audio signals for distortion reduction / removal using a trained machine learning, ML, model trained according to the method set out above in relation to the first approach, the method comprising: receiving a distorted input audio signal; generating, using a sound transformation module of the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the user of the user device that is similar to an auditory brain response to a clean / undistorted version of the distorted input audio signal; and outputting the modified audio signal. In a fourth approach of the present techniques, there is provided a user device for generating modified audio signals for a hearing-impaired person using a trained machine learning, ML, model on a user device trained according to the method set out above in relation to the first approach, the user device comprising: storage for storing the trained ML model; and a processor coupled to memory, for: receiving an input audio signal; generating, using the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the hearing- impaired person that is similar to an auditory brain response to the input audio signal in a hearing-unimpaired person; and outputting the modified audio signal.
[0059] The user device may be a constrained-resource device, but which has the minimum hardware capabilities to use a trained ML model. The user device may be any of: a hearing aid; a cochlear implant; a pair of headphones; earphones; a smartphone; a tablet; a headset; a smart speaker; and a speaker.
[0060] The user device may further comprise a speaker for outputting the modified audio signal. AL Ref: P46575WO1 14 August 2025
[0061] Alternatively, the user device may further comprise electronics for outputting electrical stimulation based on the modified audio signal. For example, a cochlear implant generates electrical shocks or impulses, and therefore, the modified audio signal may be used to alter the electrical shocks / impulses that are generated and applied to stimulate the cochlear nerve.
[0062] The user device may further comprise a microphone or audio receiver for obtaining the input audio signal which is to be modified.
[0063] The features described above with respect to the first approach apply equally to the second, third and fourth approaches and therefore, for the sake of conciseness, are not repeated.
[0064] In a related approach of the present techniques, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, causes the processor to carry out any of the methods described herein.
[0065] As will be appreciated by one skilled in the art, the present techniques may be embodied as a system, method or computer program product. Accordingly, present techniques may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.
[0066] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
[0067] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise sub- components which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high- level compiled or interpreted language constructs.
[0068] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.
[0069] The techniques further provide processor control code to implement the above- described methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data AL Ref: P46575WO1 14 August 2025 carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD- ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system. Brief description of the drawings
[0070] Implementations of the present techniques will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0071] Figure 1 illustrates the disadvantages of current techniques for treating hearing impairments;
[0072] Figure 2A is a block diagram of an example process for training a ML model for generating modified audio signals;
[0073] Figure 2B shows another example process for training a ML model for generating modified audio signals;
[0074] Figure 3A is a flowchart of steps for training a ML model to generate modified audio signals;
[0075] Figure 3B is a flowchart of example steps for training a ML model to generate modified speech audio signals;
[0076] Figure 4A is a block diagram of a process for training the first and second pre-trained modules;
[0077] Figure 4B shows results of an intracranial recording process for hearing-impaired and hearing-unimpaired brains;
[0078] Figure 4C shows comparisons of simulated neural activity for hearing-unimpaired and hearing-impaired brains using the first and second pre-trained modules;
[0079] Figure 5A is a flowchart of steps in a first example process for generating the aligned second neural response using the alignment module;
[0080] Figure 5B is a flowchart of steps in a second example process for performing the alignment; AL Ref: P46575WO1 14 August 2025
[0081] Figure 5C is a flowchart of steps in a third example process for generating the aligned second neural response using the alignment module;
[0082] Figure 5D is a schematic illustration of the transformation of Figure 5C;
[0083] Figure 6A is a block diagram of a process to train a first pre-trained module which simulates auditory brain responses in a plurality of hearing-unimpaired brains;
[0084] Figure 6B is a block diagram of a process for training the ML model using the first pre- trained module of Figure 6A;
[0085] Figure 7A is a block diagram of a process to train a second pre-trained module which simulates auditory brain responses in hearing-impaired brains;
[0086] Figure 7B is a block diagram of a process for training the ML model using the second pre-trained module of Figure 7A;
[0087] Figure 8 is a flowchart of steps in a process for using a trained ML model for generating modified audio signals;
[0088] Figure 9 is a block diagram of a user device for generating modified audio signals for a hearing-impaired person using a trained ML model;
[0089] Figure 10A shows a performance comparison of the present techniques against existing techniques as a result of the experiments above; and
[0090] Figure 10B shows the ASR performance of the present techniques. Detailed description of the drawings
[0091] Broadly speaking, the present techniques generally relate to a method and user device for generating modified audio signals using a machine learning, ML, model. In particular, the present application relates to a computer-implemented method for training a ML model to generate modified audio signals for hearing-impaired persons and / or for distortion cancellation, and a user device for generating modified audio signals using a trained ML model.
[0092] Figure 1 illustrates the disadvantages of current techniques for treating hearing impairments. Current approaches to treating hearing impairments have shortcomings, particularly for complex audio signals. The present techniques take a new approach to treating hearing impairments, by operating from the perspective of the neural code – the brain activity patterns that underlie perception. If hearing is impaired, it is because the details of the neural code have become distorted. One simple way to frame the problem of hearing aid design is as a search for sound transformations that correct these distortions. If a hearing aid can process sound such that it elicits neural activity in an impaired system that matches the activity elicited by the original sound in a healthy system, auditory perception will be restored to normal. AL Ref: P46575WO1 14 August 2025
[0093] There are existing approaches for correcting distortions in the neural code. Approaches based on this idea have been developed using cochlear models, but reliance on these models imposes fundamental limitations. Existing cochlear models have not been directly fit to experimental data, but rather hand-designed to reproduce generic effects of hearing loss on specific biophysical or perceptual phenomena. It is not clear how well the simulation of hearing loss in these models matches the complex effects of hearing loss in real ears.
[0094] Furthermore, even if the models were perfect and could be used to design a hearing aid that corrected the neural code in the cochlea, hearing would not be restored to normal. While hearing loss almost always begins with cochlear damage, the problem then spreads from the ear to the brain, causing plastic changes in the central auditory pathway that are not easily reversed. When these plastic changes are extreme, they can cause disorders such as tinnitus and hyperacusis. But even moderate plastic changes ensure that even with perfect restoration of the neural code in the cochlea, neural activity at higher levels will remain distorted because of altered central auditory processing.
[0095] To compensate for both cochlear damage and downstream plasticity, the present techniques seek to identify sound transformations that can restore distorted central neural activity to normal. The present techniques focus on the inferior colliculus (IC), the hub of the central auditory pathway. The IC is an obligatory bottleneck where multiple brainstem inputs converge to create a full neural representation of the auditory scene. It is late enough in the auditory pathway to capture many of the plastic changes that follow hearing loss, but also early enough that its neural activity is still driven primarily by raw acoustics, without strong modulation by contextual factors or other sensory modalities.
[0096] Identifying the sound transformation required to correct distortions in the neural code is easier said than done. Auditory processing is highly nonlinear, and it is impossible to hand design sound transformations that can compensate for the effects of hearing loss in their full complexity. But recent advances in deep learning provide tools that are well suited to this type of challenge. Unique experimental methodologies that have been developed previously (Armstrong, A., et al.) have the capacity to provide the large-scale, high-resolution datasets required to take full advantage of machine learning and deep learning in this context. Below, the approach of the present techniques for optimal hearing aid design through large-scale electrophysiology and deep learning is described, with illustrations of the potential to correct distortions in the neural code.
[0097] Example Processes for Training an ML Model
[0098] There are many ways in which the ML model training may be performed. Two examples are now described to illustrate the training, but it will be understood that alternatives also exist. AL Ref: P46575WO1 14 August 2025
[0099] Figure 2A is a block diagram of an example process for training a ML model for generating modified audio signals. As noted above, the aim of the present techniques is to identify a sound transformation which corrects distortions in neural responses arising due to a hearing impairment. The key elements that are needed for training the ML model 100 are as follows: (1) differentiable forward models of neural coding in normal and impaired brains; (2) a method for aligning activity from two brains without correcting differences related to hearing loss or a hearing impairment; and (3) a differentiable sound transformation module that acts as a hearing aid. These individual elements can take many forms and can be combined in many ways to design the overall training process. A number of approaches are described below, with a focus on alignment and the design of the overall process.
[0100] In more detail, the ML model comprises a sound transformation module 102, a first pre- trained module 104 (denoted ^^^), a second pre-trained module 106 (denoted ^^^) and analignment module 108. The sound transformation module 102 ℎ: ^ ∈ ^^^^ → ^∗ ∈ ^^^^ thattakes an input sound ^ and produces an output sound ^∗can be trained to minimise the difference between the activity elicited by ^ in a normal brain and activity elicited by ^∗in a hearing-impaired brain. Parameters ^ of the sound transformation module 102 are trained tominimise measure ofthe difference between neural activity patterns of hearing-impaired and hearing-unimpaired brains.
[0101] Sounds are presented to pre-trained modules 104, 106 (also referred to interchangeably herein as “ICNets”) for simulating auditory brain responses. The first pre-trained module 104 simulates auditory brain responses in at least one hearing-unimpaired brain, and generates a first neural response to the input sound ^ (i.e. an audio signal). The second pre-trained module 106 simulates auditory brain responses in at least one hearing-impaired brain, and generates a second neural response to the output sound ^∗(i.e. a modified audio signal).
[0102] The sound transformation module 102 is trained to process sound to minimise the difference between normal activity elicited by the original sounds and impaired activity elicited by the processed sounds, or equivalently, between the first and second neural responses. The sound transformation module 102 is trained jointly with an alignment module 108 to align the normal and impaired activity. An automatic speech recognition module, ASR, 110 may also be used to encourage effective neural coding of speech.
[0103] Figure 2B shows another example process for training a ML model for generating modified audio signals. Here, the forward model of Figure 2A is reduced to an encoder that maps sound into a latent space (without decoding from the latent space to neural activity), while the alignment is performed by another encoder that maps neural activity into the same AL Ref: P46575WO1 14 August 2025 latent space. The ML model is then trained to minimise differences in latent representations. A full description of Figure 2B is provided below.
[0104] Figure 3A is a flowchart of steps for training a ML model to generate modified audio signals. The training process is a computer-implemented method, and may be implemented on a server or on a computing device with sufficient resources to train a ML model. The method comprises: obtaining a plurality of audio signals (step S100); for each audio signal of the plurality of audio signals: generating a first representation of a neural response to the audio signal in a first set of brains (step S102); transforming, using a sound transformation module of the ML model, the audio signal to generate a modified audio signal (step S104); generating a second representation of a neural response to the modified audio signal in a second set of brains (step S106); and generating an aligned second representation of a neural response in the second set of brains by minimising differences in the generated second representation relative to the generated first representation other than differences arising from one or both of: a difference in hearing ability between the first set of brains and the second set of brains, and a difference between the audio signal and the modified audio signal (step S108); and training the sound transformation module of the ML model by: minimising a loss function that is based on a difference between the aligned second representation of a neural response to the modified audio signal and the first representation of a neural response to the audio signal (step S110).
[0105] As noted above, it is important that the ML model does not also learn how to correct for individual differences between brains that are unrelated to a hearing impairment, and / or differences between the distorted and undistorted audio signals. This is achieved by the generation of the aligned second neural response above (step S108). In other words, an alignment process is performed, wherein the second neural response is modified in a way that corrects for individual differences between brains (whilst preserving the differences arising only from the hearing impairment) and / or for differences between the signals themselves (in the case of noisy / distorted and clean / undistorted signals). This is so that the ML model learns only to correct for the hearing impairment and / or distortion cancellation, and not to correct for other differences.
[0106] Step S108 of generating the aligned second neural response may be performed in multiple alternative ways, as described below in relation to Figures 5A to 5C.
[0107] It will be appreciated that steps S102 to S110 are performed for all of the obtained audio signals in the training dataset.
[0108] In some cases, the step of generating a second representation (step S106) and the step of generating an aligned second representation (step S108) may be performed one after the other. Alternatively, the step of generating a second representation and the step of generating AL Ref: P46575WO1 14 August 2025 an aligned second representation may be performed simultaneously / jointly. Different modules of the ML model may be used to perform each of these steps, or the same module may be used for both steps.
[0109] The step S100 of obtaining a plurality of audio signals may comprise obtaining a training dataset comprising a plurality of audio signals. Thus, the training dataset may be a static dataset of audio signals. The audio signals in the dataset may be the same as, or be similar to, the audio signals used in the training of the pre-trained modules described below.
[0110] Alternatively, obtaining a plurality of audio signals at step S100 may comprise using a generative model to generate a plurality of audio signals. Thus, a separate generative model may be used to generate audio signals having the required characteristics / features for the training. Again, the generated audio signals may be the same as, or be similar to, the audio signals used in the training of the pre-trained modules described below.
[0111] Figure 3B is a flowchart of example steps for training a ML model to generate modified speech audio signals. That is, the present techniques may be applied to the specific task of generating modified speech audio signals. In this case, step S100 of Figure 3B for obtaining a training dataset comprises obtaining a plurality of audio signals comprising speech (step S200). The method proceeds according to steps S102 to S108 of Figure 2A, and their description is not repeated, for conciseness. The method further comprises identifying, using an automatic speech recognition, ASR, module, phonemes from the first neural response and from the aligned second neural response (step S202). Furthermore, step S110 of Figure 3A for minimising a loss function based on a difference between the aligned second neural response to the modified audio signal and the first neural response to the audio signal comprises using differences between the identified phonemes (step S204).
[0112] More generally, in the case where the plurality of audio signals comprise speech, the method may comprise using a perceptual loss for optimisation of the training process. That is, ASR is one non-limiting example technique that may be used to perform the training, and other example perception-based techniques, such as sound quality, intelligibility, listening experience, listening effort, and so on may be used to perform the training (via incorporation into the loss function used for training the ML model).
[0113] In some cases, the method in Figure 3A is for training a machine learning, ML, model to generate modified audio signals for hearing-impaired individuals. In such cases, the first set of brains may be a set of hearing-unimpaired brains, and the second set of brains may be a set of hearing-impaired brains, wherein the hearing-impaired brains are changed due to a hearing impairment. In such cases, the step of transforming the audio signal to generate a modified audio signal may comprise transforming the audio signal for hearing-impaired individuals. Thus, the modified audio signal that is generated by the sound transformation AL Ref: P46575WO1 14 August 2025 module may be modified so that a neural response to the modified audio signal in hearing- impaired brains is similar to a neural response to the (original) audio signal in hearing- unimpaired brains.
[0114] In a first example, generating the first representation of a neural response to the audio signal may comprises using a first pre-trained forward model of the ML model, and generating the second representation of a neural response to the modified audio signal may comprise using a second pre-trained forward model of the ML model. See also Figure 2A for a block diagram illustrating this example.
[0115] In this example, the first pre-trained forward model may be for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation comprises generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains. Similarly, the second pre-trained forward model may be for simulating neural responses in the second set of hearing-impaired brains, and generating the second representation comprises generating a second neural response to the modified audio signal for each brain in the second set of hearing-impaired brains. Advantageously, while the first representation is a single first neural response reflective of responses by multiple brains in the first set of brains, the second representation is individualised, i.e. is specifically for each brain in the second set of brains. This is advantageous because the model is learning how to generate modified audio signals suitable for individuals with different hearing impairments or different effects of hearing impairment. This is in contrast to standard hearing aids which deliver the same signal or same set of signals to all users, irrespective of their specific hearing impairments.
[0116] The term “neural response” is used herein to mean a simulated auditory brain response generated using ML techniques. For example, the first pre-trained module of the ML model has been trained to simulate auditory brain responses in at least one hearing-unimpaired brain, thereby generating a first neural response. In other words, the first pre-trained module simulates neural activity in brain(s) without a hearing impairment. Similarly, the second pre- trained ML model above simulates neural activity in brain(s) with a hearing impairment. In some cases the neural response is a “predicted neural response” which uses simulated activity to determine a response to an input audio signal. In other cases, the neural response is a “latent representation” of a neural response.
[0117] The pre-trained forward models may be or may comprise encoders and decoders. In this example, generating the single first neural response to the audio signal may comprise: generating, using an encoder of the first pre-trained forward model, a representation of the audio signal, and generating, using a decoder of the first pre-trained forward model and the representation of the audio signal, a single prediction of the neural activity in the first set of AL Ref: P46575WO1 14 August 2025 hearing-unimpaired brains. Similarly, generating the single second neural response to the modified audio signal may comprise: generating, using an encoder of the second pre-trained forward model, a representation of the modified audio signal, and generating, using a decoder of the second pre-trained forward model and the representation of the modified audio signal, a prediction of the neural activity for each brain in the second set of hearing-impaired brains.
[0118] In a second example, generating the first representation of a neural response to the audio signal may comprise using a first pair of pre-trained encoders of the ML model; and generating the second representation of a neural response to the modified audio signal may comprise using a second pair of pre-trained encoders of the ML model. Thus, instead of using pre-trained forward models comprising encoders and decoders, in this example, the decoders are ‘reversed’ in direction to act as encoders, such that sound and neural activity are mapped into a shared latent space between the pairs of encoders. See also Figure 2B for a block diagram showing this example.
[0119] In this example, generating the first representation may comprise: generating, using a first encoder of the first pair of pre-trained encoders, a first latent representation of features of the audio signal; generating, using a second encoder of the first pair of the pre-trained encoders, a second latent representation of features of a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing- unimpaired brains, wherein the first and second latent representations are in the same latent space. Thus, generating a first representation comprises minimising a difference between the first and second latent representations. Similarly, generating the second representation may comprise: generating, using a first encoder of the second pair of pre-trained encoders, a third latent representation of features of the modified audio signal; and generating, using a second encoder of the second pair of the pre-trained encoders, a fourth latent representation of features of a single second neural response to the audio signal for each brain of the second set of hearing-impaired brains, wherein the third and fourth latent representations are in the same latent space. Thus, generating a second representation comprises minimising a difference between the third and fourth latent representations.
[0120] In some cases, the method of Figure 3A is for training a machine learning, ML, model to generate less noisy / distorted audio signals. In such cases, the first set of brains may be a set of hearing-unimpaired brains; and the second set of brains may also be a set of hearing- unimpaired brains. The first and second set of brains may be the same or different.
[0121] The first representation may be generated by determining neural responses to the audio signals that are considered “clean” audio signals. That is, these audio signals do not have any distortion or have minimal distortion. The second representation may be generated by determining neural responses to distorted versions of the same clean audio signals. That is, AL Ref: P46575WO1 14 August 2025 the distorted versions contain distortion. Thus, the goal of the training in these cases is to transform the distorted audio signals so that the modified distorted audio signals generate a similar response in the second set of brains as the clean / undistorted audio signals generate in the first set of brains.
[0122] In such cases, the method may further comprise: modifying, prior to transforming the audio signal, the audio signal of the plurality of audio signals to include a distortion or distortions, thereby generating a modified distorted audio signal. That is, the signal that is used to generate a second representation to the modified audio signal in the second set of brains may be additionally pre-altered to include distortions. This is in additional to the modifications which take place to cause the responses in the second set of brains to be similar to the responses in the first set of brains. Any sort of distortion may be added to generate the modified (pre-distorted) audio signal such as, for example, low frequency noise, impulsive noise, intermittent noise, continuous noise, background noise, speech, babble noise, ambient or environmental noise, reverberation, compression, non-noise distortion, and so on.
[0123] In one example, generating the first representation of a neural response to the audio signal may comprise using a first pre-trained forward model of the ML model for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation may comprise generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains. Similarly, generating the second representation of a neural response to the modified audio signal may comprise using a second pre-trained forward model of the ML model for simulating neural responses in the second set of hearing-unimpaired brains, and generating the second representation may comprise generating a second neural response to the modified distorted audio signal for each brain in the second set of hearing-unimpaired brains. It will be understood that the ML model may alternatively comprise pre-trained encoder pairs, as described in relation to the hearing impairment case, and that such an architecture may be used for distortion reduction.
[0124] In some cases, the techniques for distortion cancellation and the techniques for hearing impairments may be combined. Thus, the method in Figure 3A is for training a machine learning, ML, model to generate less distorted audio signals. In such cases, the first set of brains may be a set of hearing-unimpaired brains; and the second set of brains may be a set of hearing-impaired brains.
[0125] In such cases, the method may further comprise: modifying, prior to transforming the audio signal, the audio signal of the plurality of audio signals to include distortion, thereby generating a modified distorted audio signal. AL Ref: P46575WO1 14 August 2025
[0126] In such cases, generating the first representation of a neural response to the audio signal may comprise using a first pre-trained forward model of the ML model for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation may comprise generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains. Similarly, generating the second representation of a neural response to the modified audio signal may comprise using a second pre-trained forward model of the ML model for simulating neural responses in the second set of hearing-impaired brains, and generating the second representation may comprise generating a second neural response to the modified distorted audio signal for each brain in the second set of hearing-impaired brains.
[0127] As noted above, the ML module may comprise first and second pre-trained modules to respectively generate first and second representations of the neural responses for training the ML model. For clarity and to provide relevant information for the description that follows, the training of the pre-trained models will now be described with reference to Figures 4A and 4B.
[0128] Forward models of neural coding. Figure 4A is a block diagram of a process for training the first and second pre-trained modules. The first and second pre-trained modules 106, 108 labelled “ICNet” in Figure 4A (also referred to as the first and second modules herein, for conciseness) are trained using intracranial recordings from silicon probe electrode arrays 200 with hundreds of recording channels 202, as described in Armstrong, A., et al. That is, training datasets for the first 106 and second modules 108 are obtained through large-scale recordings from the IC of hearing-unimpaired and hearing-impaired brains, respectively.
[0129] Figure 4B shows results of the intracranial recording process for hearing-impaired and hearing-unimpaired brains. To observe neural activity patterns at the required scale and resolution (e.g. at high resolution), intracranial recordings are made. This is because optimizing a sound transformation in real-time while these recordings are being made is not feasible. These electrode arrays 200 allow the observation of the timing of action potentials from neurons across the extent of the IC, providing a comprehensive view of the neural code in an individual brain. Such invasive recordings are not feasible in humans; instead gerbils may be used, a common animal model for studies of human hearing.
[0130] To achieve a comprehensive sampling of the neural code from an individual animal, activity may be recorded for several hours (e.g. more than 10 hours) in response to a wide range of sounds including, speech, music, environmental noises, and artificial sounds with unnatural acoustic properties. The multi-unit activity is extracted (spikes from a small group of neurons) for each channel on the electrode array and is represented as spike counts (typically zero or one, but up to four) in 1.3ms time bins. AL Ref: P46575WO1 14 August 2025
[0131] In other words, large-scale, high-resolution recordings of neural activity are made. Figure 4B shows multi-unit activity recorded in the IC during the presentation of tones to hearing-unimpaired (labelled “normal” in Figure 4B) and hearing-impaired brains. Each subplot shows the frequency-response area for one channel (the average activity recorded during the presentation of tones with different frequencies and intensities). The sharp frequency tuning and low intensity thresholds that are characteristic of normal hearing are not present in animals with hearing loss.
[0132] Returning to Figure 4A, the first and second modules may be regarded as forwardmodels ^: ^ ∈ ^^^^ which take input ^, a sound waveform with ^ samples, andproduce output ^̂, a pattern of spike counts across ^ time bins and channels. If these models are sufficiently accurate, then they can be used as simulators in place of the auditory systems that they are trained to mimic. Deep convolutional neural networks ^^may be used withparameters ! optimised so that the simulated neural activity ^^̂ = ^^(^; !) is as close aspossible to the real neural activity ^ recorded from a particular brain ". The problem of neural activity simulation may be framed as a classification problem: given the recent history of sound input, the model is trained using a cross-entropy loss to produce an estimate of the full probability distribution #(^̂|^) of the activity for each channel in each time bin, with classes corresponding to the possible spike count values.
[0133] That is, the first and second pre-trained modules are trained on electrode data from real brains. In particular, the first and second pre-trained modules may be trained on data obtained from a plurality of electrode channels. Accordingly, the first pre-trained module may simulate auditory brain responses in electrode channels measuring auditory brain response in real brains to audio inputs. In this case, generating the first neural response comprises generating a response for each electrode channel in the first plurality of electrode channels.
[0134] Similarly, the second pre-trained module may simulate auditory brain responses produced in a second plurality of electrode channels measuring auditory brain responses in real brains to audio inputs. In this case, generating the second neural response comprises: generating a response for each electrode channel in the second plurality of electrode channels.
[0135] Figure 4C shows comparisons of simulated neural activity for hearing-unimpaired and hearing-impaired brains using the first and second pre-trained modules. The first and second modules may be deep convolutional neural networks, and may be deep encoder-encoder models. In this case, sound-activity pairs are used to train the deep encoder-decoder models to predict the conditional distributions of IC spike counts over time and across units, and activity is simulated by sampling from the predicted distributions. The top two rows of the figure show example sound waveforms and spectrograms. The third and fourth rows show recorded AL Ref: P46575WO1 14 August 2025 and simulated activity for a hearing-unimpaired brain. The fifth and six rows show recorded and simulated activity for hearing-impaired brain. As can be seen from Figure 4C, the first and second pre-trained modules (ICNets) exhibit good performance in predicting activity in both normal and impaired brains across a wide range of sound classes.
[0136] As noted above, step S108 of Figure 3A for generating the aligned second neural response may be performed in multiple alternative ways. This will now be described in more detail below in relation to Figures 5A to 5C.
[0137] Alignment through projection matrix. Figure 5A is a flowchart of steps in a first example process for generating the aligned second neural response using the alignment module. As noted above, the present techniques require aligning the activity from the normal (i.e. hearing unimpaired) and hearing-impaired brains. If there are first ^^^and second ^^^pre-trained modules (i.e. a pair of ICNets) trained on hearing-unimpaired and hearing-impaired brains, then these models can be used to identify the sound transformation that best corrects the distortions in the impaired brain.
[0138] It is impossible to align activity for normal and impaired brains based on functional properties. In normal brains, neurons in the IC are organised tonotopically according to their preferred sound frequency, with neurons that prefer low frequencies located dorsolaterally and neurons that prefer high frequencies located ventromedially. This tonotopy is well preserved across individuals and could, in principle, be used to align activity from two normal brains by matching channels with the same preferred frequency. But hearing loss causes changes in preferred frequency, so when aligning normal and impaired brains based on preferred frequency, it becomes unclear which channels should be matched.
[0139] A simple channel-by-channel comparison of activity is not sufficient, as activity recorded on the same electrode channel in different brains will differ for reasons that have nothing to do with hearing impairment. While the overall structure of auditory brain areas is similar across brains, there is substantial variation at the scale of individual neurons. Neurons at the same location in different brains may respond similarly to some sound features but differently to others. And it is impossible to guarantee that recording electrodes are placed in the same exact location in different brains.
[0140] There are existing approaches to aligning recordings of neural activity from different brains (e.g., canonical correlation analysis), but these approaches lack the appropriate constraints, i.e., they will correct not only differences related to recording idiosyncrasies but also those related to hearing loss. Thus, a key challenge in using the neural code as a basis for hearing aid optimization is solving this alignment problem, i.e., finding a method for aligning activity patterns from two brains that removes differences related to recording idiosyncrasies but preserves differences that are a result of hearing impairment. AL Ref: P46575WO1 14 August 2025
[0141] To ensure that the present techniques did not correct differences related to hearing loss, channel rearrangement via optimal transport can be used. In other words, the alignment is implemented through a projection matrix %^→&that is applied to the original activity to producethe aligned activity ^^̂→& = ^^̂%^→&, with %^→& optimised to provide the best alignment of ^^̂ to^&̂ and learned jointly with the parameters ^ of the sound transformation ℎ.
[0142] The projection matrix % may be constrained to be a permutation matrix (i.e. an orthogonal matrix with each column containing a single one and zeros otherwise, as noted above) so that each channel is used exactly once. This constraint results in a very conservative form of alignment in which the actual values of the activity patterns remain unchanged, but the ordering of channels is rearranged such that the difference between corresponding pairs of channels is minimised. Because the alignment is learned jointly with the sound transformation, it is continuously updated along with the sound transformation to always produce the best match between the activity from the two brains.
[0143] In other words, generating an aligned second neural response may comprise: applying, using the alignment module, a projection matrix to the second neural response, where values of the projection matrix are optimised to align first and second neural responses (step S300). Applying a projection matrix to the second neural response may comprise: constraining the projection matrix to be a permutation matrix having a plurality of columns, where each column has a plurality of zero elements and a single non-zero element (step S302). In this case, training the alignment module may comprise adjusting a position of the non-zero elements in each column of the permutation matrix (S304).
[0144] Alignment through Maximum Covariance Analysis. Figure 5B is a flowchart of steps in a second example process for generating the aligned second neural response using the alignment module. A common approach to aligning two datasets is to project them into a latent space of reduced dimensionality in which their similarity is maximised. This is most commonly achieved in neuroscience (and beyond) through canonical correlation analysis (CCA), which finds the linear projection of each dataset that maximises their correlation after projection. To align normal and impaired responses ^^̂^and ^^̂^, CCA would identify a pair of weight matrices %^^→^^ that align the responses within a new space: ^^̂^→^^=^^̂^%^^→^^ and ^^̂^→^^ = ^^̂^%^^→^^. The first pair of columns in %^^→^^ and %^^→^^ containthe weights that maximise the Pearson correlation coefficient of the first dimension of the aligned responses and ^^̂^→^^. Each successive pair of columns in %^^→^^and %^^→^^contain the weights that maximise the correlation of the corresponding dimension of the aligned responses, subject to the constraint that each new dimension of the aligned responses is orthogonal to its existing dimensions. AL Ref: P46575WO1 14 August 2025
[0145] CCA is unsuitable for the alignment problem here because it corrects differences in scale (or, put differently, any measure of the similarity of CCA-aligned responses would be insensitive to changes in the scale of the original responses). One of the most obvious effects of hearing impairment is to reduce the overall level of neural activity. If ^^̂^were a scaled down version of ^^̂^, CCA would correct the difference in scaling and align them perfectly. Thus, thedifference ^(^^̂^→^^, ^^̂^→^^) would be zero and an sound transformation module trained tominimise the difference between ^^̂^and ^^̂^after alignment by CCA would learn to do nothing. But CCA is only one specific form of the latent space approach to alignment; there may be others that are more suitable for solving the present problem.
[0146] One possibility is Maximum Covariance Analysis (MCA), which is similar to CCA except that instead of maximizing the correlation between aligned datasets, it maximises their covariance. MCA has been around for decades but it is not widely used (in neuroscience or beyond). It is simple to implement. First the cross-covariance matrix between two datasets iscomputed, in our case '^^,^^ = ^^̂ ^^ ^^̂^ / ^ − 1, where ^ is the number of time bins in theresponse. Next, the weights for the projection of each response into the new space are obtained by singular value decomposition '^^,^^ are the weight matrices and Σ is a diagonal matrix, the entries of which reflect the magnitude of the covariance between the responses along each dimension in the new space.
[0147] As with CCA, the weight matrices are used to produce responses that are aligned withinthe new space: ^^̂^→^^ = and ^^̂^→^^ = ^^̂^%^^→^^. The first pair of columns in contain the weights that maximise the covariance of the first dimension of the aligned responses and ^^̂^→^^. Each successive pair of columns in %^^→^^and contain the weights that maximise the covariance of the corresponding dimension of the aligned responses. Unlike CCA, the constraint that the alignment must satisfy is not on the aligned responses but on the weight matrices themselves, which must be orthonormal.
[0148] If the simple example is considered where ^^̂^is a scaled down version of ^^̂^, it can be seen that the properties of MCA are much more suitable than those of CCA. Because the MCA weight matrices are orthonormal the alignment is scale preserving. (A more subtle point is also that because the measure that the alignment is optimizing, covariance, is sensitive to scale, the weights will reflect the difference in scale, which is desirable; so the difference between MCA and CCA is not simply the preservation of scale in the new space but also the sensitivity to scale differences in the optimization of the weights.)
[0149] While the alignment achieved by MCA seems suitable for the present techniques, the metric that is optimised, covariance, cannot be used directly for the optimization of the sound transformation module. A simple way to increase the covariance between two datasets is to scale them up, so the sound transformation module that is trained to maximise the covariance AL Ref: P46575WO1 14 August 2025 of responses aligned via MCA (or, for that matter, any other method), would simply learn to modify incoming sounds in order to make the impaired responses as large as possible.Instead, we must use a difference measure with the appropriate sensitivity.There are many possibilities, including simple measures of differences between the aligned responses in the new space (e.g., MAE or MSE), or measures of differences between one of the original responses and the reconstruction of that response from the latent representation of the other response. For example, a reconstruction of ^^̂^could be obtained by putting ^^̂^→^^through the inverse projection defined by the weights %^^→^^, and the difference between ^^̂^and this reconstruction could be measured via MAE and MSE. Each of these difference measures (and the countless variations on them) will emphasize different features but, because we do not know exactly how hearing impairment distorts neural activity or which of these distortions are most important for perception, it is difficult to know in advance which difference measure will lead to the best sound transformation performance.
[0150] That is, generating an aligned second neural response may comprise: decomposing, using the alignment module, a covariance matrix into a first weight matrix and a second weight matrix, wherein the covariance matrix corresponds to a covariance of the first and second neural responses (step S400); and applying the first and second weight matrices to the second neural response (step S402). That is, a statistical measure of the difference between the first and second neural responses, i.e. the covariance matrix, may be used to generate the aligned second neural response. The covariance matrix may be decomposed into weight matrices which encode these differences, and applied to their respective neural responses. That is in MCA, rather than the second neural response being aligned to the first neural response, both the first and second neural responses are effectively aligned to each other. To do so, the first weight matrix is applied to the first neural response, and the second weight matrix is applied to the second neural response. In this case, training the alignment module may comprise adjusting values of elements in the first and second weight matrices (step S404). In other words, values of elements in the first and second weight matrices may be iteratively adjusted, in a similar way to the example of the permutation matrix above described in relation to Figure 5A.
[0151] Alignment through affine transformation. Figure 5C is a flowchart of steps in a third example process for generating the aligned second neural response using the alignment module. In the approach described above in relation to Figure 5A, the alignment is implemented through a projection matrix %^^→^^that is applied to the original activity toproduce the aligned activity ^^̂^→^^ = constrained to be a permutationmatrix. The conservative nature of this approach is a strength because it ensures that the alignment does not correct differences in activity due to hearing loss, but it is also a weakness AL Ref: P46575WO1 14 August 2025 because it prevents the alignment from correcting for other differences that are not due to hearing loss. Aside from hearing loss, the main reason for the discrepancy in activity recorded in different brains is the difference in the position of the recording electrode. This difference in position can be well approximated by a two-dimensional rigid affine transformation, as schematically illustrated in Figure 5D, with three parameters -- one for horizontal translation, one for vertical translation, and one for rotation – and the optimal values of the parameters for aligning activity from any two brains can be learned along with the optimal sound transformation during AidNet training.
[0152] When viewed this way, the alignment problem is similar to an image registration problem (or, more accurately, a video registration problem with each time bin of neural activity corresponding to a frame). We assume that the electrode position in each brain is fixed throughout time so that the same alignment transformation can be applied to all time bins. Forany electrode channel position {", -} in the second brain, the corresponding position {"′, -′} inthe first brain is given by: "1 cos α -sin α 3^ "0-12 = 0 sin α cos α 3& 25-6 (1)1 0 0 1 1where the parameters 7, 3^, and 3&control the rotation, horizontal translation, and vertical translation, respectively.
[0153] Without any constraints on the parameter values, the position {"′, -′} is likely to fallbetween the locations of electrode channels in the first brain. A simple solution to this problemis to choose the closest channel to {"′, -′} in the first brain, but a better solution is to interpolatebetween channels, which can be achieved by applying a sampling kernel. If we are aligning an impaired brain to a normal brain, we obtain the aligned activity for channel 8 located at by applying a sampling kernel centered at {"′9, -′9 where the sum is taken over all electrode channels, {"^, -^} is the location of channel ?, and; is any sampling kernel that is differentiable with respect to "1and -1. (Note that this approach is similar to using a Spatial Transformer without the localization network.) Once the alignedactivity is obtained for all channels, the difference ^(^^̂^, ^^̂∗^→^^ ) can be computed acrosscorresponding channel pairs.
[0154] Thus, in this third example of generating the aligned second neural response, it is necessary to generate a response in electrode channels to the audio signals. Accordingly, the first pre-trained module simulates auditory brain responses produced in a first plurality of AL Ref: P46575WO1 14 August 2025 electrode channels measuring auditory brain responses in real brains to audio inputs, and generating the first neural response comprises: generating a response for each electrode channel in the first plurality of electrode channels to the audio signal (step S500). Similarly, the second pre-trained module simulates auditory brain responses produced in a second plurality of electrode channels measuring auditory brain responses in real brains to audio inputs, and generating the second neural response comprises: generating a response for each electrode channel in the second plurality of electrode channels to the modified audio signal (step S502).
[0155] Furthermore, in this third example, generating an aligned second neural response comprises: applying, using the alignment module, a transformation matrix to the second neural response, wherein the transformation matrix translates and / or rotates electrode channel positions in the second neural response (step S504). In this case, each channel in the second plurality of electrode channels may have a neural position in the hearing-impaired brain and at step S504, applying the transformation matrix may comprise transforming, using the translation and / or rotation of the transformation matrix, the neural positions of each electrode channel to generate a plurality of transformed electrode channel positions. Generating an aligned second neural response may comprise a further step S50 of modifying the response in each electrode channel of the second plurality of electrode channels to generate an aligned response in each electrode channel.
[0156] At step S50, modifying the response in each electrode channel may comprise: for each response generated in the second plurality of electrode channels: applying a sampling kernel to the response to generate an electrode channel contribution, wherein the sampling kernel is based on the transformed electrode channel position and a corresponding position of an equivalent electrode channel in the first plurality of electrode channels; and summing the electrode channel contributions to generate the aligned response in the electrode channel.
[0157] Further at step S50, minimising a loss function based on a difference between the aligned second neural response to the modified audio signal and the first neural response to the audio signal may comprise: minimising a difference between the aligned second neural response and the first neural response in corresponding electrode channels by adjusting translation and / or rotation parameters of the transformation matrix.
[0158] Further examples will now be described, directed at improving the generalizability and enhancing the personalisation of the present techniques.
[0159] Multi-brain models. Figure 6A is a block diagram of a process to train a first pre-trained module which simulates auditory brain responses in a plurality of hearing-unimpaired brains. Rather than steering impaired activity toward the specific patterns recorded in any one normal brain as described above, a training process that minimises the difference between impaired AL Ref: P46575WO1 14 August 2025 activity and a generic normal neural code can instead be used. A simple approach would be to train the ML model to minimise the difference between the activity from a hearing-impaired brain and the activity from a collection of normal brains by using a different normal brain for comparison during each iteration of the training loop. However, a more effective approach is to learn a generic normal neural representation that reflects the common features of neural activity across many normal brains. This is an alternative approach to the training process described in relation to Figures 4A to 4C above for the (pre-) training of the first and second pre-trained modules, which learn to simulate activity in a single hearing-impaired / unimpaired brain.
[0160] To achieve this, a first pre-trained module in the form of a multi-branch forward model300 A: ^ ∈ ^^^^ → ^ ∈ ^ ≥ 0^^^^C may be trained that takes a sound waveform andproduces patterns of spike counts for D different brains. The sound may be processed by a single deep convolutional neural network ^^Eto create a latent representation that then undergoes individualised alignments to produce the final output activity patterns.
[0161] The alignment, as always, is constrained such that it cannot correct differences between brains due to hearing loss. (While this constraint may not be important when training this model of hearing-unimpaired brains, it becomes important in the next step when training AidNet.) Because of the constraints on the alignment, the latent representation is not an abstract feature representation but rather a generic conditional distribution of neural activity that can be transformed to predict the activity from any individual brain via alignment.
[0162] In other words, the first pre-trained module may simulate auditory brain responses of a plurality of real hearing-unimpaired brains to the audio signal. In this case, generating a first neural response comprises generating a first neural response of each hearing-unimpaired brain of the plurality of hearing-unimpaired brains to the audio signal. That is, the response of each hearing-unimpaired brain may be simulated.
[0163] Figure 6B is a block diagram of a process for training the ML model using the first pre- trained module of Figure 6A. A second pre-trained module 106 may be trained as described in relation to Figure 4A and the two modules may be incorporated, ^^Eand ^^^, into a process for training of the sound transformation module 102. The sound transformation module 102 may be trained jointly with the alignment module %^^→^Ethat finds the best alignment between the impaired activity and the generic normal neural activity with an overall objective ofminimising the difference between them ^(^^̂E, ^^̂∗^→^E ). As mentioned above, in the casewhere the first pre-trained module simulates auditory brain responses for a plurality of real hearing-unimpaired brains to the audio signal, the first pre-trained module has learned how to generate a single latent representation (the “generic representation” in Figure 6A) that is sufficient for generating all of the responses for the multiple brains, but does not need to AL Ref: P46575WO1 14 August 2025 actually generate the individual responses themselves. The single latent representation is then used to provide the first, single neural response for the multiple brains. The same also applies for the second pre-trained module that simulates auditory brain responses for a plurality of hearing-impaired brains to the modified audio signal.
[0164] Toward personalization. Using the processes described above, a sound transformation module can be trained to learn the optimal sound transformation for any impaired brain from which we have neural recordings. For human users, there are no such recordings; instead, the optimal sound transformation must be identified indirectly. One simple solution would be to find the closest match to the audiogram (the result of the standard clinical hearing test) of a new user from the collection of audiograms corresponding to impaired brains on which sound transformation modules have been trained. A better solution would be to incorporate a parameterization of hearing loss, F, and train the sound transformation module to modify its sound transformation according to these parameters. One option would be to use the results of a hearing test (e.g., an audiogram) to determine F. However, no hearing test can capture the effects of hearing impairment as fully as large-scale neural recordings, so a better option is to learn F for each brain when training the second pre-trained module. With such a framework, optimization for any new user would entail searching the learned F space for the optimal values, using, for example, human-in-the-loop optimization (Granley, J., et al.).
[0165] Figure 7A is a block diagram of a process to train a second pre-trained module which simulates auditory brain responses in hearing-impaired brains. For the personalisation approach mentioned above, a second pre-trained module in the form of a multi-branch modelfor hearing loss 400 AG: ^ ∈ ^^^^ → ^ ∈ ^ ≥ 0^^^^C may be trained. This is similar to the multi-branch model 300 described above for hearing-unimpaired brains, but with additional parameters F 402 that are learned for each brain.
[0166] An encoder and decoder are split into a generic encoder 404 ^EHand a generic decoder 406 IEH, both of which are shared across animals. The splitting is necessary so that the parameters 402 F for each animal can be included as input to the decoder 406 along with the generic latent representation of the sound, JEH(^), that is produced by the generic encoder 404. Since the decoders receive both the generic sound representation and the parameters F, they can use the parameters as needed to produce activity that reflects the impairment in each individual brain.
[0167] The output of each decoder is individualised with respect to impairment, but is still generic with respect to the other idiosyncrasies of each brain. So, as for the multi-branch normal hearing model described above, the output of each decoder is passed through an alignment transformation %EH→^to produce the final activity patterns for each brain. For the learned space of parameters F to provide a suitable basis for personalization, it is important AL Ref: P46575WO1 14 August 2025 to ensure that its entire domain is valid, i.e., it must be ensured that the model performs as desired not only for the specific values of F corresponding the brains on which the model is trained, but also for all values in between. One way to achieve this is to introduce a variational framework so that instead of directly learning the optimal F^for each brain ", the model learns the mean and variance of a normal distribution from which F^is obtained by sampling. This results in each decoder producing neural activity that varies smoothly with respect to F within its local domain, and with suitable regularization, this smoothness can be extended over the whole domain of F. It may also be important to impose constraints on the relationship between F and other network parameters to ensure that the input {JEH(^), F} to the decoder remains disentangled, i.e., to ensure that the components reflect only the generic neural representation of sound and individualised impairment, respectively.
[0168] Figure 7B is a block diagram of a process for training the ML model using the second pre-trained module of Figure 7A. Once the multi-branch model 400 for impaired brains has been trained, a sound transformation module 102 can be trained to minimise the difference^(^^̂E , ^Ê∗H→^E ) between the generic normal neural activity produced by the multi-branch model300 for normal brains from above and the activity of the impaired model with varying values of F, where the alignment %EH→^Ebetween the generic normal activity and the generic impaired activity is again jointly learned with the sound transformation module 102. The impairment parameters 402 are provided to the sound transformation module 102 as well as to the decoder 406 of the impaired forward model so that the sound transformation performed by the sound transformation module is matched to the needs of any specific impairment. As with the multi-branch impaired model, the introduction of variational framework into the sound transformation module 102 along with suitable regularization can ensure that the sound transformation module 102 performs as desired across the entire domain of F and, thus, is appropriate for any user.
[0169] That is, the method for training the ML model may further comprise personalising the training for a specific hearing-impaired person by: generating, using an encoder of the second pre-trained module, a latent representation of the audio signal; and transforming, using a personalisation parameter for each hearing-impaired brain and a decoder of the second pre- trained ML model, the latent representation to generate a neural response for the specific hearing-impaired person, wherein the personalisation parameter corresponds to a quantification of the effects of the hearing impairment in the brain of the specific hearing- impaired person.
[0170] Returning to Figure 2B, which shows an alternative process for training an ML model, a full explanation of this process is now described. First, as shown in image (i) of Figure 2B, thedecoders in the forward models that previously mapped sound to neural activity, ^: ^ → ^,̂ are AL Ref: P46575WO1 14 August 2025 reversed in direction to act as encoders so that sound and neural activity are mapped into ashared latent space with the two encoders, JK: ^ → LK and JM: ^ → LM, trained jointly to minimisethe difference ^(LK, LM) between the latent representations of the sound and neural activity.
[0171] If these encoders perform well, the latent representations LKand LMwill be equivalent and can be used interchangeably, as shown in image (ii) of Figure 2B. As with the original forward models, a multi-brain system can be built and trained to learn a generic normal latent representation using activity from a collection of normal brains. In this case, a set of encoders are optimised jointly such that the differences between the latent representationsof the sound and neural activity across all brains ∑CPQ^ ^(LK, LMP) is minimised.
[0172] The considerations for the sound encoder JKin this alternative system are the same as for the original forward models that mapped sound to neural activity. If the activity encoders{JM^, … , JMC} are constrained to be orthonormal linear projections, then the mapping from activityto the latent space will have similar properties to alignment via MCA as described above, i.e., the mapping will align the activity from different brains within a shared latent space without risk of correcting differences in activity that are caused by hearing loss (which, again, is not important when training a system with normal brains, but becomes important when training the ML model.)
[0173] After this normal multi-brain system is trained, the sound encoder JKis frozen and can be used to generate the generic normal latent representation for any sound. In the next step (shown as image (iii) in Figure 2B), a parametric form of hearing impairment is incorporated (described in more detail below).
[0174] In this system, hearing impairment is imposed on the generic normal latent representation LKto create a modified latent representation RK. With data from a collection ofhearing impaired brains, a set of encoders {JM^, … , JMC} are optimised jointly with the impairmentparameters F for each brain such that the differences between the pairs of latentrepresentations of the sound and neural activity across all brains ∑CPQ^ ^(RKP, LMP) is minimised.The activity encoders {J^M, … , JCM } are again constrained to be orthonormal linear projections sothat the activity from different brains is aligned in the latent space without correcting hearing impairment, while the effects of hearing impairment are imposed on the generic normal latent representation as needed to match the latent representation of the impaired activity from each brain.
[0175] Finally, with the impairment parameters from this system and the same frozen sound encoder, the ML model can be trained (see image (iv) in Figure 2B). In one branch of the training system, sound is mapped to a generic normal latent representation LK. In the other branch, sound is first modified by the sound transformation module (also referred to herein as “AidNet”) to produce ^∗, then mapped to a generic normal latent representation LK∗, and finally AL Ref: P46575WO1 14 August 2025 mapped to the impaired latent representation for a particular brain RKP∗ using the impairment parameters FPfor that brain. AidNet is trained to minimise the difference between the twolatent representations ^(L PK, RK∗ ).
[0176] Note that neural activity is never generated or even used in this AidNet training system. Because the two previous steps ensure that the sound and activity encoders produce latent representations that are equivalent (or as close to equivalent as possible), AidNet can be optimised using only the latent representations produced by the sound encoder.
[0177] The AidNet training system can also be extended to multiple brains, as shown in image (v) of Figure 2B. The impairment parameters F are provided to the AidNet so that the sound transformation performed by AidNet is matched to the needs of any specific impairment. In this system, the generic normal latent representation LKserves as the target and a single AidNet is trained to process sound based on the impairment parameters for each brain in order to restore all of the impaired latent representations to normal, i.e., so that the difference minimised.
[0179] Figure 8 is a flowchart of steps in a process for using a trained ML model for generating modified audio signals. That is, Figure 8 describes the process for using the ML model, trained according to the present techniques, to generate modified audio signals for a hearing-impaired person. The process is a computer-implemented method, and is implemented on a user device. Typically, the user device is a resource-constrained device but with the minimum set of hardware capabilities required to use a trained ML model for generating modified audio signals. The method comprises: receiving an input audio signal (step S600); generating, using a sound transformation module of the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the hearing-impaired person that is similar to an auditory brain response to the input audio signal in a hearing-unimpaired brain (step S602); and outputting the modified audio signal (step S604).
[0180] Figure 9 is a block diagram of a user device for generating modified audio signals for a hearing-impaired person using a trained ML model. The user device 500 comprises: storage 502 for storing the trained ML model 504; and a processor 508 coupled to memory 510, for: receiving an input audio signal; generating, using the trained ML model 504, a modified audio signal for generating an auditory brain response in the brain of the hearing-impaired person that is similar to an auditory brain response to the input audio signal in a hearing-unimpaired person; and outputting the modified audio signal.
[0181] The user device may be any of: a hearing aid; a cochlear implant; a pair of headphones; earphones; a smartphone; a tablet; a headset; a smart speaker; and a speaker. In each case, the user device has the minimum hardware capabilities to use a trained ML model. AL Ref: P46575WO1 14 August 2025
[0182] The user device may further comprise a speaker 506 for outputting the modified audio signal.
[0183] Alternatively, the user device may further comprise electronics for outputting electrical stimulation based on the modified audio signal. For example, a cochlear implant generates electrical shocks or impulses, and therefore, the modified audio signal may be used to alter the electrical shocks / impulses that are generated and applied to stimulate the cochlear nerve.
[0184] The user device may further comprise a microphone or audio receiver (not shown) for obtaining the input audio signal which is to be modified.
[0185] A series of experiments conducted to validate the present techniques will now be described.
[0186] Experimental protocol. Experiments were performed on young-adult gerbils of both sexes that were born and raised in standard laboratory conditions. Hearing impairment was induced through exposure to noise when at an age of 16-18 weeks. IC recordings were made at an age of 20-24 weeks. All experimental protocols were approved by the UK Home Office (PPL P56840C21).
[0187] Noise exposure. Mild-to-moderate sensorineural hearing loss was induced through exposure to high-pass filtered noise with a 3 dB / octave roll-off below 2 kHz at 118 dB SPL for 3 hours. For anaesthesia, an initial injection of 0.2 ml per 100 g body weight was given with fentanyl (0.05 mg per ml), medetomidine (1 mg per ml), and midazolam (5 mg per ml) in a ratio of 4:1:10. A supplemental injection of approximately 1 / 3 of the initial dose was given after 90 minutes. Internal temperature was monitored and maintained at 38.7° C.
[0188] Preparation for large-scale IC recording. Animals were placed in a sound-attenuated chamber and anesthetised for surgery with an initial injection of 1 ml per 100 g body weight of ketamine (100 mg per ml), xylazine (20 mg per ml), and saline in a ratio of 5:1:19. The same solution was infused continuously during recording at a rate of approximately 2.2 μl per min. Internal temperature was monitored and maintained at 38.7° C. A small metal rod was mounted on the skull and used to secure the head of the animal in a stereotaxic device. The pinnae were removed and speakers (Etymotic ER-10X) coupled to tubes were inserted into both ear canals. Two craniotomies were made along with incisions in the dura mater, and a 256-channel multi-electrode array was inserted into the central nucleus of the IC in each hemisphere.
[0189] Multi-unit activity. Multi-unit activity (MUA) was measured from recordings on each channel of the electrode array as follows: (1) a bandpass filter was applied with cutoff frequencies of 700 and 5000 Hz; (2) the standard deviation of the background noise in the bandpass-filtered signal was estimated as the median absolute deviation / 0.6745 (this estimate is more robust to outlier values, e.g., neural spikes, than direct calculation); (3) times AL Ref: P46575WO1 14 August 2025 at which the bandpass filtered signal made a positive crossing of a threshold of 3.5 standard deviations were identified and grouped into bins with a width of 1.3 ms.
[0190] Characteristic frequency analysis. To determine the preferred frequency of the recorded units, 50 ms tones were presented with frequencies ranging from 300 Hz to 16000 Hz in 0.2 octave steps and intensities ranging from 4 dB SPL to 85 dB SPL in 9 dB steps with 10 ms cosine on and off ramps and 75 ms pause between tones. Tones were presented 8 times each in random order with a sampling rate of 44.1 kHz. The characteristic frequency (CF) of each unit was defined as the frequency that elicited a significant response at the lowest intensity (significance was defined as p < 0.0001 for the total spike count elicited by a tone given a unit’s distribution of spike counts for a time window of the same duration in the absence of sound).
[0191] Mapping sound to neural activity (ICNet). DNN models of neural coding, ICNets, were trained to map sound input ^ to an estimated conditional distribution of spike counts #(^). The architecture that was used comprised: (1) a SincNet layer (Ravanelli, M. et al.) with 48 bandpass filters of size 64 and stride 1, followed by a symmetric logarithmic activation - =^I?(") ∗ STI(|"| + 1); (2) a stack of five 1-D convolutional layers with 128 filters of size 64and stride 2, each followed by a PReLU activation; (3) a 1-D bottleneck convolutional layer with 64 filters of size 64 and stride 1, followed by a PReLU activation; (4) a cropping layer to eliminate convolutional edge effects as described below; and (5) a 1-D decoder convolutional layer with 512 ×Wfilters of size 1 and stride 1, with no bias term, followed by a softmax activation, whereWis the number of possible spike counts. All convolutional layers in the encoder included a bias term and used a causal kernel.
[0192] ICNets were trained to transform 24,414.0625 Hz sound input frames of 10,240 samples into 762.9395 Hz neural activity frames of 256 samples. Context of 2048 samples was added on the left side of the sound input frame and was cropped after the bottleneck layer (2048 divided by a decimation factor of 2Xresulted in 64 cropped samples after the bottleneck). Sound inputs were scaled such that an RMS of 0.04 corresponded to a level of 94 dB SPL. The number of classesWin the decoder was set to 5, as the percentage of spike counts above 4 across all datasets was less than 0.02%. During training, a cross-entropy loss was used to estimate the probabilities of each class for each time bin and channel (256 × 512 categorical distributions withWdiscrete classes corresponding to the probability of each unitfiring 0 to W − 1 spikes within a given time bin). All MUA data were clipped during trainingand inference to maximum values of 4.
[0193] ICNets were trained on NVidia RTX 4090 GPUs using Python and Tensorflow. A batch size of 50 was used with the Adam optimizer and a starting learning rate of 0.0004. The learning rate was halved if the loss in the validation set did not decrease for 2 consecutive AL Ref: P46575WO1 14 August 2025 epochs. Early stopping was used and determined the total number of training epochs if the validation loss did not decrease for 5 consecutive epochs.
[0194] When evaluating the trained ICNets, the decoder parameters were reshaped to form a categorical probability distribution withWclasses using the Tensorflow Probability Toolbox. The distribution was then sampled to yield simulated neural activity across the time bins and neural units.
[0195] Training dataset. All sounds that were used for training the ICNets had a total duration of 7.83 hr and are described below.10% of the sounds were randomly chosen and formed the validation set during training. ^ Speech: Sentences were taken from the TIMIT corpus (Garofolo, John S. et al.) that contains speech read by a wide range of US English talkers. The entire corpus excluding ‘SA’ sentences was used and was presented either in quiet or with background noise. The intensity for each sentence was chosen at random from 45, 55, 65, 75, 85, or 95 dB SPL. The speech-to-noise ratio (SNR) was chosen at random from either 0 or 10 when the speech intensity was 55 or 65 dB SPL (as is typical of a quiet setting such as a home or an office) or –10 or 0 when the speech intensity was 75 or 85 dB SPL (as is typical of a noisy setting such as a pub). The total duration of speech in quiet was 1.25 hr and the total duration of speech in noise was 1.58 hr. ^ Noise: Background noise sounds were taken from the Microsoft Scalable Noisy Speech Dataset (Reddy, C.K. et al.), which includes recordings of environmental sounds from a large number of different settings (e.g., café, office, roadside) and specific noises (e.g., washer-dryer, copy machine, public address announcements). The intensity of the noise presented with each sentence was determined by the intensity of the speech and the SNR as described above. ^ Augmented speech: Speech from the TIMIT corpus was also processed in several ways: (1) the speed was increased by a factor of 2 (via simple resampling without pitch correction of any other additional processing); (2) linear multi-channel amplification was applied, with channels centered at 0.5, 1, 2, 4, and 8 kHz and gains of 3, 10, 17, 22, and 25 dB SPL, respectively; or (3) the speed was increased and linear amplification was applied. The total duration of augmented speech was 2.66 hr. ^ Music: Pop music was taken from the musdb18 dataset (Rafii, Z., et al.), which contains music in full mixed form as well as in stem form with isolated tracks for drums, bass, vocals and other (e.g., guitar, keyboard). The total duration of the music presented from this dataset was 1.28 hr. Classical music was taken from the musopen dataset (including piano, violin and orchestral pieces) and was presented either in its original form; after its speed was increased by a factor of 2 or 3; after it was high-pass filtered with a cutoff AL Ref: P46575WO1 14 August 2025 frequency of 6 kHz; or after its speed was increased and it was high-pass filtered. The total duration of the music presented from this dataset was 0.45 hr. ^ Ripples: Dynamic moving ripple (DMR) sounds were created by modulating a series of sustained sinusoids to achieve a desired distribution of instantaneous amplitude and frequency modulations (Depireux, D. A., et al.). The lowest frequency sinusoid was either 300 Hz, 4.7 kHz, or 6.2 kHz. The highest frequency sinusoid was always 10.8 kHz. The series contained sinusoids at frequencies between the lowest and the highest in logarithmic steps of 0.02 octaves with the phase of each sinusoid chosen randomly from between 0 and 2π. The modulation envelope was designed so that the instantaneous frequency modulations ranged from 0 to 4 cycles / octave, the instantaneous amplitude modulations ranged from 0 to 10 Hz, and the modulation depth was 50 dB. The total duration of the ripples was 0.22 hr.
[0196] Evaluation dataset. To evaluate ICNets, sounds that were not part of the training dataset were solely used. All sounds had at least two consecutive trials to estimate their test-retest variability. The total duration of each sound segment for the evaluation was 30 s. ^ Speech in quiet: A speech segment from the UCL SCRIBE dataset (phon.ucl.ac.uk / resource / scribe) consisting of sentences spoken by a male talker was presented at 60 dB SPL. ^ Speech in noise: A speech segment from the UCL SCRIBE dataset consisting of sentences spoken by a female talker was presented at 85 dB SPL in hallway noise from the Microsoft Scalable Noisy Speech dataset at 0 dB SNR. ^ Ripples: Dynamic moving ripples with a lowest frequency sinusoid of 300 Hz was presented at 60 dB SPL. ^ Music: Three seconds from each of 10 mixed pop songs from the musdb18 dataset were presented at 75 dB SPL. ^ Violin: A solo violin recording from the musopen dataset was presented at 85 dB SPL.
[0197] Mapping sound to sound (AidNet). The hearing aid model, AidNet, comprised a symmetric encoder-decoder DNN architecture based on the Wave-U-Net model (Stoller, D., et al.) with 18 convolutional layers in total. The encoder consisted of 91-D convolutional layers of size 12 and stride 2, with [32, 32, 32, 64, 64, 64, 128, 128, 128] filters and PReLU activations between them. The decoder layers mirrored the layers of the encoder, with skip connections between each encoder layer and the corresponding decoder layer. All convolutional layers included a bias term and used a causal kernel.
[0198] AidNets were trained to process 24,414.0625 Hz sound input frames of 65,536 samples. Context of 16384 samples was added on the left side of the sound input frame and was cropped at the output of the forward models (16384 divided by a decimation factor of 2X AL Ref: P46575WO1 14 August 2025 resulted in 512 cropped MUA samples). Sound inputs were scaled such that an RMS of 0.04 corresponded to a level of 94 dB SPL.
[0199] AidNets were trained to restore impaired neural activity to normal as follows: (1) the earth mover's distance algorithm was used to derive the transformation that aligned the impaired activity to the normal activity ^^̂^, comprising the optimal channel reordering that minimised the Wasserstein distance between YW[^^̂^] and YW[^^̂∗^ ], the expected value of the normal and impaired activity, where YWdenotes expectation over spike counts; (2) the aligned impaired activity ^^̂∗^→^^was computed by applying the optimal channel reordering to the impaired activity is a permutation matrix with a unique mapping of each channel of the impaired activity to the normal activity ^^̂^; (3) a loss comprising several terms was computed as described below and was used to train the AidNet via backpropagation. The alignment in step 1 was performed using the Optimal Transport toolbox in Python.
[0200] Hearing aid optimization. Once the activity from the normal and impaired brains is aligned, a straightforward difference measure such as Euclidean distance can be applied. But inasmuch as even a truly optimal AidNet will not be able to correct all of the distortions in the neural code, it is important that the difference measure encourages the correction of those distortions that matter most. Neural activity patterns are complex and it is not clear which of their features are most important for perception, so it is also not clear how to ensure that the difference measure is providing the right emphasis.
[0201] Given this uncertainty, a loss function was chosen that combined several metrics including the MSE and the correlation of the mean activity patterns YW[^̂|^], where YWdenotes expectation over possible spike counts, and the KL divergence of the predicted spike count distributions #(^̂|^). A term that was focused on the neural coding of speech may also be added. As noted above, an automatic speech recognition (ASR) network may be used as a backend to identify phonemes from the mean impaired activity patterns YW[^^̂∗^→^^]. The ASR backend was trained jointly with the alignment and sound transformation and its performance was included in the overall loss function, as described in more detail below.
[0202] The loss term ^Mcombined several metrics to assess different aspects of the similarity between the normal and impaired activity: (1) the KL divergence between #(^^̂^|^) and #(^^̂^→^^|^∗); (2) the MSE between YW[^^̂^|^] and (3) the correlation between YW[^^̂^|^] and YW[^^̂^→^^|^∗]. The three individual loss components were multiplied with weighting factors of 0.05, 0.1 and 0.1, respectively, to achieve approximately equal contribution.
[0203] An additional loss term ^K(^, ^∗) was included to minimise differences between theunprocessed and processed sound and comprised: (1) a spectral loss %\Y(|\[;]|, |\∗[;]|), AL Ref: P46575WO1 14 August 2025 where \ and \∗are the Fourier transforms of the unprocessed and processed sound and ;are the indices for frequencies < 20 Hz and > 12200 Hz; (2) a sound intensity loss %DY(20 ∗STI^^(^%\(^)), 20 ∗ STI^^(^%\(^∗)) ); and (3) an SI-SDR loss (Roux, J. L., et al.) computedin the Mel domain, i.e., after passing the unprocessed and processed sound through 32 filters spaced logarithmically between 80 and 12000 Hz. The three individual loss components weremultiplied with weighting factors of 5 ⋅ 10ab, 5 ⋅ 10ac and 5 ⋅ 10ac, respectively, to ensure alower contribution compared to the loss term ^M.
[0204] Speech recognition. To ensure that AidNet enhanced the features of neural activity that are most relevant for speech perception, another loss term was added to assess the neural coding of speech. Cross-entropy loss was used to train an ASR backend to classify phonemes from YW[^^̂∗^→^^]. This cross-entropy loss was added to the loss terms ^Mand ^Kwith a weighting factor of 1. An ASR architecture based on ConvTasNet (Luo, Y. et al.) was used that predicted phonemes across time in response to 762.9395 Hz frames of neural activity. The ASR architecture comprised a block of 8 dilated convolutional layers (dilation from 1 to 2d) with 128 filters and a kernel size of 3. A PReLU activation was used after each layer except for the last layer, which used 40 output filters followed by a softmax activation to predict the probabilities of 40 phoneme classes.
[0205] Training dataset. The train subset of the TIMIT dataset was used to train AidNet. Sentences were calibrated to randomly selected levels between 40 and 90 dB SPL, and were mixed with noise from the MSNS dataset at SNRs of -10, -5, 0, 5, 10, 20, and 100 dB. The phonetic transcriptions of the TIMIT dataset were downsampled to the sampling frequency of the neural activity (762.9395 Hz) and were grouped into 40 classes (15 vowels and 24 consonants plus the glottal stop).
[0206] Evaluation. The test subset of the TIMIT dataset was used to evaluate AidNet. A sentence was randomly chosen and was calibrated at 50 dB SPL, and mixed with Cafeteria noise from the MSNS dataset at 5 dB SNR. Two standard amplification strategies were used as a baseline and were compared against the restoration achieved by AidNet: (1) a linear amplification strategy (NAL-R) that applied a frequency-dependent amplification and (2) a compressive amplification strategy (NAL-R,WDRC) that applied a frequency- and level- dependent amplification. The implementations of the two strategies on Python from the Cadenza Challenge code (github.com / claritychallenge / clarity) were used. Speech recognition performance was evaluated with the ASR models that were used in the training of AidNet. Overall performance was grouped to percentages of correct classification of vowels and consonants.
[0207] In summary of the experiments described above, the efficacy of the present techniques was demonstrated. In particular, ICNets were trained on one normal brain and one impaired AL Ref: P46575WO1 14 August 2025 brain. The hearing impairment was induced through overexposure to high-intensity broadband noise, resulting in sloping mild-to-moderate hearing loss (20 dB at 1 kHz increasing to 50 dB at 8 kHz). The ICNets were used to train an AidNet as described above. After training was complete, the AidNet’s performance was assessed in restoring neural coding to normal and compared its performance against that of two other common hearing aid algorithms: NAL-R, a linear algorithm that provides frequency-weighted amplification, and NAL-R-WDRC, a nonlinear algorithm that provides frequency-weighted amplification as well as amplitude compression.
[0208] Figure 10A shows a performance comparison of the present techniques against existing techniques as a result of the experiments above. In particular, Figure 9A shows simulated multi-unit activity for a segment of speech under different conditions. The top two rows show example sound waveforms and spectrograms. The third row shows simulated activity resolved across individual units. The neural activity differed dramatically across the normal and impaired brains and these differences were largely corrected by AidNet and NAL-R, but not by NAL-R-WDRC. This is consistent with existing results (Armstrong, A., et al.) showing that the WDRC algorithm fails to correct distortions in the central neural coding of speech. Compressive hearing aid algorithms are designed to compensate for a loss of physiological compression within the cochlea after hearing loss but, in practice, they provide little additional benefit beyond linear amplification and may even introduce additional distortions.
[0209] To estimate the degree to which the restoration of normal neural activity could be expected to improve speech perception, the same ASR backend that was included in the AidNet training to identify phonemes from mean activity patterns (with retraining to optimise the ASR for each condition) was used.
[0210] Figure 10B shows the ASR performance of the present techniques. In particular, Figure 9C shows the performance of an automatic speech recognition system trained to identify phonemes from neural activity under different conditions. Average results across 15 vowels and 25 consonants are shown. The symbols from left to right show the performance for normal activity, impaired activity, and impaired activity with AidNet, NAL-R, or NAL-R-WDRC processing, respectively. The performance for the impaired brain was much lower than that for the normal brain, especially for consonants. Both AidNet and NAL-R increased phoneme recognition to normal levels, with AidNet achieving slightly higher performance.
[0211] The present techniques demonstrate that ML techniques can be used to develop sound transformations that correct distortions in neural coding after hearing loss is both feasible and effective. The present techniques involve first using large-scale neural recordings to train forward models of neural coding in the brain, ICNets (i.e. the first and second pre-trained modules), and then using the trained ICNets to train a sound transformation module, AidNet. AL Ref: P46575WO1 14 August 2025 Such an approach has only recently become possible because of advances in deep learning for nonlinear modelling and in electrophysiology for collecting the required datasets. Because the present models are trained directly from brain recordings, the present techniques bypass the gaps in our understanding of auditory processing and the limitations of existing computational models. However, it also presents significant challenges.
[0212] One of the biggest challenges that is unique to the present techniques is the problem of aligning neural activity from two different brains. There are many contexts in which activity from different brains must be aligned, but the present case is unique in that it must be ensured that the alignment preserves differences related to hearing impairment. (Alignment via canonical correlation analysis, for example, would be insensitive to differences in overall activity.) It has been shown that a conservative form of alignment that allows only channel rearrangement can be effective, but there are alternative forms of alignment, such as the others described above (e.g. affine transformation and MCA) that may result in better performance.
[0213] References ^ Sabesan, S., Fragner, A., Bench, C., Drakopoulos, F. & Lesica, N. A. Large-scale electrophysiology and deep learning reveal distorted neural signal dynamics after hearing loss. eLife 12, e85108 (2023). ^ Armstrong, A. G., Lam, C. C., Sabesan, S. & Lesica, N. A. Compression and amplification algorithms in hearing aids impair the selectivity of neural responses to speech. Nat. Biomed. Eng.6, 717–730 (2022). ^ Granley, J., Fauvel, T., Chalk, M. & Beyeler, M. Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. Adv. Neural Inf. Process. Syst.36, 79376–79398 (2023). ^ Stoller, D., Ewert, S. & Dixon, S. Wave-U-Net: A Multi-Scale Neural Network for End-to- End Audio Source Separation. Preprint at https: / / doi.org / 10.48550 / arXiv.1806.03185 (2018). ^ Roux, J. L., Wisdom, S., Erdogan, H. & Hershey, J. R. SDR - half-baked or well done? Preprint at https: / / doi.org / 10.48550 / arXiv.1811.02508 (2018). ^ Luo, Y. & Mesgarani, N. Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation. IEEEACM Trans. Audio Speech Lang. Process. 27, 1256–1266 (2019). ^ Garofolo, John S. et al. TIMIT Acoustic-Phonetic Continuous Speech Corpus.715776 KB Linguistic Data Consortium https: / / doi.org / 10.35111 / 17GK-BN40 (1993). AL Ref: P46575WO1 14 August 2025 ^ Rafii, Z., Liutkus, A., Stöter, F.-R., Mimilakis, S. I. & Bittner, R. MUSDB18-HQ - an uncompressed version of MUSDB18. Zenodo https: / / doi.org / 10.5281 / ZENODO.3338373 (2019). ^ Reddy, C.K., Beyrami, E., Pool, J., Cutler, R., Srinivasan, S. and Gehrke, J., 2019. A scalable noisy speech dataset and online subjective test framework. arXiv preprint arXiv:1909.08050. ^ Ravanelli, M. & Bengio, Y. Speaker Recognition from Raw Waveform with SincNet. arXiv:1808.00158 (2019). ^ Depireux, D. A., Simon, J. Z., Klein, D. J. & Shamma, S. A. Spectro-temporal response field characterization with dynamic ripples in ferret primary auditory cortex. J.Neurophysiol. 85, 1220–1234 (2001).
[0214] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present techniques, the present techniques should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.
Claims
1. AL Ref: P46575WO1 14 August 2025 CLAIMS 1. A computer-implemented method for training a machine learning, ML, model to generate modified audio signals for individuals, the method comprising: obtaining a plurality of audio signals; for each audio signal of the plurality of audio signals: generating a first representation of a neural response to the audio signal in a first set of brains; transforming, using a sound transformation module of the ML model, the audio signal to generate a modified audio signal; generating a second representation of a neural response to the modified audio signal in a second set of brains; and generating an aligned second representation of a neural response in the second set of brains by minimising differences in the generated second representation relative to the generated first representation other than differences arising from one or both of: a difference in hearing ability between the first set of brains and the second set of brains, and a difference between the audio signal and the modified audio signal; and training the sound transformation module of the ML model by: minimising a loss function that is based on a difference between the aligned second representation of a neural response to the modified audio signal and the first representation of a neural response to the audio signal.
2. The method as claimed in claim 1 wherein generating an aligned second representation of a neural response comprises minimising differences in the generated second representation relative to the generated first representation arising from one or both of: anatomical differences between the first set of brains and the second set of brains that are unrelated to hearing impairment; and differences in position of apparatus for recording neural responses to audio signals.
3. The method as claimed in claim 1 or 2 wherein obtaining a plurality of audio signals comprises obtaining a training dataset comprising a plurality of audio signals.
4. The method as claimed in claim 1 or 2 wherein obtaining a plurality of audio signals comprises using a generative model to generate a plurality of audio signals.AL Ref: P46575WO1 14 August 2025 5. The method as claimed in claim 1, 2, 3 or 4 wherein: the method is for training a machine learning, ML, model to generate modified audio signals for hearing-impaired individuals; the first set of brains is a set of hearing-unimpaired brains; the second set of brains is a set of hearing-impaired brains, wherein the hearing- impaired brains are changed due to a hearing impairment; and transforming the audio signal to generate a modified audio signal comprises transforming the audio signal for hearing-impaired individuals.
6. The method as claimed in claim 5 wherein: generating the first representation of a neural response to the audio signal comprises using a first pre-trained forward model of the ML model; and generating the second representation of a neural response to the modified audio signal comprises using a second pre-trained forward model of the ML model.
7. The method as claimed in claim 6 wherein: the first pre-trained forward model is for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation comprises generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains; and the second pre-trained forward model is for simulating neural responses in the second set of hearing-impaired brains, and generating the second representation comprises generating a second neural response to the modified audio signal for each brain in the second set of hearing-impaired brains.
8. The method as claimed in claim 7 wherein: generating the single first neural response to the audio signal comprises: generating, using an encoder of the first pre-trained forward model, a representation of the audio signal, and generating, using a decoder of the first pre-trained forward model and the representation of the audio signal, a single representation of the neural activity in the first set of hearing-unimpaired brains; and generating the second neural response to the modified audio signal comprises: generating, using an encoder of the second pre-trained forward model, a representation of the modified audio signal, andAL Ref: P46575WO1 14 August 2025 generating, using a decoder of the second pre-trained forward model and the representation of the modified audio signal, a representation of the neural activity in each brain of the second set of hearing-impaired brains.
9. The method as claimed in claim 5 wherein: generating the first representation of a neural response to the audio signal comprises using a first pair of pre-trained encoders of the ML model; and generating the second representation of a neural response to the modified audio signal comprises using a second pair of pre-trained encoders of the ML model.
10. The method as claimed in claim 9 wherein: generating the first representation comprises: generating, using a first encoder of the first pair of pre-trained encoders, a first latent representation of features of the audio signal; generating, using a second encoder of the first pair of the pre-trained encoders, a second latent representation of features of a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing- unimpaired brains, wherein the first and second latent representations are in the same latent space; and generating a first representation by minimising a difference between the first and second latent representations; and generating the second representation comprises: generating, using a first encoder of the second pair of pre-trained encoders, a third latent representation of features of the modified audio signal; and generating, using a second encoder of the second pair of the pre-trained encoders, a fourth latent representation of features of a single second neural response to the audio signal for each brain of the second set of hearing-impaired brains, wherein the third and fourth latent representations are in the same latent space; and generating a second representation by minimising a difference between the third and fourth latent representations.
11. The method as claimed in any one of claims 5 to 10 further comprising personalising the training for a specific hearing-impaired brain by: generating a latent representation of the audio signal; and obtaining a personalisation parameter for the specific hearing-impaired brain;AL Ref: P46575WO1 14 August 2025 wherein generating the second representation of a neural response comprises transforming, using the obtained personalisation parameter and a personalisation parameter for each hearing-impaired brain in the second set of brains, the latent representation to generate the second neural response for the specific hearing-impaired brain, wherein the personalisation parameter corresponds to a quantification of effects of the hearing impairment of the specific hearing-impaired brain.
12. The method as claimed in claim 1, 2, 3 or 4 wherein: the method is for training a machine learning, ML, model to generate less distorted audio signals; the first set of brains is a set of hearing-unimpaired brains; and the second set of brains is a set of hearing-unimpaired brains.
13. The method as claimed in claim 12 further comprising: modifying, prior to transforming the audio signal, the audio signal of the plurality of audio signals to include distortion, thereby generating a modified distorted audio signal.
14. The method as claimed in claim 13 wherein: generating the first representation of a neural response to the audio signal comprises using a first pre-trained forward model of the ML model for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation comprises generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains; and generating the second representation of a neural response to the modified audio signal comprises using a second pre-trained forward model of the ML model for simulating neural responses in the second set of hearing-unimpaired brains, and generating the second representation comprises generating a second neural response to the modified distorted audio signal for each brain in the second set of hearing-unimpaired brains.
15. The method as claimed in claim 1, 2, 3 or 4 wherein: the method is for training a machine learning, ML, model to generate less distorted audio signals; the first set of brains is a set of hearing-unimpaired brains; and the second set of brains is a set of hearing-impaired brains.
16. The method as claimed in claim 15 further comprising:AL Ref: P46575WO1 14 August 2025 modifying, prior to transforming the audio signal, the audio signal of the plurality of audio signals to include distortion, thereby generating a modified distorted audio signal.
17. The method as claimed in claim 16 wherein: generating the first representation of a neural response to the audio signal comprises using a first pre-trained forward model of the ML model for simulating neural responses in the first set of hearing-unimpaired brains, and generating the first representation comprises generating a single first neural response to the audio signal that reflects common features of neural response across the first set of hearing-unimpaired brains; and generating the second representation of a neural response to the modified audio signal comprises using a second pre-trained forward model of the ML model for simulating neural responses in the second set of hearing-impaired brains, and generating the second representation comprises generating a second neural response to the modified distorted audio signal for each brain in the second set of hearing-impaired brains.
18. The method as claimed in any preceding claim wherein obtaining a plurality of audio signals comprises obtaining a plurality of audio signals comprising speech.
19. The method as claimed in claim 18 wherein training the sound transformation module of the ML model comprises: identifying, using an automatic speech recognition, ASR, module, phonemes from the first representation of a neural response and from the aligned second representation of a neural response; wherein minimising a loss function comprises minimising the loss function using differences between the identified phonemes.
20. A computer-implemented method for generating, on a user device, modified audio signals for a hearing-impaired person using a trained machine learning, ML, model trained according to the method of any of claims 1 to 19, the method comprising: receiving an input audio signal; generating, using a sound transformation module of the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the hearing-impaired person that is similar to an auditory brain response to the input audio signal in a hearing- unimpaired brain; and outputting the modified audio signal.AL Ref: P46575WO1 14 August 2025 21. A computer-implemented method for generating, on a user device, modified audio signals for distortion reduction using a trained machine learning, ML, model trained according to the method of any of claims 1 to 19, the method comprising: receiving a distorted input audio signal; generating, using a sound transformation module of the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the user of the user device that is similar to an auditory brain response to an undistorted version of the distorted input audio signal; and outputting the modified audio signal.
22. A user device for generating modified audio signals for a hearing-impaired person using a trained machine learning, ML, model on a user device trained according to the method of any of claims 1 to 19, the user device comprising: storage for storing the trained ML model; and a processor coupled to memory, for: receiving an input audio signal; generating, using the trained ML model, a modified audio signal for generating an auditory brain response in the brain of the hearing-impaired person that is similar to an auditory brain response to the input audio signal in a hearing-unimpaired person; and outputting the modified audio signal.
23. The user device as claimed in claim 22 wherein the user device is any of: a hearing aid; a cochlear implant; a pair of headphones; earphones; a smartphone; a tablet; a headset; a smart speaker; and a speaker.
24. The user device as claimed in claim 22 or 23 further comprising a speaker for outputting the modified audio signal, or electronics for outputting electrical stimulation based on the modified audio signal.
25. A computer-readable storage medium comprising instructions which, when executed by a processor, causes the processer to carry out the method of claims 1 to 21.
Citation Information
Patent Citations
Signal processing in a hearing device
US20210185465A1
A neural network model for cochlear mechanics and processing
US20220248148A1
Closed-loop method to individualize neural-network-based audio signal processing
US20230156413A1