Signal processing method and device for hearing aid
A neural network-based hearing aid system dynamically adjusts amplification using DNN, LSTM, CRNN, and Transformer architectures to address flexibility and personalization issues in traditional hearing aids, improving speech intelligibility and comfort.
Patent Information
- Application Number
- PCT/US2025/036299
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-08
AI Technical Summary
Existing hearing aid technologies struggle with limited flexibility in noisy environments, often leading to discomfort or auditory overload due to equal amplification of background noise and speech, and lack personalized adaptability for varying degrees of hearing loss.
A neural network-based signal processing method and device that utilizes deep learning to dynamically adjust amplification strategies based on individual hearing profiles, incorporating neural networks like DNN, LSTM, CRNN, and Transformer, and integrates with WDRC for personalized sound amplification.
Enhances speech intelligibility and quality, providing adaptable and comfortable sound amplification tailored to individual hearing needs, overcoming limitations of traditional methods.
Smart Images

Figure US2025036299_08012026_PF_FP_ABST
Abstract
Description
SIGNAL PROCESSING METHOD AND DEVICE FOR HEARINGAIDFIELD OF INVENTION
[0001] The present description relates to hearing aid device and method.BACKGROUND
[0002] The global prevalence of hearing loss has increased steadily in recent years. The World Health Organization (WHO) also highlights that millions people worldwide currently face problems related to hearing loss. If left untreated, individuals with hearing loss may encounter height-end risks, including social isolation, depression, diminished quality of life, and cognitive decline. Surprisingly, only some of eligible individuals in the United States consistently use hearing aids (HA). Despite the increasing prevalence of hearing loss and widespread awareness of its risks when left untreated, the use of hearing aids as an effective intervention method remains relatively low.SUMMARY
[0003] The disclosure provides a signal processing method for hearing aid, comprising: receiving an incoming sound signal; providing a listener audiogram; amplifying the incoming sound signal based on a representation of the listener audiogram and features of the incoming sound signal by using a neural network; and outputting an amplified sound signal from the neural network.
[0004] The disclosure provides a device for hearing aid, comprising: at least one processor; and at least one storage media including a listener audiogram and instructions operable to be executed by the at least one processor to perform a set of operations comprising: receiving an incoming sound signal; amplifying the incoming sound signal based on a representation of the listener audiogram and features of the incoming sound signal by using a neural network; and outputting an amplified sound signal from the neural network.
[0005] In some embodiments, amplifying the incoming sound signal comprises: applying feature extraction on the incoming sound signal by time-to-frequency conversion to output the features of the incoming sound signal; applying a dense layer for receiving the listener audiogram to output the representation of the listener audiogram; and concatenating the features of the incoming sound signal and the representation of the listener audiogram as a comprehensive input of the neural network.
[0006] In some embodiments, the listener audiogram is provided according to the incoming sound signal. The neural network is trained based on loss function, wherein the loss function is calculated according to an estimated spectral representation of the neural network and a ground-truth spectral representation, the ground-truth spectral representation is obtained by applying a prescriptionformula and WDRC to the incoming sound signal. In some embodiments, the listener audiogram is provided according to a clean sound signal, and the incoming sound signal has noise. The neural network is trained based on loss function, wherein the loss function is calculated according to an estimated spectral representation of the neural network and a ground-truth spectral representation, the ground-truth spectral representation is obtained by applying a prescription formula and WDRC to the clean sound signal.
[0007] In some embodiments, the WDRC in amplifying the incoming sound signal further comprises: applying a filterbank to group individual time-to-frequency conversion bins into a predefined number of channels and generating an estimation of short-term level; calculating frequency-specific gain functions based on the estimation of short-term level; interpolating channelspecific gain functions to the individual time-to-frequency conversion bins; and applying the channel-specific gain functions to the features of the incoming sound signal.
[0008] In some embodiments, amplifying the incoming sound signal includes insertion gain.
[0009] In some embodiments, amplifying the incoming sound signal further comprises: applying a final dense layer for the neural network to output amplified features; and converting the amplified features by frequency -to-time conversion to output the amplified sound signal.
[0010] In some embodiments, the time-to-frequency conversion is short time Fourier transform (STFT). The frequency-to-time conversion is inverse short time Fourier transform (iSTFT).
[0011] In some embodiments, the neural network is constructed with at least one of Convolutional Neural Network (CNN), deep neural network (DNN), long short-term memory (LSTM), convolutional recurrent neural network (CRNN), and Transformer.
[0012] In some embodiments, the neural network is trained with speech data and music data. In some embodiments, the neural network is integrated with prescription fitting formula and wide dynamic range compression (WDRC).
[0013] In some embodiments, the neural network includes neural speech enhancement and neural echo cancellation, thereby enabling end-to-end optimization.BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to sufficiently understand the essence, advantages and the preferred embodiments of the present invention, the following detailed description will be more clearly understood by referring to the accompanying drawings.
[0015] FIGS. 1A and IB depict schematic diagrams showing signal processing with integrative neural network for hearing aid according to some embodiments.
[0016] FIG. 2 depicts a schematic diagram showing signal processing with integrative neural network for hearing aid according to some embodiments.
[0017] FIGS. 3A-3C show tables 1-3 according to some embodiments.
[0018] FIGS. 4-11 show scatter plots according to some embodiments.
[0019] FIG. 12 depicts a schematic diagram showing signal processing with neural network for hearing aid according to some embodiments.
[0020] FIG. 13 depicts a schematic diagram showing signal processing with neural network for hearing aid according to some embodiments.
[0021] FIGS. 14A-14E show tables 4-8 according to some embodiments.
[0022] FIGS. 15-20 show scatter plots according to some embodiments.
[0023] FIG. 21 shows waveform comparison according to some embodiments.
[0024] FIG. 22 shows spectrogram comparison according to some embodiments.
[0025] FIG. 23 shows frequency band energy over time comparison according to some embodiments.
[0026] FIG. 24 shows frequency-specific gain comparison according to some embodiments.
[0027] FIG. 25 shows table 9 according to some embodiments.
[0028] FIGS. 26-29 show bar charts according to some embodiments.DETAILED DESCRIPTION
[0029] The following description shows the preferred embodiments of the present invention. The present invention is described below by referring to the embodiments and the figures. Thus, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the principles disclosed herein. Furthermore, that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims.
[0030] There are many factors that influence the rejection of hearing aids, which can be roughly divided into two categories: non-audiological determinants and audiological determinants. In the case of non-audiological determinants, the factors include self-perceived hearing problems, expectations, demo-graphics, group consultation, support from significant others, self-perceived benefit, and satisfaction. On the other hand, for audiological determinants, the factors include the severity of hearing loss, type of hearing aid, tolerance to background noise, and adjustment to better match prescription targets. Furthermore, considering such factors is certainly important to obtain the best-fitting hearing aid systems and to gain a better understanding of users in both physiological and technical aspects.
[0031] Some hearing aids operate primarily by amplifying sounds to make them audible based on an individual’s specific hearing loss profile. The process begins with an audio-metric evaluation to measure hearing thresholds across various frequencies. From these results, audiologists useprescriptive formulas, such as the National Acoustic Laboratories’ Revised (NAL-R) or the Desired Sensation Level (DSL) method, to derive a personalized compensation gain curve. An example of one amplification method is linear amplification, which applies a consistent amount of gain to all sound levels. This approach is straightforward and cost-effective, making it a reliable choice in environments where noise levels are stable and predictable. However, it offers limited flexibility as it does not adjust to the loudness of incoming sounds, which can result in discomfort or auditory overload in variable noise environments. Additionally, linear amplification can struggle in noisy settings because it amplifies background noise equally with speech, making it difficult for users to focus on conversations. Like many types of aids, systems using linear amplification can also produce feedback, particularly at higher volumes or frequencies, which can be disruptive and uncomfortable for the wearer.
[0032] In contemporary hearing aid fitting, the gain is initially calculated based on the hearing thresholds of individuals with hearing loss, in conjunction with compensation prescriptions fitting formula. This process yields a compensation gain curve, which is then implemented within commonly used hearing aid amplification architectures, such as Wide Dynamic Range Compression (WDRC). Consequently, the appropriateness of the compensation gain provided to individuals with hearing loss is directly influenced by both the compensation prescription and the amplification architecture employed in the hearing aid. The process of amplifying audio in hearing aids is not a simple step. It involves the calculation of hearing aid fitting formulas and various processing steps of WDRC.
[0033] In recent years, the field of deep learning has witnessed remarkable advancements, permeating various domains and revolutionizing approaches. Among these domains, the domain of hearing aid technology has emerged as a promising field for the integration of neural networks. In the disclosure, NeuroAMP (Neural Amplifier) is provided, which is an innovative framework in the field of hearing aid technology, leveraging neural networks for advanced signal processing and amplification. Traditional hearing aid processing techniques have long relied on conventional signal processing algorithms, often constrained by their linear nature. However, the advent of deep learning has paved the way for transformative changes in this domain. NeuroAMP represents a paradigm shift by replacing conventional processing methodologies with neural networks, thereby providing unparalleled flexibility and unlocking opportunities for nonlinear processing approaches in hearing aid technology. The limitations inherent in existing nonlinear hearing aid processes have posed significant challenges for researchers and practitioners alike. However, with the introduction of NeuroAMP, these limitations are poised to be overcome. By leveraging the capabilities of neural networks, NeuroAMP enables the realization of advanced nonlinear processing techniques thatwere previously unattainable, thereby pushing the boundaries of what can be achieved in hearing aid technology. One of the topics of NeuroAMP lies in its ability to provide robust and differentiable processes tailored specifically for hearing loss compensation. The inherent adaptability and learning capabilities of neural networks empower NeuroAMP to dynamically adjust and optimize its processing strategies, thereby offering personalized solutions for individuals with varying degrees of hearing impairment.
[0034] The development of NeuroAMP involved a comprehensive exploration of various neural network architectures commonly employed in signal processing tasks, followed by a comparative analysis among them. These architectures served as the foundational building blocks for NeuroAMP. To ensure the versatility and adaptability of NeuroAMP, extensive experimentation is conducted by using a diverse array of datasets. These datasets encompassed English speech data, Taiwanese Mandarin speech data, and music data, thereby capturing the linguistic and acoustic nuances inherent in different languages and musical genres. To leverage these datasets, the procedure includes a thorough data augmentation process, augments the original data with various forms of noise, and enhances their quality using speech enhancement (SE) models. The augmentation process not only augments the size of the dataset but also imbues NeuroAMP with resilience against real-world environmental noise, thereby enhancing its practical utility in real- world scenarios. Furthermore, to assess the generalization ability of NeuroAMP, evaluations are conducted by using unseen datasets comprising speech and music samples distinct from those used during training. This evaluation strategy is enabled to gauge NeuroAMP’s ability to generalize across different domains and adapt to unseen data, thereby affirming its efficacy in real-world applications. Through iterative refinement and validation against benchmark datasets, NeuroAMP emerges as a versatile and robust framework capable of addressing the diverse challenges encountered in hearing aid technology. By harnessing the power of neural networks and leveraging diverse datasets augmented through sophisticated techniques, NeuroAMP stands poised to usher in a new era of innovation in the field of hearing aid processing.
[0035] To overcome the difficulty in amplification optimization composed of multiple modular components which are difficult to optimize collectively, the signal processing with the neural network leverages deep neural network models to perform amplification within hearing aids. The signal processing with the neural network is constructed with different model architectures, including Deep Neural Network (DNN), Long Short-Term Memory (LSTM), Convolutional Recurrent Neural Network (CRNN), and Transformer. The loss functions to train the parameters in the neural network are tested. The capability of the neural network in diverse auditory tasks, spanning English speech, Mandarin speech, and music are explored. The neural network using twospeech (TIMIT and TMHINT) and two music (FMA and MTG-Jamendo) datasets is evaluated and a comparative analysis of the performance of the neural network is conducted. One purpose is on the assessment of intelligibility and quality using the Hearing Aid Speech Perception Index (HASPI) and the Hearing Aid Speech Quality Index (HASQI) for speech data and Hearing Aid Audio Quality Index (HAAQI) for music data. Experimental results validate the efficacy of the neural network and demonstrate comparable amplification performance to traditional methods. The signal processing with the neural network’s adaptability is reflected in its seamless integration with other neural network-based modules, including neural speech enhancement and neural echo cancellation, enabling end-to-end optimization. Furthermore, the signal processing with the neural network can easily adapt to varying degrees of hearing loss, highlighting its potential as a personalized hearing solution.
[0036] FIGS. 1 A and IB depict schematic diagrams showing signal processing with neural network for hearing aid according to some embodiments. The signal processing with neural network can be applied to NeuroAMP. The signal processing with neural network can be utilized to develop individualized amplification strategies, meeting specific needs and preferences. The signal processing with neural network can enhance speech intelligibility / quality, and music quality in audio. Manufactures and technology companies can integrate the signal processing into their products or services, offering personalized audio enhancement and auditory solutions. The signal processing with neural network achieves end-to-end amplification. The adaptability of the signal processing with neural network to varying degrees of hearing loss makes it a unique solution for personalized amplification.
[0037] In some embodiments, the signal processing with neural network may be applied to a device for hearing aid. The device includes at least one processor and at least one storage media. The storage media includes the listener audiogram and instructions operable to be executed by the at least one processor to perform a set of operations of the signal processing with neural network.
[0038] The signal processing with neural network is dedicated to improving patient audibility, sound quality, and speech enhancement, and it can utilize deep learning networks and fully differentiable hearing-aid strategies and tend to overlook WDRC or the integration between various functions. WDRC helps prevent sounds from being too loud or too soft, compresses a wide range of volumes into the dynamic range of individuals with hearing loss, and ensures that the amplified sound remains within the individual’s comfortable range. Some individuals with hearing loss may be more sensitive to high volumes, leading to the risk of feedback or acoustic feedback. Therefore, while calculated gain is crucial for improving audibility, integrating WDRC is essential for effectively enhancing individual wearing comfort.
[0039] In FIG. 1A, the signal processing with integrative neural network includes: receiving an incoming sound signal (11); providing a listener audiogram (12); amplifying (13) the incoming sound signal based on a representation of the listener audiogram and features of the incoming sound signal by using an integrative neural network; and outputting an amplified sound signal (14) from the integrative neural network. In FIG. IB, the integrative neural network is integrated with prescription fitting formula (131) and WDRC (132). In some embodiments, the amplified incoming sound signal includes insertion gain and WDRC. In some embodiments, the WDRC (132) in amplifying the incoming sound signal further comprises: applying a filterbank to group individual time-to-frequency conversion bins into a predefined number of channels and generating an estimation of short-term level; calculating frequency-specific gain functions based on the estimation of short-term level; interpolating channel-specific gain functions to the individual time-to-frequency conversion bins; and appliyng the channel-specific gain functions to the features of the incoming sound signal.
[0040] Prescription fitting formula (131) refers to configuring appropriate amplification for hearing aids based on the user’s degree of hearing loss and individual needs through a specific prescription procedure. The prescription fitting formula aims to ensure that users can correctly hear speech signals at different frequencies and volume levels, enhancing speech recognition and auditory comfort. The prescription fitting formula can be categorized into linear amplification and non-linear amplification based on the amplification strategy. The main difference between linear amplification and nonlinear amplification lies in the fact that nonlinear amplification compresses the input sound based on its volume, frequency, and the degree of hearing loss. In contrast, linear amplification provides a constant gain across different volume levels.
[0041] The benefit of nonlinear amplification lies in its ability to preserve natural speech quality through a certain degree of volume variation. Linear amplification, on the other hand, may amplify all input speech to a uniform level, and the resulting sound quality from such amplification strategies may not be as widely accepted.
[0042] Common linear amplification methods include NAL-R which is used for mild to moderate hearing loss. Another variant based on NAL-R is NAL-RP, designed to generate the same prescription gains as NAL-R but applicable to a broader range of hearing levels, from mild to profound hearing loss. These two fitting formulas are suitable for linear hearing aids. However, some individuals still incorporate linear formulas with compression for nonlinear hearing aids.
[0043] Commonly used nonlinear fitting formulas include NAL-NL1, NAL-NL2, as well as DSL m[i / o]. These formulas are often applied to nonlinear hearing aids. NAL-NL1 is a nonlinear prescription procedure. In contrast to NAL-R, NAL-NL1 does not normalize relative volume levelsacross frequency components but tends to equalize them while optimizing predicted speech intelligibility for a given volume. NAL-NL2 is an improved version based on NAL-NL1, sharing the same objective. However, NAL-NL2 considers additional factors in calculating gains, such as gender, age, and whether the native language is tonal, which can impact the volume output. NAL- NL2 is also currently the most widely used fitting formula. As for DSL m[i / o] can be adjusted according to the different hearing needs of children and adults, providing varying levels of hearing aid gain for quiet and noisy environments. Additionally, DSL m[i / o] aims to avoid uncomfortable loudness during the hearing aid listening process, ensuring the audibility of important information in conversations as much as possible.
[0044] In some embodiments, NAL-R is chosen as the prescription fitting formula. The primary reason is that NAL-R is considered a suitable foundation for extension, given it research-validated. Moreover, the majority of nonlinear fitting formulas are not open-source, posing limitations on research. Despite evidence suggesting that NAL-R is not suitable for nonlinear hearing aids and WDRC, some individuals compress NAL-R separately to adapt it to the characteristics of nonlinear hearing aids.
[0045] WDRC (132) is a compression strategy that helps prevent sounds from being too loud or too soft. With WDRC, a wide range of volumes can be compressed into the dynamic range of individuals with hearing loss, ensuring that the amplified sound remains within the individual’s comfortable range. Such compression is necessary because the dynamic range of individuals with hearing loss is much smaller than that of those with normal hearing. The use of nonlinear WDRC in hearing aids brings many benefits, such as achieving consistent audibility of speech signals at a safe and comfortable hearing level. Additionally, WDRC is not a one-step audio processing system. The main steps of WDRC are as follows:The input audio signal undergoes analysis through Short-Time Discrete Fourier Transform (STFT): The input audio signal undergoes thorough analysis using STFT. Filterbank Application: A filterbank can be applied to group individual Discrete Fourier Transform (DFT) bins into a predefined number of channels, and the short-term level is estimated.Calculation of Frequency-Specific Gain Functions: Based on this level estimation, frequency-specific gain functions are calculated.Interpolation and Application: Channel-specific gain functions are interpolated to individual DFT bins and applied to the STFT representation of the input signal.Signal Reconstruction: The processed output signal is reconstructed from the modified STFT representation by applying an Inverse Short-Time Discrete Fourier Trans- form (ISTFT).
[0046] These steps ensure that the WDRC function can compress a wide range of volumes while maintaining the listener’s comfort range. Additionally, the system considers the listener’s hearing thresholds and dynamic range variations across frequencies to provide improved speech intelligibility.
[0047] The input to the signal processing with integrative neural network encompasses a diverse range of speech signals, including English speech data, Taiwanese Mandarin speech data, and music data. One purpose is to capture the variability in speech across languages and the complexity inherent in music signals. Within speech data, the inputs are categorized into three distinct categories: clean speech, speech with noise, and enhanced speech. To generate noisy speech, 100 types of environmental noise with different signal-to-noise ratios are introduced into the clean speech. The enhanced speech is derived by applying two widely used deep learning models. The first model of the integrative neural network of the signal processing is based on LSTM, which operates on Short Time Fourier Transform (STFT) features to enhance noisy speech. The second model is based on a Fully Convolutional Network (FCN), which inputs the raw waveform to enhance noisy speech. Employing these two methods results in two distinct types of enhanced speech, both of which serve as input to the signal processing with integrative neural network. For music data, two categories are primarily considered: clean music and music with noise. Clean music represents pristine audio recordings, while music with noise incorporates various environmental disturbances to simulate real-world listening scenarios, thereby enhancing input diversity. Each input signal is then paired with a listener audiogram, and the signal processing with integrative neural network processes the combined input to provide an amplified signal to the listener. Given a combined input, the ground truth hearing aid processed speech reference is obtained in two steps: the prescription insertion gain is calculated using the NAL-R hearing-aid fitting formula based on the input listener audiogram, followed by compression using WDRC.First Embodiment:
[0048] FIG. 2 depicts a schematic diagram showing signal processing with integrative neural network for hearing aid according to some embodiments. In some embodiments, amplifying (13) the incoming sound signal includes: applying feature extraction on the incoming sound signal by time-to-frequency conversion (133) to output the features of the incoming sound signal; applying a dense layer (134) for receiving the listener audiogram to output the representation of the listeneraudiogram; concatenating (135) the features of the incoming sound signal and the representation of the listener audiogram as a comprehensive input of the integrative neural network; inputting the result of the concatenating into the integrative neural network (136); applying a final dense layer (137) for the integrative neural network to output amplified features; and converting the amplified features by frequency -to-time conversion (138) to output the amplified sound signal.
[0049] The time-to-frequency conversion may be short time Fourier transform (STFT) or other conversions. The frequency -to-time conversion may be inverse short time Fourier transform (iSTFT) or other conversions depending on the type of the time-to-frequency conversion applied to the incoming sound signal.
[0050] In some embodiments, the integrative neural network (136) is constructed with at least one of deep neural network (DNN), long short-term memory (LSTM), convolutional recurrent neural network (CRNN), and Transformer. The integrative neural network is trained with speech data and music data. The integrative neural network includes neural speech enhancement and neural echo cancellation, thereby enabling end-to-end optimization. In some embodiments, the integrative neural network (136) is constructed with a combination of at least two of deep neural network (DNN), long short-term memory (LSTM), convolutional recurrent neural network (CRNN), and Transformer.
[0051] Regarding the model architecture of signal processing with integrative neural network, as shown in FIG. 1A or FIG. 3, the signal processing with integrative neural network takes a sound signal (Y = [yl , . . . , yn . . . , yN ]) and a listener audiogram (Z = [zl , . . . , zn . . . , zN ]) as input, processes this combined input through a neural network (NN) model, and generates a processed output signal. FIG. 3 illustrates the overall architecture of the signal processing with integrative neural network. In some embodiments, four neural network architectures are used: Deep Neural Network (DNN), Long Short-Term Memory (LSTM), Convolutional Recurrent Neural Network (CRNN), and Transformer. The latter is a combination of Convolutional Neural Network (CNN) and LSTM. Both architectures are widely recognized for their effectiveness in signal processing tasks. Feature extraction of the input signal can be accomplished through Short-Time Fourier Transform (STFT). The STFT decomposes the input signal into its time-frequency representation, capturing both the magnitude and phase components with emphasis on utilizing the magnitude, which decomposes the input signal into a time-frequency representation. However, the magnitude is specifically utilized as the primary input feature. Simultaneously, the audiogram is processed through a single dense layer. These two processed components are then concatenated before being forwarded into the subsequent layers of the model. The rationale for processing the audiogram through the dense layer is rooted in extensive experiments demonstrating its efficacy over directcombination of the audiogram and STFT features. The feedforward process of the signal processing with integrative neural network is defined as follows:PS = ST FT(Y)Zproj = DenseiL'Concat=[PS| Zproj]Y = NeuralNet(C oncat ) (1)
[0052] In the LSTM model architecture, two LSTM layers are employed, while the CRNN architecture utilizes 12 convolutional layers followed by 2 LSTM layers. The output of the LSTM or CRNN modules is directed to the final dense layer to generate amplified features, which are then converted into the final processed signal using phase information from the input via the Inverse Short-Time Fourier Transform (iSTFT).
[0053] To optimize NeuroAMP, Mean Squared Error (MSE) is adopted as the loss function. This facilitates the optimization process by penalizing deviations between predicted (yi) and actual (yi) spectral representations. Specifically, it guides the network towards learning spectral mappings that accurately capture the intricacies of the input features, ensuring effective training of the signal processing with integrative neural network for tasks reliant on spectral information. The MSE loss is computed as follow, where N is the number of samples:
[0054] For model training and evaluation phase, a combination of datasets is utilized to train and test the model. In some embodiments, the TIMIT, TMHINT, and MUSIC datasets are employed for training, while for testing and evaluating the generalization of the model, the VoiceBank dataset and MUSDB18-HQ dataset are included.
[0055] The training dataset are described hereinafter.
[0056] 1) TIMIT Dataset: The TIMIT dataset is a widely acknowledged speech corpus for tasks such as speech recognition and speech enhancement. It comprises recordings from 630 speakers representing eight major dialects of American English, with each speaker reading 10 phonetically rich sentences. The TIMIT training set consists of 4,620 audio files, while the test set comprises 1,690 audio files. Additionally, the TIMIT dataset is augmented by generating noisy and enhanced speech copies to diversify the training data.
[0057] 2) TMHINT Dataset: In the model training, data are incorporated from the Taiwan Mandarin Hearing in Noise Test (TMHINT), augmenting the dataset with diverse speech samples to enhance the robustness of the models. The TMHINT dataset consists of recordings from multiple speakers, featuring 16 phonemically balanced lists, each containing 20 sentences. Each sentence iscrafted with ten syllables, representing various phonetic elements of Mandarin, including four primary lexical tones and a neutral tone. By utilizing the TMHINT dataset, one purpose is to explore speech perception in noise, a critical aspect of auditory communication. This dataset facilitates the investigation into the effects of background noise on speech intelligibility and allows assessing the performance of advanced signal processing with integrative neural network techniques, such as noise reduction algorithms, in enhancing speech clarity. Additionally, the TMHINT dataset is extended by introducing environmental noises at different signal-to-noise ratio levels (-5 dB, 0 dB, and 5 dB) and generating enhanced speech copies using state-of-the-art methods like Long Short-Term Memory (LSTM) and Fully Convolutional Network (FCN). This augmentation process aims to diversify the training data and better capture the complexities of real- world listening scenarios, ultimately enhancing the effective- ness of the proposed approaches
[0058] 3) MUSIC Dataset: In the model training, the music data provided by the Cedenza challenge are incorporated, leveraging its diverse collection of musical samples to enhance the versatility of the dataset. Additionally, to simulate real-world listening conditions, noise is introduced to the music data by varying the Signal-to-Noise Ratio (SNR) levels (-5 dB, 0 dB, and 5 dB) and incorporating different types of environmental noises. This augmentation process aimed to replicate common scenarios where music is played amidst various background sounds, such as street noise, crowd chatter, or household disturbances. By integrating music data with augmented noise, the models sought to be trained to better handle noisy musical environments.
[0059] For model testing, the test sets derived from the same datasets used in training, namely TIMIT, TMHINT, and the Cedenza challenge music data are employed, maintaining consistency in evaluation methodology. Specifically, the models are evaluated on test samples from these datasets, considering various Signal-to-Noise Ratio (SNR) levels (-6 dB, 0 dB, and 6 dB) to assess their performance across different noise conditions comprehensively. Additionally, to gauge the generalization capability of the models, unseen speech and music datasets are included for testing. These unseen datasets introduced novel samples not encountered during model training, allowing evaluating the models’ ability to adapt to new data and perform effectively in real-world scenarios beyond the training domain. By testing the models on diverse datasets with varying noise levels and incorporating unseen data for evaluation is to validating their robustness, efficacy, and generalization capacity in addressing speech and music processing tasks in practical settings. The testing dataset are described hereinafter.
[0060] 1) VoiceBank Dataset: During the model testing phase, the test set from the VoiceBank- DEMAND dataset is utilized, a widely recognized resource in the field of speech processing. This dataset comprises recordings of isolated English digits spoken by multiple speakers, capturingdiverse accents and conditions. Specifically, the test set consists of 824 utterances from 2 speakers, encompassing four Signal-to-Noise Ratio (SNR) levels: 17.5 dB, 12.5 dB, 7.5 dB, and 2.5 dB. These varying SNR levels represent different degrees of noise added to the speech signals, providing a comprehensive evaluation of the models’ performance under diverse acoustic conditions.
[0061] 2) MUSDB18-HQ Dataset: the evaluation is expanded to include the MUSDB18 dataset, another prominent resource in the field of audio processing. The MUSDB18 dataset comprises multi-track music recordings, facilitating tasks such as source separation and music transcription. By incorporating the MUSDB18 dataset into the evaluation process is to assess the performance of the models in handling music signals.
[0062] These datasets collectively provide diverse speech signal data for training and evaluating the model’s performance across various domains and conditions.
[0063] Hearing Loss Patterns: In the experiment, audiograms provided by the Clarity Challenge are utilized, which offer a comprehensive dataset for both training and testing phases. These audiograms serve as critical tools for assessing hearing capabilities across a standard set of frequencies. Specifically, audiograms that cover six key frequencies: 250Hz, 500 Hz, 1000 Hz, 2000 Hz, 4000 Hz, and 6000 Hz are analyzed. These frequencies are crucial for evaluating speech intelligibility and understanding the auditory challenges that individuals with hearing loss face. By examining these audiograms is to better tailor the hearing aid models to the unique auditory profiles of different users, ensuring that the experiments are grounded in realistic scenarios that mimic the varied hearing conditions encountered in daily life. This approach allows refining the performance of the novel signal processing with integrative neural network technology, aiming for optimal effectiveness and adaptability in real-world hearing enhancement.Experimental Example
[0064] In the experimental example, the input signal undergoes preprocessing using a Short-Time Fourier Transform (STFT) with a window length of 512 samples, a window increment of 256 samples, and an FFT length of 512 samples. The prescription input, structured as a 1x6 matrix representing the listener audiogram and consisting of intensity values for provided frequencies, which undergoes processing through eight dense layers in the model to generate a refined representation. This processed audiogram information is then concatenated with the STFT features to form a comprehensive input for further model processing.
[0065] The various architectures can be employed as follows:- Deep Neural Network (DNN): This model utilizes two dense layers, each with 512 units, effectively processing and learning from the audiogram data.- Long Short-Term Memory (LSTM): LSTM configuration of the example includes two layers, each with 256 units, designed to capture temporal dependencies in the audio signal.Convolutional Recurrent Neural Network (CRNN): The CRNN model combines four convolutional layers with filter sizes of 16, 32, 64, and 128, followed by two LSTM layers, each with 256 units. This structure is particularly effective for handling both spectral features and temporal sequences.Transformer: The transformer model features four encoder blocks, each with 16 heads. Its attention mechanism enhances the model’s ability to discern subtle nuances in complex auditory environments.
[0066] Each model configuration includes a 257-unit dense layer for the final output, ensuring robust feature extraction and signal processing with integrative neural network. This architecture enables the signal processing with integrative neural network to effectively integrate both the spectral features from the STFT and the detailed information from the prescription input. The combination of DNN, LSTM, CRNN, and Transformer layers enhances the system’s capacity to learn intricate patterns and relationships within the data, ultimately providing a powerful framework for customized hearing aid performance.Experimental Results:
[0067] One comprehensive evaluation strategy examines the performance of the proposed model, emphasizing both intelligibility and quality aspects. To measure the effectiveness of the model in terms of speech perception and maintaining speech quality, two key metrics are employed: the Hearing Aid Speech Perception Index (HASPI) and the Hearing Aid Speech Quality Index (HASQI). These metrics play a key role in quantifying the impact of the model on the perceived intelligibility and quality of speech. For music, the Hearing Aid Audio Quality Index (HAAQI) is employed to assess quality aspects. For robust benchmarking, the IG+WDRC processed signal is considered as the reference or benchmark for the model’s output. To conduct a thorough comparison, three evaluation metrics are utilized: Linear Correlation Coefficient (LCC), Spearman Rank Correlation Coefficient (SRCC), and Mean Squared Error (MSE). Lower MSE scores indicate a closer alignment of predicted scores with the ground-truth HASPI and HASQI scores, with lower values being desirable. Conversely, higher LCC and SRCC scores indicate superior correlations between the predicted and ground-truth assessment scores, with higher values signifying better alignment.
[0068] FIGS. 3A-3C show tables 1-3 according to some embodiments.
[0069] In FIG. 3A, table 1 summarizes the evaluation results of the overall test dataset for the speech dataset, which includes test sets from the training datasets TIMIT and TMHINT, as well as the additional unseen VoiceBank dataset. Using the key metrics LCC, SRCC, and MSE, the evaluation assesses the perfor- mance of four models: DNN, LSTM, CRNN, and Transformer. The assessment covers both HASQI and HASPI. The LCC and SRCC scores reveal strong correlations between the predicted and ground truth assessment scores, while the MSE values provide insights into the accuracy of the models in predicting HASQI and HASPI scores.
[0070] In FIG. 3B, table 2 presents the evaluation results for the music dataset, encompassing both the training dataset test sets and the unseen music dataset test set. The assessment is conducted by calculating HAAQI scores using the model’s output and comparing them to the ground truth HAAQI scores. Key metrics such as LCC, SRCC, and MSE are employed to measure the correlation and accuracy between the predicted and ground truth HAAQI scores across the various datasets.
[0071] For a more detailed analysis, the evaluation is broken down into different speech conditions: clean, noisy, and enhanced. In FIG. 3C, table 3 consolidates the results for clean, noisy, and enhanced speech, focusing on LCC, SRCC, and MSE for both HASQI and HASPI. The structure of the table allows for a quick comparison of how each model performs under various conditions and metrics. The numerical values in Table 3 represent the specific results obtained during the evaluation process. Interpretation of these values will focus on identifying patterns, strengths, and potential areas for improvement in each model’s performance across different speech conditions and evaluation metrics.
[0072] Scatter plots are utilized to provide a visual comparison and assess the alignment between the model’s calculated values and the ground truth for both the speech and music datasets. For the speech dataset, HASPI and HASQI serve as the assessment metrics, while HAAQI scores are calculated for the music dataset. These scatter plots offer an intuitive way to understand the performance of the models across different datasets and architectures, helping to identify any distinct patterns or disparities.
[0073] In some embodiments, FIG. 4 displays scatter plots for HASPI scores on the TIMIT dataset for each model architecture: (a) DNN; (b) LSTM; (c) CRNN; and (d) Transformer. FIG. 5 presents similar plots for HASQI scores on the TIMIT dataset for each model architecture: (a) DNN; (b) LSTM; (c) CRNN; and (d) Transformer. For the TMHINT dataset, FIGS. 6 and 7 show scatter plots for HASPI and HASQI scores, respectively, featuring the same model architectures, for each model architecture: (a) DNN; (b) LSTM; (c) CRNN; and (d) Transformer. FIGS. 8 and 9 illustrate scatterplots for HASPI and HASQI scores on the VCTK dataset, again across all model architectures, for each model architecture: (a) DNN; (b) LSTM; (c) CRNN; and (d) Transformer.
[0074] A comprehensive visual analysis for the Music dataset is also conducted, using scatter plots to evaluate the models’ performance in predicting HAAQI scores. FIG. 10 features scatter plots for the Clarity-Cadenza Challenge dataset, detailing the results for each model architecture: (a) DNN; (b) LSTM; (c) CRNN; and (d) Transformer. FIG. 11, on the other hand, showcases scatter plots for HAAQI scores on the MUSDB18-HQ dataset, offering insights into the performance of these models across another music dataset, for each model architecture: (a) DNN; (b) LSTM; (c) CRNN; and (d) Transformer.
[0075] The signal processing with integrative neural network addresses the critical challenges in hearing aid customization through the introduction of the integrative neural network, an innovative end-to-end neural network model that excels in nonlinear amplification, replacing traditional methods. The methodology incorporates four distinct model architectures DNN, LSTM, CRNN, and Transformer — each employing the Short-Time Fourier Transform (STFT) for spectral feature extraction, which serves as input for further processing. The effectiveness of integrative neural network is consistently highlighted by the evaluation metrics employed: the Hearing Aid Speech Perception Index (HASPI) and the Hearing Aid Speech Quality Index (HASQI) for speech data, along with the Hearing Aid Audio Quality Index (HAAQI) for music data, all supported by the Linear Correlation Coefficient (LCC), Spearman’s Rank Correlation Coefficient (SRCC), and Mean Squared Error (MSE). These findings underscore the integrative neural network as a superior alternative for customization, with the potential to significantly improve auditory experiences for individuals with hearing impairments. The signal processing with integrative neural network may be applied to integration with speech enhancement and hearing loss modules to provide even greater optimization.
[0076] The above embodiments and functional operations can be implemented in digital electronic circuitry, in tangibly-embodied software or firmware, in hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The operations and logic flows described in FIGS. 1 A, IB and 2 can be performed by one or more programmable computing devices executing one or more programs to perform their functions, or by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Computing devices suitable for the execution of the one or more programs include, by way of example, can be based on general or special purpose microcontrollers or both, or general or special purpose microprocessors or both, or any other kind of central processing unit. Computer-readable media suitable for storing program instructions anddata include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0077] Second Embodiment:
[0078] Another deep learning-based neural amplifier model also called NeuroAMP is designed to provide personalized amplification in an end-to-end manner. The non-linearity of the neural network would allow NeuroAMP to capture a greater variety of acoustic information and more effectively adjust gain and compression based on the user’s specific degree of hearing loss at different frequencies. The NeuroAMP involved exploring various neural network architectures, including convolutional neural networks (CNN), long short-term memory (LSTM), convolutional recurrent neural networks (CRNN), and Transformers. The training objective of NeuroAMP is to minimize the loss between the estimated amplified audio and the corresponding ground-truth audio. To ensure robust performance across diverse conditions, a data-augmentation strategy using datasets containing multiple languages and music is applied. Experimental results confirm the superiority of the Transformer model over other models, achieving the best performance in almost all metrics (with Spearman’s Rank Correlation Coefficient (SRCC) scores of 0.9927 for Hearing Aid Speech Quality Index (HASQI) and 0.9905 for Hearing Aid Speech Perception Index (HASPI) on the TIMIT dataset, and 0.9738 for HAAQI on the Cadenza dataset). Additionally, the effectiveness of data augmentation are verified in maintaining good performance on unseen datasets (with SRCC scores of 0.9810 for HASQI and 0.9914 for HASPI on the VCTK dataset, and 0.9892 for HAAQI on the MUSDDB18-HQ dataset). NeuroAMP is able to mimic conventional amplification methods and provide personalized amplification in an end-to-end manner.
[0079] Denoising NeuroAMP combines NeuroAMP with a speech enhancement module, showcasing direct integration between amplification and enhancement. Denoising NeuroAMP involves combining amplification and noise reduction in one module. Specifically, the model takes noisy audio input and outputs enhanced-amplified audio, and the training objective is to minimize the loss between the predicted enhanced-amplified output from the noisy input and the corresponding ground-truth clean audio. Experimental results confirm that Denoising NeuroAMP can achieve a HASPI score of 0.90 and a HASQI score of 0.59, compared to 0.89 and 0.55 for NAL-R+WDRC, and 0.85 and 0.30 for the two-stage Denoising>NAL-R+WDRC baseline, which further confirms the advantages of Denoising NeuroAMP as an integrated amplification and denoising approach.
[0080] Prescription fitting formulas aim to amplify sound based on the user’s degree of hearing loss, making speech signals audible across different frequencies while maintaining comfort. These formulas can be broadly categorized into linear and non-linear approaches. Some examples of linear amplification are NAL-R and NAL-RP. Some examples of non-linear amplification are NAL-NL1, NAL-NL2, and DSL m[i / o],
[0081] Wide Dynamic Range Compression (WDRC) is a widely adopted compression strategy designed to prevent sounds from being either too loud or too soft. The use of WDRC in hearing aids provides several benefits, such as ensuring consistent audibility of speech signals at safe and comfortable hearing levels. Moreover, WDRC involves multiple steps of audio processing, which includes signal analysis through short-time Fourier transform (STFT), filterbank construction, calculation of filterbank-specific gain functions, interpolation and modification, and signal reconstruction. These steps ensure that the WDRC system effectively compresses a wide range of volume levels while maintaining listener comfort. In addition, the WDRC considers the listener’s hearing thresholds and dynamic range variations across frequencies, aiming to enhance speech intelligibility.
[0082] The processing workflow of NeuroAMP involves taking an audio signal and the listener’s audiogram as inputs, processing them through a neural network, and producing an amplified output signal. The detailed architecture of the NeuroAMP model is illustrated in FIG. 12. Specifically, given the input signal y, the STFT is processed to extract y into its time frequency spectral features (Y = [yi,..., yt,..., yr]). Simultaneously, the audiogram z is transformed into a vector representation via a dense layer and then replicated into a sequence of length T (Z = [zi,..., zt,..., ZT]). The two input streams are concatenated along the y-axis feature dimension and then passed to the subsequent layers of the model. In one exemplary setup, two sets of features are concatenated in the latent space, leveraging its ability to capture semantic feature similarity and perform feature fusion. This approach has been shown to be effective for building personalized deep-learning models. The feedforward process of NeuroAMP is defined as follows:Y = STFT(y) Z = Dense(z) Concat = [Y | Z]Y = NeuroAMP(Concat) (3)
[0083] In this embodiment, for the core NN module in NeuroAMP, four neural network architectures are selectd, including the Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Convolutional Recurrent Neural Network (CRNN), and Transformer.
[0084] The Convolutional Neural Network (CNN) is integrated into the NeuroAMP model due to its ability to extract hierarchical features from spatial or spectral data. CNNs are particularly adept at recognizing patterns and features in audio spectrograms from complex audio signals. In the NeuroAMP model, the CNN is used to process spectral features extracted by the STFT, with the goal of capturing detailed representations of the frequency components of the audio signal. The operation process of the CNN in the NeuroAMP model can be described by the following equations: hi = ReLU(Conv(Concat, Wi) + bi) hi = ReLU(Conv(hi, W2) + b?) hi = ReLU(Conv(hi-i, Wi) + bi)Y = Dense(Flatten(hL-i)) (4) where Wi and bi are the convolutional weights and biases of the 1-th layer, respectively, and L represents the number of convolutional layers. This architecture ensures a comprehensive feature extraction process, allowing the NeuroAMP model to more effectively distinguish between auditory elements.
[0085] Long Short-Term Memory (LSTM) is chosen for its proficiency in modeling temporal dependencies in sequential data. In the context of hearing aids, preserving the temporal structure of the audio signal is essential to maintaining speech intelligibility and naturalness. LSTM is effective at capturing long-range dependencies and temporal patterns, which is crucial for processing continuous audio signals and adapting to varying acoustic environments. The process of the LSTM in the NeuroAMP model is described as follows, ht, Ct = LSTM(Concat, ht-i, Ct-i)Y = DenseQrr) (5) where htand Ct are the hidden state and cell state at time step t, respectively.
[0086] Convolutional Recurrent Neural Network (CRNN) combines the strengths of CNN and RNN, making it ideal for signal processing tasks that involve both spatial and temporal feature extraction. Its convolutional layers capture local spectral features, while the recurrent layers model temporal dynamics. This hybrid approach enhances the network’s ability to manage audio signal complexity, thereby improving performance in tasks like speech enhancement. The process of the CRNN in the NeuroAMP model is described as follows: hc= CNN(Concat) ht, Ct = LSTM(Concat, ht-i, Ct-i)Y = DenseQrr) (6)where hcis the output of the convolutional layer.
[0087] The Transformer is included for its state-of-the-art performance in sequence modeling tasks. Unlike traditional RNNs, the Transformer uses a self-attention mechanism to process entire sequences in parallel, allowing for efficient handling of long-range dependencies. This capability is particularly beneficial in audio signal processing, where capturing the global context and dependencies can enhance performance. The process of the Transformer in the NeuroAMP model is described as follows:Q, K, V = WqConcat, WKConcat, WvConcatAttention = Softmax h = FFN( Attention)Y = Dense(h)where Q, K, and V are the query, key, and value matrices, WQ, WK, and Wv are the weight matrices for the query, key, and value, dk is the dimension of the key, and FFN is the feedforward network.
[0088] To optimize NeuroAMP, Mean Squared Error (MSE) is used as the loss function. This approach penalizes deviations between predicted (Y) and ground-truth (Y) spectral representations, guiding the network to learn spectral mappings. The loss function is calculated as follow, where yt and Ytare the ground-truth amplified and estimated amplified spectral representations of the t-th frame of the training audio signal y, and T is the number of frames.
[0089] Denoising NeuroAMP system is provided to address the challenges of noisy environments, which severely impact the effectiveness of hearing aids. This system is built on the NeuroAMP architecture, allowing simultaneous denoising and amplification by seamlessly integrating speech enhancement within the neural network. The architecture of the Denoising NeuroAMP model is illustrated in FIG. 13.
[0090] As shown in FIG. 13, Denoising NeuroAMP takes a noisy audio signal (X) as input, represented as a magnitude spectrogram. Simultaneously, the listener’s audiogram (Z) is processed through a dense layer to construct a vector representation corresponding to the hearing profile. These two features, including the noisy magnitude spectrogram and the audiogram vector, are concatenated and then fed into the core Denoising NeuroAMP neural network. This network, while conceptually similar to the original NeuroAMP model, is trained to perform both speech enhancement and personalized amplification in a single unified process. It learns to adapt itsprocessing based on the characteristics of the noisy input and the specific hearing loss profile, using the audiogram to guide optimization toward the amplified clean audio output.
[0091] The training objective of Denoising NeuroAMP, as shown in Eq. 10, is to minimize the MSE between the ground-truth magnitude spectrogram (obtained by applying the NAL-R prescription formula and WDRC to the clean input audio) and the magnitude spectrogram of the predicted enhanced-amplified output from the noisy input. This ensures that the system learns to produce amplified outputs that are both enhanced and aligned with established hearing aid fitting practices.X = STFT(x) Z = Dense(z) Concat = [X | Z] X = Den-NeuroAMP(Concat) (9)
[0092] The output of the Denoising NeuroAMP network is a predicted magnitude spectrogram of the amplified-enhanced signal. The phase information of the input noisy signal is combined with the predicted magnitude spectrogram to reconstruct the time-domain, enhanced-amplified audio signal through the Inverse Short-Time Fourier Transform (ISTFT).
[0093] EXPERIMENTS
[0094] In this embodiment, the input to NeuroAMP includes two types of speech signals (English and Taiwanese Mandarin speech data) and music data. The proposed models are evaluated under diverse acoustic properties of cross-linguistic speech and music signals. The speech data contains three conditions: clean, noisy, and enhanced speech. To generate noisy speech, 100 types of environmental noise are added to clean speech at different signal-to-noise ratio (SNR) levels. Enhanced speech was then obtained by applying two deep learning-based speech enhancement models to the noisy data. The first model is based on LSTM, which operates on STFT features to enhance noisy speech. The second model is based on a Fully Convolutional Network (FCN), which operates on the raw waveform to enhance noisy speech. For music data, two categories are considered: clean music and noisy music (i.e., music with added noise). Clean music represents the original recording, while noisy music incorporates various environmental disturbances to simulate real world listening scenarios. To train the NeuroAMP model, TIMIT, TMHINT, and the MUSIC dataset provided by the Cadenza Challenge are used. To test and evaluate the model’s performance and generalization, the unseen test set (the acoustic conditions and audio content are entirelydifferent from the training conditions) from the VoiceBank and MUSDB18-HQ datasets is also included.
[0095] Denoising NeuroAMP Dataset: For training and evaluating the Denoising NeuroAMP framework, the VoiceBank+DEMAND dataset is utilized. This dataset is widely recognized and employed in the field of speech enhancement, making it particularly suitable for the task of combining denoising with personalized amplification. The VoiceBank+DEMAND dataset comprises the VoiceBank corpus, which contains clean speech recordings from multiple speakers, and the DEMAND corpus, which provides a diverse set of environmental noises. This combination allows the creation of noisy speech samples by mixing clean speech with various types of noise at different SNR levels. The official split of the dataset was used, which includes a training set of 11,572 utterances from 28 speakers and a test set of 824 utterances from 2 speakers. The availability of both clean and noisy speech makes it ideal for training models to perform both speech enhancement and amplification.
[0096] Hearing Loss Patterns: audiograms from the Clarity Challenge is used, which provides comprehensive samples for both training and testing. These audiograms are essential for evaluating hearing abilities across a standard set of frequencies. Audiograms are specifically examined that cover six fundamental frequencies: 250, 500, 1000, 2000, 4000, and 6000 Hz. These frequencies are important for assessing speech intelligibility and understanding the auditory challenges faced by individuals with hearing loss.
[0097] Experimental Setup
[0098] In the NeuroAMP experimental setup, the input signal is preprocessed with a 512-point STFT, a Hamming window of 32 ms, and a hop of 16 ms. The prescription input, which is a 1x6 matrix representing the listener’s audiogram with intensity values for specific frequencies, is processed through dense layers in the model to create a refined representation. This processed audiogram information is then concatenated with the spectral features to form the input vector for further modeling. The various model architectures and their configurations employed in the study are as follows:Convolutional Neural Network (CNN): The CNN configuration in NeuroAMP includes four convolutional layers with filter sizes of 32, 64, 128, and 256. This setup aims to extract detailed spatial features from the audio signals.- Long Short-Term Memory (LSTM): The LSTM model configuration includes two layers, each with 256 units, designed to capture temporal dependencies in the audio signal.Convolutional Recurrent Neural Network (CRNN): The CRNN model combines four convolutional layers with 16, 32, 64, and 128 filter sizes, followed by two LSTM layers, each with 256 units. This structure is particularly effective for handling both spectral features and temporal sequences.Transformer: The transformer model includes four encoder blocks, each with 16 heads. These blocks use an attention mechanism that helps the model better understand subtle details in complex audio settings.
[0099] Each model configuration includes a 257-unit dense layer for the final output, ensuring robust feature extraction and signal processing. This architecture enables NeuroAMP to effectively integrate the spectral features from the STFT and the detailed information from the prescription input.
[0100] To optimize the model, the Adam optimizer is utilized with the following hyperparameters: a learning rate of 0.0001, pi of 0.9, [32 of 0.999, and e of le-07. The training process is configured with a batch size of 1 and 100 epochs. Early stopping is enabled to prevent overfitting, ensuring the system achieves optimal performance.
[0101] For the Denoising NeuroAMP experiments, the setup largely mirrors that of the NeuroAMP, with a few key distinctions. The input signal is similarly preprocessed with a 512-point STFT using a 32 ms Hamming window and a 16 ms hop. The listener’s audiogram, represented as a 1x6 matrix, is processed through a dense layer and concatenated with the spectral features of the noisy input signal. In the Denoising NeuroAMP, an LSTM-based neural network architecture is specifically employed. The LSTM model includes two layers, each with 256 units. This configuration is chosen for its effectiveness in capturing temporal dependencies in the audio signal, which is crucial for both speech enhancement and personalized amplification. The training target for Denoising NeuroAMP is generated by applying the NAL-R prescription formula combined with WDRC to the clean version of the input audio signal. The Mean Squared Error (MSE) is used as the loss function, calculated between the predicted magnitude spectrogram and the target magnitude spectrogram.
[0102] Similar to NeuroAMP, the Denoising NeuroAMP model is optimized using the Adam optimizer. The hyperparameters are identical: a learning rate of 0.0001, [31 of 0.9, [32 of 0.999, and e of le-07. The model is trained with a batch size of 1 for 100 epochs, and early stopping is used to prevent overfitting.
[0103] Evaluation Results
[0104] The evaluation strategy focuses on assessing the intelligibility and quality of amplified speech processed by the proposed NeuroAMP model and NAL-R+WDRC (baseline). Two established metrics are employed in hearing aid auditory research: HASPI and HASQI for speech,along with the Hearing Aid Audio Quality Index (HAAQI) for music. These metrics are important for quantifying the model’s impact on both speech and music quality. For benchmarking purposes, the metrics derived from NeuroAMP are compared with those obtained from the NAL-R+WDRC processed signal, using three statistical measures: Linear Correlation Coefficient (LCC), Spearman Rank Correlation Coefficient (SRCC), and Mean Squared Error (MSE). Higher LCC values, higher SRCC values, and lower MSE values indicate greater similarity between NeuroAMP and NAL- R+WDRC.
[0105] Furthermore, the relationship between model parameter scale and performance is explored through two experimental settings: a basic parameter configuration for all neural networks to establish a baseline (denoted as basic (same)) and an optimized configuration to evaluate the full potential of each architecture (denoted as optimized).
[0106] 1) Impact of Model Parameter Scale on Performance: Table IV in FIG. 14A summarizes the number of parameters for each model under basic (same) and optimized settings. This comparison aims to evaluate the trade-offs between model complexity and performance.
[0107] Table V in FIG. 14B presents the results under the basic (same) parameter setting. With these constraints, the Transformer model consistently exhibited superior performance across most datasets and metrics. It achieved the highest LCC and SRCC values and the lowest MSE, indicating its inherent architectural efficiency and ability to generalize well even with limited parameters.
[0108] Table VI in FIG. 14C shows the results after optimizing the parameters. With increased parameter counts, almost all models exhibited consistent performance improvements. Notably, the CNN and LSTM models showed substantial gains in their ability to generalize, particularly on the TMHINT dataset. The CRNN model, which initially struggled under the minimal setting, demonstrated a notable improvement and achieved comparable performance after tuning. Furthermore, the Transformer model achieves the best performance and further improves out-of- domain generalization in the unseen VCTK dataset.
[0109] Table VII in FIG. 14D provides a detailed performance evaluation of the CNN, LSTM, CRNN, and Transformer models on the VoiceBank dataset (unseen) under different speech conditions: clean, noisy, and enhanced (using FCN and LSTM). The CRNN model excels in clean and enhanced conditions, achieving high LCC and SRCC scores along with notably low MSE values. The Transformer model also performs well, particularly in clean and enhanced conditions.
[0110] Scatter plots are utilized to visually compare and assess the prediction distribution between the NeuroAMP and the ground truth (NAL-R+WDRC) for various speech datasets. Specifically, two datasets are selected: the TMHINT dataset, which is included in the training data, and the VoiceBank dataset, which is unseen during training. For the evaluation, HASPI and HASQI scoresare used as assessment metrics. FIG. 15 displays scatter plots of the HASPI scores for the TMHINT dataset, separated by each model architecture: (a) CNN, (b) LSTM, (c) CRNN, and (d) Transformer. FIG. 16 shows scatter plots of the HASQI scores for the same dataset. FIG. 17 and FIG. 18 illustrate scatter plots for HASPI and HASQI scores on the VoiceBank dataset, respectively, across all model architectures. Each scatter plot is color-coded to distinguish between different types of data: noisy, enhanced by FCN, enhanced by LSTM, and clean, allowing for a clear visualization of score distributions using HASPI and HASQI scores. The density of points along the diagonal in these plots reflects a high degree of agreement between the model outputs (NeuroAMP) and the ground truth (NAL-R+WDRC), suggesting superior model performance. Specifically, the Transformer and CRNN models exhibit distinctive patterns in these scatter plots. The Transformer model ((d) of FIG. 15 and (d) of FIG. 18) shows a particularly dense clustering of points along the diagonal across all types of data, suggesting robust performance and effective handling of variations in audio quality. Moreover, the CRNN model ((c) of FIG. 15 and (c) of FIG. 18) shows better alignment with the ground truth but exhibits slightly more dispersion, which may be due to the CRNN’s varying ability to handle certain dataset characteristics or noise levels.
[0111] 2) Performance Evaluation on Music Dataset: Table VIII in FIG. 14E presents a comprehensive evaluation of the models using the HAAQI metric to assess music quality against ground-truth scores across two test sets: the Cadenza Challenge and MUSDB18-HQ. LCC, SRCC, and MSE are employed to assess correlation and accuracy. The Cadenza Challenge dataset, which is seen during the training phase, and the unseen MUSDB18-HQ dataset, which is not exposed to the model during training, are used to evaluate the models’ abilities to generalize to new data.
[0112] Notably, the performance metrics across these datasets demonstrate that the models, particularly the Transformer and CRNN architectures, can handle unseen scenarios effectively. This is evident from their notable performances in LCC, SRCC, and particularly low MSE values in the MUSDB18-HQ dataset. For example, the Transformer model achieves an LCC of 0.9888 and an SRCC of 0.9892 on the unseen MUSDB18-HQ dataset, indicating notable prediction performance and correlation with the ground truth, despite having no prior exposure during training. These findings confirm the capability of the models to adapt and perform robustly across different types of audio content, such as speech and music, which inherently possess distinct acoustic characteristics. In addition, the models show good generalization performance on the unseen data. Furthermore, the effectiveness of the models across both seen and unseen datasets highlights their potential for practical applications in various real-world scenarios.
[0113] Next, scatter plots are utilized to visually analyze how well the models predicted HAAQI scores across two different music datasets. In FIG. 19, scatter plots are shown for the Clarity-Cadenza Challenge dataset, showing the results for each model architecture: (a) CNN, (b) LSTM, (c) CRNN, and (d) Transformer. FIG. 20 presents scatter plots of HAAQI scores on the MUSDB18- HQ dataset, providing insights into the performance of the same models on this unseen music dataset. The plots indicate that all models performed well, with the Transformer architecture particularly excelling in handling both noisy and clean data, as evidenced by its densely clustered points along the diagonal line. These visualizations align with the numerical results in Table II and demonstrate the models’ ability to maintain high accuracy and consistency across diverse scenarios, including seen and unseen datasets.
[0114] 3) Waveform Analysis: To further validate the model’s performance, an in-depth analysis is conducted encompassing waveform comparison, spectral analysis, and frequency band energy over time. The unprocessed NAL-R, NAL-R+WDRC, and NeuroAMP amplified signals are compared. FIG. 21 presents the waveform comparison, revealing that the Transformer model’s output closely aligns with the ground truth (NAL-R+WDRC) signal. The similarity in waveform patterns demonstrates the effectiveness of the model in mimicking the ground-truth waveform, as highlighted in the red boxes.
[0115] 4) Spectrogram Analysis: FIG. 22 presents the spectral comparison for unprocessed NAL-R, NAL-R+WDRC, and NeuroAMP amplified signals. The spectral analysis aligns the findings from the waveform comparison, indicating that the Transformer model’s spectral characteristics closely resemble those of the NAL-R+WDRC amplified signal. This consistency across both waveform and spectral domains highlights NeuroAMP’s capability to replicate the performance characteristics of the ground-truth signals.
[0116] 5) Frequency Band Analysis: FIG. 23 illustrates the comparative analysis of frequency band energy over time for unprocessed, NAL-R, NAL-R+WDRC, and NeuroAMP -processed signals. This analysis provides insights into how each processing method affects the time-frequency distribution, focusing on specific frequency bands and their energy distribution. In the low- frequency band (0-500 Hz), the unprocessed signal shows the lowest energy levels with significant fluctuations over time. The NAL-R processed signal exhibits higher energy levels, indicating effective amplification in this region. The NAL-R+WDRC processing maintains these elevated energy levels while adding dynamic range compression, which is visible as slight variations over time. Notably, the NeuroAMP processed signal closely mirrors the energy pattern of the NAL- R+WDRC, demonstrating its ability to effectively replicate low-frequency enhancements.
[0117] In the mid-frequency band (500-2000 Hz), the unprocessed signal shows the lowest energy levels. The NAL-R processing significantly increases energy in this range, enhancing mid frequency components crucial for speech intelligibility. While similar to NAL-R in energy levels,the NAL-R+WDRC processing introduces slight variations due to dynamic range compression. The NeuroAMP model closely aligns with the NALR+WDRC energy distribution, demonstrating a high degree of accuracy in replicating mid-frequency enhancements.
[0118] In the high-frequency band (2000-8000 Hz), the unprocessed signal maintains low energy levels. The NAL-R processed signal shows significant amplification, enhancing high frequency components crucial for speech intelligibility. As with other bands, the NAL-R+WDRC processing introduces minor variations due to compression effects. The NeuroAMP processed signal closely mirrors the NAL-R+WDRC pattern, effectively replicating both high-frequency amplification and compression.
[0119] Overall, this detailed analysis across specific frequency bands confirms that the NeuroAMP model performs comparably to traditional NAL-R and NAL-R+WDRC processing techniques. By effectively preserving the time-frequency energy distribution, the NeuroAMP model closely mirrors the ground truth provided by the NAL-R+WDRC processed signal.
[0120] 6) Gain Function Analysis: Additionally, in the comparative study, gain patterns across different listener audiograms are analyzed to evaluate the performance of various signal processing methods. FIG. 24 presents an analysis of audiograms and frequency-specific gain comparisons for two subjects with different hearing loss profiles. Panels (a) and (c) show the audiograms, illustrating the hearing thresholds of the subjects relative to the normal hearing threshold. Panels (b) and (d) display the frequency-specific gain comparisons for the NAL-R, NAL-R+WDRC, and NeuroAMP -processed signals.
[0121] In panel (a), the audiogram displays the hearing threshold (red line) for the first subject, indicating significant hearing loss at higher frequencies. The normal hearing threshold (green dashed line) serves as a reference, highlighting the extent of hearing loss. Panel (c) shows the audiogram for the second subject, who has more pronounced hearing loss, particularly at higher frequencies, compared to the first subject. Panels (b) and (d) compare the gains provided by different processing methods. In Panel (b), the NAL-R processed gain (blue line) shows substantial amplification at lower frequencies with a gradual decrease towards higher frequencies. The NAL- R+WDRC processed gain (red line) closely follows the NAL-R curve but exhibits a more stable gain distribution across frequencies due to compression effects. The NeuroAMP processed gain (green line with triangles) closely mirrors the NAL-R+WDRC gain, demonstrating NeuroAMP’s effective replication of traditional processing methods.
[0122] In Panel (d), the NAL-R processed gain (blue line) shows a more pronounced increase, particularly at higher frequencies, to address the greater hearing loss shown in Panel (c). The NAL- R+WDRC processed gain (red line) continues to follow the NAL-R curve, with compression effects 1providing a stable gain distribution. The NeuroAMP -processed gain (green line with triangles) closely aligns with the NAL-R+WDRC gain pattern, demonstrating NeuroAMP’s effective replication of the gain characteristics. This analysis highlights that the NeuroAMP model successfully replicates the gain features of NAL-R+WDRC processing across both hearing loss profiles. Its ability to offer appropriate amplification and compression tailored to specific hearing loss profiles suggests it as a viable alternative for hearing aid signal processing. By matching the performance of established methods, the NeuroAMP model ensures comparable speech intelligibility and listening comfort for users, similar to traditional methods
[0123] Denoising NeuroAMP Performance Evaluation
[0124] The performance of the Denoising NeuroAMP model, which combines speech enhancement and personalized amplification in a single unified process, is evaluated. The evaluation focuses on the VoiceBank+DEMAND dataset, a widely used benchmark for speech enhancement tasks. Denoising NeuroAMP is compared to two baselines:NAL-R+WDRC: This represents the conventional approach of applying the NAL-R prescription formula and WDRC directly to the noisy speech signal.- Denoising>NAL-R+WDRC: This is a two-stage baseline that first uses a separate LSTM-based speech enhancement model to enhance noisy speech, and then uses NAL- R+WDRC for amplification.
[0125] HASPI and HASQI are used as evaluation metrics. These metrics are commonly used to assess the intelligibility and quality of processed speech for hearing aid users. Table VI in FIG. 25 lists the HASPI and HASQI scores of Denoising NeuroAMP and the two baselines on the noisy test set of the VoiceBank+DEMAND dataset.
[0126] The results show that Denoising NeuroAMP outperforms both baselines in both HASPI and HASQI scores. Specifically, Denoising NeuroAMP achieves a HASPI score of 0.90 and a HASQI score of 0.59, while NAL-R+WDRC has scores of 0.89 and 0.55, and the two-stage Denoising>NAL-R+WDRC baseline has scores of 0.85 and 0.30. The improvements of Denoising NeuroAMP over Denoising>NAL-R+WDRC (a standard two-stage approach commonly used in hearing aids) are confirmed by a t-test to be statistically significant. These findings suggest that Denoising NeuroAMP, an integrated approach that jointly optimizes speech enhancement and personalized amplification, is more effective than applying these processes separately. The lower performance of the two stage baseline, particularly with respect to HASQI, highlights the potential drawbacks of a disjointed approach, where errors in the speech enhancement stage may negatively impact the subsequent amplification stage. Denoising NeuroAMP’s superior performance confirms its potential to improve speech intelligibility and quality for hearing aid users in noisy environments.By learning a unified denoising and amplification process tailored to an individual’s hearing profile, Denoising NeuroAMP provides an alternative solution to more effective hearing aid technology.
[0127] Subjective Evaluations
[0128] 1) Perceptual Similarity Analysis: a preference test is presented comparing the proposed NeuroAMP model with the traditional NAL-R+WDRC processing. The evaluation involved 20 subjects, evenly split between males and females, all confirmed to have normal hearing through audio metric testing. To simulate hearing loss, the MSBG hearing loss simulator is used, enabling a realistic assessment of processed audio as perceived by individuals with hearing loss. Each subject listened to pairs of audio samples processed by the NAL-R+WDRC and NeuroAMP models, followed by the MSBG hearing loss simulator. They were asked to determine whether sample A was better, sample B was better, or if the samples were similar and indistinguishable. The sample order was randomized, and the tests were conducted in a double blind manner. A diverse range of samples is selected from the test dataset to ensure a comprehensive evaluation, including 10 clean speech samples, 10 noisy speech samples, 10 speech samples enhanced by the FCN model, 10 speech samples enhanced by the LSTM model, 5 clean music samples, and 5 noisy music samples. Participants evaluated each pair based on intelligibility (clarity of speech) and quality (overall audio quality, including naturalness and pleasantness).
[0129] FIG. 26 shows the evaluation results, aggregating preferences and ratings across different sample types. Detailed visual representations of the evaluation results for each sample type -clean speech, noisy speech, enhanced speech by FCN, enhanced speech by LSTM, clean music, and noisy music - are shown in FIG. 27. From FIG. 26, it is observed that “Similar” responses significantly outnumber those for “NAL-R+WDRC” and “NeuroAMP”. FIG. 27 shows a consistent trend: “Similar” decisions prevail across all conditions, including speech, noisy speech, enhanced speech, and music scenarios. These results confirm that the proposed NeuroAMP model effectively simulates NAL-R+WDRC and demonstrates its capability to deploy a deep learning-based amplification approach.
[0130] 2) Mean Opinion Score (MOS): To further quantitatively assess the perceived audio quality of the processed samples, a MOS evaluation involving 30 listeners is conducted. Similar to the previously described Perceptual Similarity Analysis, audio samples were processed using three amplification approaches: the traditional NAL-R+WDRC processing, the proposed NeuroAMP model, and Denoising NeuroAMP. Each audio sample was further processed using the MSBG hearing loss simulator to emulate auditory experiences representative of individuals with hearing impairment. The listeners are requested to evaluate audio samples from three distinct categories: clean speech, noisy speech, and music. Each participant rated the perceived audio quality using aMOS scale from 1 (’’poor” quality) to 5 (’’excellent” quality). The evaluation results are listed in FIG. 28.
[0131] As shown in FIG. 28, the three processing approaches demonstrate comparable performance for clean speech. Specifically, Denoising NeuroAMP can achieve the highest MOS of 3.71, closely followed by NeuroAMP and the conventional NAL-R+WDRC, with MOS values of 3.64 and 3.62, respectively. Next, in noisy speech conditions, Denoising NeuroAMP achieves the best performance with an average MOS of 3.57, demonstrating the effectiveness of its denoising in noisy environments. Furthermore, NeuroAMP and NAL-R + WDRC exhibit comparable performance, with NeuroAMP achieving a higher MOS than NAL-R+WDRC (2.90 vs. 2.82). For the music scenario, the MOS of NeuroAMP and the conventional NAL-R+WDRC are specifically evaluated. The results of the listening test indicate that both approaches achieve closely matched results, with NeuroAMP slightly outperforming NAL-R+WDRC by obtaining a MOS of 2.80 versus 2.77.
[0132] Overall, the MOS evaluation confirms that the proposed NeuroAMP model and the conventional NAL-R+WDRC processing offer similar perceptual quality across various audio categories. Interestingly, the introduction of denoising capabilities through Denoising NeuroAMP significantly improves audio quality in noisy speech scenarios, highlighting the potential for further exploration of such enhancements in deep learning-based amplification frameworks.
[0133] Finally, 7 individuals are recruited with hearing loss to participate in the MOS listening test. Each subject had a specific audiogram, and the evaluation employed the same set of speech samples as used in FIG. 28. Prior to the listening tests, the audiogram of each participant’s target ear are estimated. Based on these audiograms, the speech signals were processed using NAL-R + WDRC, NeuroAMP, and Denoising NeuroAMP, and then presented to the participants without the use of HAs. The resulting MOS scores are shown in FIG. 29. The trends observed in FIG. 29 align with those in FIG. 28. In clean and music conditions, NAL-R + WDRC and NeuroAMP demonstrated comparable performance, with NeuroAMP slightly outperforming NAL-R + WDRC. Under noisy conditions, Denoising NeuroAMP exhibited significant improvements over the other two methods.
[0134] In the embodiment, NeuroAMP is a novel deep learning based system for personalized hearing aid amplification. Denoising NeuroAMP is an extension that integrates noise reduction for better performance in adverse environments. Four neural network architectures: CNN, LSTM, CRNN, and Transformer are investigated, using spectral features and audiograms as inputs. Extensive evaluation using HASPI, HASQI, and HAAQI, together with statistical measures like LCC, SRCC, and MSE, demonstrated the superior performance of the Transformer within the NeuroAMP framework. Notably, NeuroAMP achieved SRCC scores of 0.9810 (HASQI) and0.9914 (HASPI) on the unseen VCTK dataset, and 0.9892 (HAAQI) on the MUSDB18-HQ dataset, validating the effectiveness of data augmentation strategy. Next, Subjective listening tests confirmed that both NeuroAMP produce outputs that are perceptually similar to those of traditional methods while offering better adaptability. Furthermore, Denoising NeuroAMP outperformed conventional NAL-R+WDRC and two-stage baselines on the VoiceBank+DEMAND dataset, highlighting its potential for enhanced speech intelligibility in noisy environments. In conclusion, this work demonstrates the potential of NeuroAMP and Denoising NeuroAMP to perform a personalized, user-centric approach to amplification and noise reductions.
Claims
Claims1. A signal processing method for hearing aid, comprising: receiving an incoming sound signal; providing a listener audiogram; amplifying the incoming sound signal based on a representation of the listener audiogram and features of the incoming sound signal by using a neural network; and outputting an amplified sound signal from the neural network.
2. The method of claim 1, wherein amplifying the incoming sound signal comprises: applying feature extraction on the incoming sound signal by time-to-frequency conversion to output the features of the incoming sound signal; applying a dense layer for receiving the listener audiogram to output the representation of the listener audiogram; and concatenating the features of the incoming sound signal and the representation of the listener audiogram as a comprehensive input of the neural network.
3. The method of claim 2, wherein the listener audiogram is provided according to the incoming sound signal.
4. The method of claim 3, further comprising: training the neural network based on loss function, wherein the loss function is calculated according to an estimated spectral representation of the neural network and a ground-truth spectral representation, the ground-truth spectral representation is obtained by applying a prescription formula and WDRC to the incoming sound signal.
5. The method of claim 2, wherein the listener audiogram is provided according to a clean sound signal, and the incoming sound signal has noise.
6. The method of claim 5, further comprising: training the neural network based on loss function, wherein the loss function is calculated according to an estimated spectral representation of the neural network and a ground-truth spectral representation, the ground-truth spectral representation is obtained by applying a prescription formula and WDRC to the clean sound signal.
7. The method of claim 1, wherein amplifying the incoming sound signal further comprises: applying a final dense layer for the neural network to output amplified features; and converting the amplified features by frequency -to-time conversion to output the amplified sound signal.
8. The method of claim 1, wherein the neural network is constructed with at least one of Convolutional Neural Network (CNN), deep neural network (DNN), long short-term memory (LSTM), convolutional recurrent neural network (CRNN), and Transformer.
9. The method of claim 1, wherein the neural network is integrated with prescription fitting formula and wide dynamic range compression (WDRC).
10. The method of claim 1, wherein the neural network includes neural speech enhancement and neural echo cancellation, thereby enabling end-to-end optimization.
11. A device for hearing aid, comprising: at least one processor; and at least one storage media including a listener audiogram and instructions operable to be executed by the at least one processor to perform a set of operations comprising: receiving an incoming sound signal; amplifying the incoming sound signal based on a representation of the listener audiogram and features of the incoming sound signal by using a neural network; and outputting an amplified sound signal from the neural network.
12. The device of claim 11, wherein amplifying the incoming sound signal comprises: applying feature extraction on the incoming sound signal by time-to-frequency conversion to output the features of the incoming sound signal; applying a dense layer for receiving the listener audiogram to output the representation of the listener audiogram; and concatenating the features of the incoming sound signal and the representation of the listener audiogram as a comprehensive input of the neural network.
13. The device of claim 12, wherein the listener audiogram is provided according to the incoming sound signal.
14. The device of claim 13, wherein the set of operations further comprises: training the neural network based on loss function, wherein the loss function is calculated according to an estimated spectral representation of the neural network and a ground-truth spectral representation, the ground-truth spectral representation is obtained by applying a prescription formula and WDRC to the incoming sound signal.
15. The device of claim 12, wherein the listener audiogram is provided according to a clean sound signal, and the incoming sound signal has noise.
16. The device of claim 15, wherein the set of operations further comprises: training the neural network based on loss function, wherein the loss function is calculated according to an estimated spectral representation of the neural network and a ground-truth spectralrepresentation, the ground-truth spectral representation is obtained by applying a prescription formula and WDRC to the clean sound signal.
17. The device of claim 11, wherein amplifying the incoming sound signal further comprises: applying a final dense layer for the neural network to output amplified features; and converting the amplified features by frequency -to-time conversion to output the amplified sound signal.
18. The device of claim 11, wherein the neural network is constructed with at least one of Convolutional Neural Network (CNN), deep neural network (DNN), long short-term memory (LSTM), convolutional recurrent neural network (CRNN), and Transformer.
19. The device of claim 11, wherein the neural network is integrated with prescription fitting formula and wide dynamic range compression (WDRC).
20. The device of claim 11, wherein the neural network includes neural speech enhancement and neural echo cancellation, thereby enabling end-to-end optimization.
Citation Information
Patent Citations
Speech enhancement method, speech recognition method, speaker recognition method and system
CN116092501A
System and method for acoustic echo cancellation using deep multitask recurrent neural networks
US20200312346A1
Deep neural network based audio processing method, device and storage medium
US20210074266A1
Signal processing in a hearing device
US20230127309A1
Hearing aid comprising a signal processing network conditioned on auxiliary parameters
US20230353958A1