Reinforced learning hearing aid adaptation method combined with hearing loss speech evaluation model

By combining hearing loss simulation model and reinforcement learning algorithm, the hearing aid gain adjustment is optimized, and the hearing aid is insufficient adaptability in different noise environments is solved, and personalized auditory compensation and auditory experience improvement is achieved.

CN120302226APending Publication Date: 2025-07-11EAST CHINA NORMAL UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510491644.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing hearing aid adaptation methods are difficult to accurately adjust according to individual hearing loss characteristics, especially in different noise environments, which affects the patient's hearing experience and quality of life.

Method used

Combining the hearing loss simulation model and reinforcement learning algorithm, the hearing map is obtained through hearing measurement tools, supervised pre-training is used to use the noise data set, and the hearing loss speech evaluation model is used to optimize the hearing aid gain configuration, and the gain strategy is adjusted through the hearing loss speech evaluation model feedback to achieve personalized hearing aid gain adjustment.

Benefits of technology

Accurate auditory compensation in different noise environments is achieved, the adaptation accuracy and auditory experience of hearing aids are improved, especially the speech clarity and comfort in quiet and noisy environments, and the patient's auditory communication ability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302226A_ABST
    Figure CN120302226A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning hearing aid adaptation method in combination with a hearing loss speech evaluation model, and aims to design an optimal hearing aid adaptation gain matched with a listening environment based on the common listening environment of a hearing-impaired person and improve the hearing aid use effect of the hearing-impaired person. An audiogram of a subject is obtained through a hearing measurement tool, the gain configuration of the hearing aid is optimized by using a reinforcement learning algorithm, and finally a personalized gain adjustment scheme is provided for the subject. The method is based on a reinforcement learning algorithm iteration strategy network, the hearing aid gain is automatically adjusted, and it is ensured that optimal hearing experience is provided under different sound pressure levels and noise environments. By means of the hearing loss speech model, the gain adjustment effect can be evaluated in real time, and the speech perception effect is optimized. Compared with a traditional method, the method has the advantages that the common acoustic environment requirements of a patient can be accurately matched, the adaptation precision and effect are improved, and higher adaptability and wide clinical application prospects are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hearing rehabilitation, and particularly to a method for adjusting and adapting the gain of a hearing aid based on individual hearing loss characteristics. Specifically, the present invention combines audiological theory with reinforcement learning algorithms to optimize the gain adjustment of hearing aids, aiming to improve the adaptability and performance of hearing aids in different types of hearing loss and various noise environments. By combining the reinforcement learning algorithm with the speech evaluation model of hearing loss, the present invention can achieve accurate assessment and personalized debugging of patients' hearing loss, thereby providing innovative technical means and theoretical support for clinical applications. Background Art

[0002] Hearing aids, as the most effective hearing intervention devices in the field of hearing rehabilitation, are widely used in various hearing-impaired patients. In recent years, with the continuous progress of technology, the functions and application scopes of hearing aids have been continuously expanded. During the fitting process of hearing aids, audiologists rely on audiograms, speech spectra, and the auditory loudness information of patients to develop a series of regular parameters. These parameters together constitute the hearing aid fitting method, enabling audiologists to more accurately compensate for the hearing loss of each patient, thereby improving the quality of life of patients.

[0003] Hearing loss simulation, as an important tool in psychoacoustics and auditory research, can not only provide researchers with an intuitive method for evaluating the impact of hearing loss, but also effectively simulate the perceptual differences of normal-hearing people when facing different degrees of hearing loss. By making specific adjustments to aspects such as the time-frequency characteristics and volume of sounds, a hearing loss simulator can simulate the auditory perception of hearing-impaired people in complex environments. The application of this technology not only helps to improve the fitting accuracy of hearing aids, but also provides strong support for the research and development of new hearing aids. Here, auditory perception includes multiple dimensions, such as speech intelligibility and loudness. Intelligibility refers to the amount of speech sounds that can be understood. Speech intelligibility is an important factor affecting whether hearing-impaired people are willing to wear hearing aids for a long time in daily life. Loudness refers to the size of the sound heard, which is mainly affected by the hearing threshold, uncomfortable threshold, and loudness growth law, and is also affected by the binaural integration effect.

[0004] Reinforcement learning (RL for short) is a key area in machine learning and an important branch of machine learning. Its core goal is to optimize the action strategy through the interaction between an agent and the environment to maximize the cumulative reward or return. Different from supervised learning that relies on labeled data and unsupervised learning that explores data structures, the unique feature of reinforcement learning is that it can gradually learn the optimal strategy through a trial-and-error process without explicit guidance. This feature makes it particularly suitable for solving decision-making problems in complex dynamic systems, such as the gain adjustment of hearing aids. Summary of the Invention

[0005] The object of the present invention is to provide a reinforcement learning hearing aid fitting method combined with a hearing loss simulation model. This method obtains the audiogram of a patient by using a hearing measurement tool, and calculates the preliminary hearing aid gain value based on the mainstream air conduction fitting formula. Then, it performs supervised pre-training using an audio signal dataset containing different signal-to-noise ratios and noise types, enabling the policy network to quickly obtain a reasonable initial policy. Next, the agent predicts the hearing aid gain through the policy network according to the current state, i.e., the input audio signal and the audiogram, and applies this gain to the audio signal to generate a new state. This process simulates the working mode of a hearing aid in an actual usage scenario. The audio signal after gain is input into the hearing loss speech evaluation model to calculate the auditory perception score after simulating hearing loss, such as the speech intelligibility score. This score is used as a reward to evaluate the effect of the current gain policy. In this way, the reinforcement learning algorithm can guide the policy network to gradually optimize the gain setting to enhance the user's auditory experience. By interacting with the environment multiple times and collecting state-action-reward sequences, the reinforcement learning algorithm updates the parameters of the policy network. Repeating this process, a highly personalized hearing aid gain policy is finally obtained.

[0006] The specific technical steps to achieve the object of the present invention are as follows:

[0007] A method for a reinforcement learning hearing aid fitting model combined with a hearing loss speech evaluation model includes the following specific steps:

[0008] Step 1: Use a hearing measurement tool to measure the hearing of hearing loss subjects in a soundproof booth to obtain the audiogram; based on the mainstream air conduction fitting formula in the hearing aid fitting software, calculate the hearing aid gains (G L0 、G M0 、G H0 ) at multiple frequency points of low, medium, and high sound pressure levels;

[0009] Step 2: Use an open-source noise dataset and speech dataset to synthesize audio signals with different signal-to-noise ratios and different noise types, further generate audio signals of low, medium, and high sound pressure levels, and divide them into a pre-training set for supervised training, a training set for reinforcement learning PPO interaction, a validation set for early stopping judgment and model selection, and a test set only for final evaluation according to the ratio of 6∶2∶1∶1;

[0010] Step 3: Pre-train the policy network; use the pre-training set obtained in Step 2 to pair with the fixed audiogram in Step 1 as the input of the policy network, and the hearing aid gains at multiple frequency points calculated by the mainstream air conduction fitting formula in Step 1 as the true values output by the policy network, and perform supervised pre-training on the policy network through the mean square error loss MSE to initialize the parameters of the policy network;

[0011] Step 4: Fine-tune the policy network based on the Proximal Policy Optimization (PPO) algorithm for reinforcement learning, specifically including steps 4.1 - 4.4;

[0012] Step 4.1: Initialize the training parameters; set the total number of training rounds to 1000 rounds, the early stopping condition is that the average reward of the training environment set does not improve for 20 consecutive rounds, and the batch size B = 32 audio samples / round;

[0013] Step 4.2: Interactive training loop; for each audio in the training set of step 2, perform the following sub-steps:

[0014] Step 4.2.1: State initialization; pair the audiogram construction agent in step 1 to initialize the state s0 = (Audio0, Audiogram);

[0015] Step 4.2.2: Predict actions using the policy network; input s0 into the policy network to predict the hearing aid gain a0 = (G L1 、G M1 、G H1 );

[0016] Step 4.2.3: State transition; apply a0 to Audio0 in step to generate the audio signal Audio1 after gain; at this time, the conversion from state s0 to state s1 is completed, where s1 = (Audio1, Audiogram);

[0017] Step 4.2.4: Reward calculation; input the audio Audio1 after gain and the audiogram Audiogram into the hearing loss speech evaluation model to calculate the auditory perception score, that is, the speech evaluation score r, after simulating hearing loss; this score is used as the reward for the action a0 taken by the agent to evaluate the quality of the hearing aid gain adjustment;

[0018] Step 4.3: Parameter update; collect all interaction trajectories {(s, a, r, s)} in each round, and update the policy network parameters through the reinforcement learning algorithm PPO;

[0019] Step 4.4: Termination condition monitoring; after each round of training, calculate the average reward of the training set to monitor the training progress; calculate the average reward of the validation set for model selection; terminate the training when the early stopping condition is met or N rounds are reached;

[0020] Step 5: Save the policy network model parameters; select the version of the policy network parameters with the highest reward in the validation set for saving;

[0021] Step 6: Use the policy network; input the audiogram and the audio signal in the test set into the fine-tuned policy network to obtain the optimized hearing aid gain (G L2 、G M2 、G H2)。

[0022] Advantages of the present invention:

[0023] 1) By using the audiogram data of patients and combining with the reinforcement learning algorithm, the present invention can automatically adjust the gain configuration of hearing aids according to the characteristics of individual hearing loss. This process is based on the training and optimization of the policy network by the reinforcement learning algorithm, realizing personalized hearing loss compensation and providing a tailor-made solution for each patient. Compared with the traditional debugging method based on experience and standard formulas, the present invention can more accurately match the hearing needs of patients and improve the accuracy and effect of adaptation.

[0024] 2) The present invention fully considers the auditory needs of patients in different noise environments. By synthesizing an audio signal dataset containing various signal-to-noise ratios and simulating the auditory experience of patients at low, medium, and high sound pressure levels, the optimized hearing aid gain configuration is not only applicable to quiet environments but also can effectively cope with complex noise environments. The reinforcement learning agent gradually adjusts the gain according to the environmental feedback to ensure the best auditory experience in various hearing environments. This highly adaptable gain adjustment greatly improves the practical application effect of hearing aids.

[0025] 3) The present invention combines a speech evaluation model for hearing loss. By evaluating the speech intelligibility of the audio signal after adjusting the hearing aid gain, it objectively measures the effect of the hearing aid gain. This evaluation can not only provide real-time feedback on the improvement of speech recognition and clarity after gain adjustment but also optimize the gain configuration in each iteration to ensure that the hearing aid gain adjustment maximally improves the patient's speech comprehension ability, thereby improving the patient's auditory communication experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of the present invention;

[0027] Figure 2 is an architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0029] Embodiment 1:

[0030] A quiet environment such as at home or in a library is one of the common usage scenarios for hearing loss patients. In this environment, the signal-to-noise ratio is relatively high, usually greater than 20 dB. The main challenge is to ensure speech clarity and sound quality comfort while avoiding discomfort caused by excessive amplification.

[0031] Referring to Figure 1 , a reinforcement learning-based hearing aid adaptation method combined with a hearing loss simulation model in this embodiment includes:

[0032] Step 1: Use a hearing measurement tool to measure the hearing of hearing loss subjects in a soundproof booth to obtain an audiogram; based on the mainstream air conduction fitting formula NAL-NL2 in the hearing aid fitting software, calculate the hearing aid gains (G L0 、G M0 、G H0 ) at multiple frequencies under low, medium, and high sound pressure levels;

[0033] Step 2: Use an open-source noise dataset and a speech dataset to synthesize an audio signal with a signal-to-noise ratio of 30 dB to simulate the quiet environment of a home living room, generate audio signals with low, medium, and high sound pressure levels, and divide them into a pre-training set for supervised training, a training set for Proximal Policy Optimization (PPO) interaction, a validation set for early stopping judgment and model selection, and a test set only for final evaluation in a ratio of 6:2:1:1;

[0034] Step 3: Pre-train the policy network. Use the pre-training set obtained in Step 2 to pair with the audiogram of the subject obtained in Step 1 as the input of the policy network, and the hearing aid gains (G L0 、G M0 、G H0 ) at multiple frequencies under the three sound pressure levels calculated by the mainstream air conduction fitting formula NAL-NL2 in Step 1 as the true values output by the policy network, and perform supervised pre-training on the policy network through the mean squared error (MSE) loss to initialize the policy network parameters;

[0035] Step 4: Fine-tune the policy network based on the Proximal Policy Optimization (PPO) algorithm, specifically including Steps 4.1 - 4.4;

[0036] Step 4.1: Initialize the training parameters; set the total number of training rounds to 1000 rounds, the early stopping condition to no improvement in the average reward of the training environment set for 20 consecutive rounds, and the batch size B = 32 audio samples / round;

[0037] Step 4.2: Interactive training loop; for each audio in the training set of Step 2, perform the following sub-steps:

[0038] Step 4.2.1: State initialization; pair with the audiogram in Step 1 to construct the initial state of the agent s0=(Audio0, Audiogram);

[0039] Step 4.2.2: The policy network predicts an action; input s0 into the policy network and predict the hearing aid gain a0=(G L1 、G M1 、G H1 );

[0040] Step 4.2.3: State Transition; Apply a0 to step Audio0 to generate the post-gain audio signal Audio1; At this point, the transition from state s0 to state s1 is completed, where s1 = (Audio1, Audiogram).

[0041] Step 4.2.4: Reward Calculation; Input the post-gain audio Audio1 and the audiogram Audiogram into the hearing loss speech assessment model to calculate the auditory perception score, i.e., the speech assessment score r, after simulating hearing loss; This score serves as the reward for the action a0 taken by the agent and is used to evaluate the quality of the hearing aid gain adjustment.

[0042] Step 4.3: Parameter Update; Collect all interaction trajectories {(s, a, r, s)} in each round and update the policy network parameters through the reinforcement learning algorithm PPO.

[0043] Step 4.4: Termination Condition Monitoring; After each round of training, calculate the average reward of the training set to monitor the training progress; Calculate the average reward of the validation set for model selection; Terminate the training when the early stopping condition is met or when N rounds are reached.

[0044] Step 5: Save the Policy Network Model Parameters; Select the version of the policy network parameters with the highest reward in the validation set for saving.

[0045] Step 6: Use the Policy Network; Input the audiogram and the audio signal in the test set into the fine-tuned policy network to obtain the optimized hearing aid gain (G L2 、G M2 、G H2 ).

[0046] Step 7: Process the audio signals in the test set separately using the hearing aid gains (G L0 、G M0 、G H0 ) obtained from the mainstream air conduction fitting formula NAL-NL2 in the hearing aid fitting software in Step 1 and the hearing aid gains (G L2 、G M2 、G H2 ) predicted by the fine-tuned personalized policy network in Step 6, and let the subjects compare and evaluate the auditory perception quality of the two methods, including clarity, comfort, and overall satisfaction.

[0047] Example 2:

[0048] Noisy environments such as streets or wet markets pose another major challenge for hearing loss patients. In such scenarios, the signal-to-noise ratio is low (usually less than 10 dB), the background noise is complex and dynamically changing, and the main goal is to improve speech intelligibility and suppress background noise.

[0049] A reinforcement learning hearing aid fitting method combined with a hearing loss simulation model in this embodiment includes:

[0050] Step 1: Use a hearing measurement tool to measure the hearing of hearing loss subjects in a soundproof booth to obtain an audiogram; based on the mainstream air conduction fitting formula DSP-v5 in the hearing aid fitting software, calculate the hearing aid gains at multiple frequencies under low, medium, and high sound pressure levels (G L0 、G M0 、G H0 );

[0051] Step 2: Utilize an open-source noise dataset and speech dataset to synthesize audio signals with signal-to-noise ratios of 0 dB and -5 dB to simulate a noisy environment such as a street or a vegetable market, generate audio signals at low, medium, and high sound pressure levels, and divide them into a pre-training set for supervised training, a training set for reinforcement learning PPO interaction, a validation set for early stopping judgment and model selection, and a test set only for final evaluation according to the ratio of 6:2:1:1;

[0052] Step 3: Pre-train the policy network. Use the pre-training set obtained in Step 2 to pair with the audiogram of the subject obtained in Step 1 as the input of the policy network, and the hearing aid gains at multiple frequencies under three sound pressure levels (G L0 、G M0 、G H0 ) calculated through the mainstream air conduction fitting formula DSP-v5 in Step 1 as the true values of the policy network output, and perform supervised pre-training on the policy network through the mean square error loss MSE to initialize the policy network parameters;

[0053] Step 4: Fine-tune the policy network based on the reinforcement learning PPO algorithm, specifically including Steps 4.1 - 4.4;

[0054] Step 4.1: Initialize the training parameters; set the total number of training rounds to 1000 rounds, the early stopping condition is that the average reward of the training environment set has not improved for 20 consecutive rounds, and the batch size B = 32 audio samples / round;

[0055] Step 4.2: Interactive training loop; for each audio in the training set of Step 2, execute the following sub-steps:

[0056] Step 4.2.1: State initialization; pair with the audiogram in Step 1 to construct the initial state of the agent s0=(Audio0, Audiogram);

[0057] Step 4.2.2: Policy network predicts actions; input s0 into the policy network to predict the hearing aid gains a0=(G L1 、G M1 、G H1 );

[0058] Step 4.2.3: State Transition; Apply a0 to step Audio0 to generate the gain-adjusted audio signal Audio1; At this point, the transition from state s0 to state s1 is completed, where s1 = (Audio1, Audiogram);

[0059] Step 4.2.4: Reward Calculation; Input the gain-adjusted audio Audio1 and the audiogram Audiogram into the hearing loss speech assessment model to calculate the auditory perception score after simulating hearing loss, i.e., the speech assessment score r; This score serves as the reward for the action a0 taken by the agent and is used to evaluate the quality of the hearing aid gain adjustment;

[0060] Step 4.3: Parameter Update; In each round, collect all interaction trajectories {(s, a, r, s)}, and update the parameters of the policy network through the reinforcement learning algorithm PPO;

[0061] Step 4.4: Termination Condition Monitoring; After each round of training, calculate the average reward of the training set to monitor the training progress; Calculate the average reward of the validation set for model selection; Terminate the training when the early stopping condition is met or when N rounds are reached;

[0062] Step 5: Save the Parameters of the Policy Network Model; Select the version of the policy network parameters with the highest reward in the validation set for saving;

[0063] Step 6: Use the Policy Network; Input the audiogram and the audio signal in the test set into the fine-tuned policy network to obtain the optimized hearing aid gain (G L2 、G M2 、G H2 ).

[0064] Step 7: Use the hearing aid gains (G L0 、G M0 、G H0 ) obtained from the mainstream air conduction fitting formula DSP-v5 in the hearing aid fitting software in step 1 and the hearing aid gains (G L2 、G M2 、G H2 ) predicted by the fine-tuned personalized policy network in step 6 to process the audio signals in the test set respectively, and let the subjects compare and evaluate the auditory perception quality of the two methods, including speech intelligibility, background noise suppression effect, and overall satisfaction.

Claims

1. A reinforcement learning hearing aid adaptation method combined with a speech evaluation model for hearing loss, characterized in that, It includes the following specific steps: Step 1: Use a hearing measurement tool to measure the hearing of hearing loss subjects in an audiometry room to obtain an audiogram; based on the mainstream air conduction fitting formula in the hearing aid fitting software, calculate the hearing aid gains (G L0 , G M0 , G H0 ) at multiple frequency points of low, medium, and high sound pressure levels; Step 2: Using an open-source noise dataset and a speech dataset, synthesize audio signals with different signal-to-noise ratios and different noise types, further generate audio signals with low, medium, and high sound pressure levels, and divide them into a pre-training set for supervised training, a training set for Proximal Policy Optimization (PPO) interaction in reinforcement learning, a validation set for early stopping judgment and model selection, and a test set only for final evaluation according to the ratio of 6:2:1:1; Step 3: Pre-train the policy network; use the pre-training set obtained in Step 2 to pair with the fixed audiogram in Step 1 as the input of the policy network, and the hearing aid gains at multiple frequency points calculated by the mainstream air conduction fitting formula in Step 1 as the true values output by the policy network, and perform supervised pre-training on the policy network through the mean squared error (MSE) loss to initialize the policy network parameters; Step 4: Fine-tune the policy network based on the Proximal Policy Optimization (PPO) algorithm in reinforcement learning, specifically including Steps 4.1 - 4.4; Step 4.1: Initialize the training parameters; set the total number of training rounds to 1000 rounds, the early stopping condition is that the average reward of the training environment set has no improvement for 20 consecutive rounds, and the batch size B = 32 audio samples / round; Step 4.2: Interactive training loop; for each audio in the training set of Step 2, execute the following sub-steps: Step 4.2.1: State initialization; Pair the audiogram in Step 1 to construct the initial state of the agent s0=(Audio0, Audiogram); Step 4.2.2: The policy network predicts an action; input s0 into the policy network, and the predicted hearing aid gain is a0 = (G L1 , G M1 , G H1 ); Step 4.2.3: State transition; Apply a0 to Audio0 in Step to generate the post-gain audio signal Audio1; At this time, the transition from state s0 to state s1 is completed, where s1=(Audio1, Audiogram); Step 4.2.4: Reward calculation; Input the post-gain audio Audio1 and the audiogram Audiogram into the hearing loss speech evaluation model to calculate the auditory perception score after simulating hearing loss, that is, the speech evaluation score r; This score is used as the reward for the action a0 taken by the agent to evaluate the quality of the hearing aid gain adjustment; Step 4.3: Parameter update; Collect all interaction trajectories {(s, a, r, s)} in each round, and update the policy network parameters through the Proximal Policy Optimization (PPO) algorithm in reinforcement learning; Step 4.4: Termination condition monitoring; After each round of training, calculate the average reward of the training set to monitor the training progress; calculate the average reward of the validation set for model selection; terminate the training when the early stopping condition is met or N rounds are reached; Step 5: Save the policy network model parameters; Select the version of the policy network parameters with the highest reward in the validation set for saving; Step 6: Use the policy network; input the audiogram and the audio signals in the test set into the fine-tuned policy network to obtain the optimized hearing aid gain (G L2 , G M2 , G H2 ).

Citation Information

Cited By

  • Auditory speech rehabilitation training method fusing semantic noise control and difficulty self-adaption

    CN121506175A