Audio processing method and hearing aid device
By using a signal processing model and adaptive algorithm trained through simulated audio, hearing aids can effectively reduce noise and eliminate acoustic feedback in complex noisy environments, thereby improving audio quality and solving the problem of poor audio quality in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANKER INNOVATIONS TECH CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
Smart Images

Figure CN122093718A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method and a hearing aid device. Background Technology
[0002] Hearing aids are electronic devices used to pick up and amplify audio in a usage scenario to provide users with enhanced audio. In related technologies, the audio processing of hearing aids is not only affected by the audio transmission between the microphone and speaker within the device, but also cannot adapt to the complex and ever-changing noise levels of the usage scenario, resulting in poor audio quality provided by the hearing aids. Summary of the Invention
[0003] Therefore, it is necessary to provide an audio processing method and a hearing aid device to address the aforementioned technical problems.
[0004] In a first aspect, this application provides an audio processing method, including:
[0005] Acquire the raw audio signal captured by the microphone;
[0006] The original audio signal is input into the signal processing model for noise reduction and feedback elimination to obtain the model output signal. The signal processing model is obtained using simulated audio, which is obtained through closed-loop simulation and includes noise and feedback.
[0007] The audio to be played is determined based on the model's output signal.
[0008] In one embodiment, determining the audio to be played based on the model output signal includes:
[0009] The model output signal is shaped to obtain the time-frequency domain signal to be used;
[0010] Perform an inverse short-time Fourier transform on the frequency domain signal to be used to obtain the audio to be played.
[0011] In one embodiment, the model output signal includes at least two channels of time-frequency domain signals; the model output signal is shaped to obtain the time-frequency domain signal to be used, including:
[0012] An adaptive feedback cancellation algorithm is used to eliminate feedback noise in the time-frequency domain signals of each channel, resulting in feedback cancellation signals for each channel's time-frequency domain signals.
[0013] Beamforming fusion is performed on all feedback cancellation signals to obtain a single-channel enhanced signal;
[0014] The signal gain of the single-channel enhanced signal is adjusted to obtain the time-frequency domain signal to be used.
[0015] In one embodiment, signal gain adjustment includes at least one of the following processes:
[0016] Reduce the signal gain of the noise band in the single-channel enhanced signal and increase the signal gain of the speech band in the single-channel enhanced signal;
[0017] Compress the gain of a single-channel enhanced signal that is greater than the gain threshold to the gain threshold;
[0018] The signal gain of a single-channel enhanced signal can be changed by using random gain coefficients.
[0019] In one embodiment, the training process of the signal processing model includes:
[0020] Obtain the noisy time-frequency domain signal of the noiseless sample audio;
[0021] The simulated audio is determined based on the noisy time-frequency domain signal;
[0022] The signal processing model was trained using noiseless sample audio and simulated audio.
[0023] In one embodiment, acquiring the noisy time-frequency domain signal of the noiseless sample audio includes:
[0024] Noise-free sample audio and preset noisy audio are fused to obtain noisy sample audio;
[0025] Perform a Fourier transform on the noisy audio sample to obtain the noisy time-frequency domain signal.
[0026] In one embodiment, determining the simulated audio based on the noisy time-frequency domain signal includes:
[0027] The noisy time-frequency domain signal is shaped to obtain the noisy time-frequency domain signal to be used.
[0028] Acoustic feedback is added to the noisy time-frequency domain signal to be used to obtain simulated audio.
[0029] In one embodiment, acoustic feedback is added to the noisy time-frequency domain signal to be used to obtain simulated audio, including:
[0030] A transfer function is randomly selected from the set of transfer functions; the transfer function is used to characterize the conversion relationship between the output audio of the speaker and the output audio of the speaker captured by the microphone.
[0031] The transfer function is used to convolve the noisy time-frequency domain signal to be used, and the simulated audio is determined based on the signal after convolution.
[0032] In one embodiment, a signal processing model is trained using noiseless sample audio and simulated audio, including:
[0033] Obtain the noiseless time-frequency domain signal of the noiseless sample audio and the simulated time-frequency domain signal of the simulated audio;
[0034] The initial signal processing model is trained using the noiseless time-frequency domain signal and the simulated time-frequency domain signal to obtain the signal processing model.
[0035] In one embodiment, an initial signal processing model is trained based on a noiseless time-frequency domain signal and a simulated time-frequency domain signal to obtain a signal processing model, including:
[0036] Obtain the noiseless time-frequency domain signal and then shape it to obtain the noiseless frequency to be used;
[0037] The simulated time-frequency domain signal is input into the initial signal processing model for noise reduction and feedback noise elimination, and the model output signal of the initial signal processing model is obtained and shaped to obtain the simulated audio to be used.
[0038] The parameters of the initial signal processing model are adjusted based on the noise-free frequency to be used and the simulated audio to be used, and the signal processing model is obtained.
[0039] In one embodiment, the parameters of an initial signal processing model are adjusted based on the noise-free frequency to be used and the simulated audio to be used, to obtain a signal processing model, including:
[0040] The target audio is obtained by adding preset noise audio to the noiseless frequency to be used by random signal-to-noise ratio parameters;
[0041] Adjust the model parameters in the initial signal processing model according to the loss function between the target audio and the simulated audio to be used until the loss function is minimized, and obtain the signal processing model.
[0042] Secondly, this application also provides a hearing aid device, including a controller, a microphone, and a speaker, wherein the controller is used to implement the steps of any of the above-described audio processing methods.
[0043] In the aforementioned audio processing method and hearing aid device, the raw audio signal collected by the microphone is acquired, and a signal processing model is used to denoise and eliminate feedback sound to obtain the model output signal. The audio to be played is then determined based on the model output signal. The signal processing model is trained using simulated audio, which is obtained through closed-loop simulation and includes noise and feedback sound. In this method, the signal processing model trained using simulated audio denoises and eliminates feedback sound from the raw audio signal, achieving AI-based audio processing. This improves the adaptability of the processing to complex scenarios, reduces noise and acoustic feedback in the obtained audio, achieves multi-task optimization, and consequently improves audio quality. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating an audio processing method in one embodiment;
[0045] Figure 2 This is a schematic diagram of the process of obtaining the audio to be played in one embodiment;
[0046] Figure 3 This is a schematic diagram of the process for obtaining the time-frequency domain signal to be used in one embodiment;
[0047] Figure 4 This is a flowchart illustrating the training process of a signal processing model in one embodiment.
[0048] Figure 5 This is a schematic diagram of the process for obtaining a noisy time-frequency domain signal in one embodiment;
[0049] Figure 6 This is a schematic diagram of the process for obtaining simulated audio in one embodiment;
[0050] Figure 7 This is a schematic diagram of the process for obtaining simulated audio in another embodiment;
[0051] Figure 8 This is a flowchart illustrating the process of training a signal processing model in another embodiment;
[0052] Figure 9 This is a flowchart illustrating the process of training a signal processing model in another embodiment;
[0053] Figure 10 This is a flowchart illustrating the process of training a signal processing model in another embodiment;
[0054] Figure 11 This is a spectrogram of the target audio in one embodiment;
[0055] Figure 12 This is a spectrum diagram of the audio signal exhibiting feedback in one embodiment;
[0056] Figure 13 This is a spectrogram of the audio to be played in one embodiment;
[0057] Figure 14 This is a flowchart illustrating an audio processing method in one embodiment;
[0058] Figure 15 This is a structural block diagram of an audio processing device in one embodiment. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0060] This application provides an audio processing method applied to a hearing aid device including a controller, a microphone, and a speaker.
[0061] Due to the usage scenario, the audio received by hearing aids usually contains noise. Furthermore, part of the sound from the audio collected by the microphone, after being amplified by the speaker, will be transmitted through space to the microphone again, becoming the input signal once more. It will then be repeatedly amplified by the speaker, causing resonance at certain frequencies in the audio and generating acoustic feedback. Severe acoustic feedback can produce howling.
[0062] Therefore, the presence of noise and acoustic feedback significantly reduces the quality of the audio provided by hearing aids.
[0063] Based on this, to improve audio quality, in one embodiment, this application provides an audio processing method applied to a controller in a hearing aid device. For example... Figure 1 As shown, the provided audio processing method includes the following steps:
[0064] S110: Acquire the raw audio signal captured by the microphone.
[0065] The raw audio signal captured by the microphone is a single frame audio. Optionally, the raw audio signal can be a single frame audio including scene noise directly captured by the microphone, or a single frame audio including scene noise and acoustic feedback captured by the microphone after being amplified by the speaker.
[0066] Optionally, the controller can acquire the audio captured by the microphone and use each frame of the audio as the raw audio signal.
[0067] S120. The original audio signal is denoised and feedback is eliminated using a signal processing model to obtain the model output signal. The signal processing model is trained using simulated audio, which is obtained through closed-loop simulation and includes noise and feedback.
[0068] In this context, noise reduction and feedback cancellation processing refers to reducing / removing noise and acoustic feedback from a signal. The signal processing model is a neural network (NN) model used to implement noise reduction and feedback cancellation processing. It can be trained using a large number of simulated audio samples, including noise and feedback, as training samples. For example, this signal processing model can be a deep learning model.
[0069] Optionally, the controller can perform a short-time fourier transform (STFT) on the acquired raw audio signal to convert the raw audio signal from the time domain to the signal processing model carried by the hearing aid. The signal processing model is then used to perform noise reduction and feedback noise elimination on the input time-frequency domain signal to obtain the model output signal.
[0070] For example, the controller can use asymmetric window technology to perform STFT on the original audio signal, that is, to use a longer analysis window and a shorter synthesis window for STFT. The long analysis window ensures frequency resolution, while the short synthesis window reduces the algorithm delay introduced by the transform.
[0071] S130. Determine the audio to be played based on the model output signal.
[0072] Optionally, after obtaining the model output signal of the signal processing model, the controller can perform an inverse short-time fourier transform (ISTFT) on the model output signal to convert the model output signal from the time-frequency domain to the time domain, thereby obtaining the audio to be played.
[0073] For example, similar to STFT, the controller can use asymmetric windowing to perform ISTFT on the model output signal, that is, to perform ISTFT using a longer analysis window and a shorter synthesis window.
[0074] In this embodiment, the raw audio signal acquired by the microphone is processed by an input signal processing model to reduce noise and eliminate feedback, resulting in a model output signal. The audio to be played is then determined based on this output signal. The signal processing model is trained using simulated audio, which is obtained through closed-loop simulation and includes noise and feedback. This method utilizes a signal processing model trained on simulated audio to reduce noise and eliminate feedback from the raw audio signal, achieving AI-based audio processing. This improves the adaptability of the processing to complex scenarios, reduces noise and acoustic feedback in the audio, achieves multi-task optimization, and consequently improves audio quality.
[0075] To further improve audio quality, in one embodiment, such as Figure 2 As shown, the above-mentioned S130, determining the audio to be played based on the model output signal, includes:
[0076] S210. The model output signal is shaped to obtain the time-frequency domain signal to be used.
[0077] The shaping process is used to adjust the signal waveform. For example, the shaping process includes, but is not limited to, feedback noise elimination, beamforming fusion, and signal gain adjustment. The model output signal is a time-frequency domain signal.
[0078] Optionally, the controller can perform shaping processing on the model output signal of the signal processing model to obtain a processed time-frequency domain signal, which can be used as the time-frequency domain signal to be used. For example, taking the shaping processing as an example of eliminating feedback noise, the controller can use an adaptive feedback cancellation (AFC) algorithm to eliminate feedback noise in the model output signal to obtain the time-frequency domain signal to be used.
[0079] It should be noted that when using the AFC algorithm to eliminate feedback noise, not only the audio of each frame is needed, but also the audio of the previous frame after shaping, which is then used as reference information after delay compensation, and processed together with the audio of the current frame.
[0080] S220. Perform inverse short-time Fourier transform on the frequency domain signal to be used to obtain the audio to be played.
[0081] Optionally, after obtaining the shaped time-frequency domain signal to be used, the controller can perform ISTFT on the signal to convert it from the time-frequency domain to the time domain, thus obtaining the audio to be played. For example, asymmetric window technology can be used to perform ISTFT on the signal to be used.
[0082] In this embodiment, the model output signal is shaped to obtain the time-frequency domain signal to be used. Then, an inverse short-time Fourier transform is performed on the time-frequency domain signal to obtain the audio to be played. In the above method, the model output signal of the signal processing model is shaped, enabling further algorithmic processing after the artificial intelligence audio processing, thereby further improving the audio quality.
[0083] Hearing aids typically employ dual microphones on one side, or more microphones, and the raw audio signal includes at least two channels of audio signal, with the corresponding model output signal including at least two channels of time-frequency domain signal. Based on this, in one embodiment, such as... Figure 3 As shown, S210 above, which shapes the model output signal to obtain the time-frequency domain signal to be used, includes:
[0084] S310. An adaptive feedback cancellation algorithm is used to eliminate feedback noise in the time-frequency domain signals of each channel, resulting in feedback cancellation signals for the time-frequency domain signals of each channel.
[0085] Optionally, the controller can employ a preset adaptive feedback cancellation algorithm to eliminate feedback noise in the time-frequency domain signals of each channel, thereby obtaining feedback cancellation signals for each channel's time-frequency domain signals. For example, the AFC algorithm can be the Least Mean Square (LMS) algorithm or the Recursive Least Squares (RLS) algorithm.
[0086] S320. Beamforming and fusing all feedback cancellation signals to obtain a single-channel enhanced signal.
[0087] Optionally, after obtaining the feedback cancellation signal of each channel's time-frequency domain signal, the controller can perform beamforming (BF) fusion on all feedback cancellation signals to synthesize a highly sensitive directional beam, thus obtaining a single-channel enhanced signal.
[0088] S330: Adjust the signal gain of the single-channel enhancement signal to obtain the signal to be used.
[0089] Signal gain adjustment is used to change the gain of different frequency bands of the signal.
[0090] Optionally, after obtaining the single-channel enhanced signal, the controller can adjust the signal gain of the single-channel enhanced signal to obtain the time-frequency domain signal to be used.
[0091] In an optional embodiment, signal gain adjustment includes at least one of the following processes:
[0092] Reduce the signal gain of the noise band in the single-channel enhanced signal and increase the signal gain of the speech band in the single-channel enhanced signal;
[0093] Compress the gain of a single-channel enhanced signal that is greater than the gain threshold to the gain threshold;
[0094] The signal gain of a single-channel enhanced signal can be changed by using random gain coefficients.
[0095] For example, the random gain coefficient ranges from [-3, 1], with 0 as the critical stable gain. When it is greater than 0, a whistling sound will occur.
[0096] Optionally, after obtaining the single-channel enhanced signal, the controller can activate a Wide Dynamic Range Compression (WDRC) strategy to reduce the signal gain of the noise frequency band in the single-channel enhanced signal and increase the signal gain of the voice frequency band in the single-channel enhanced signal. After obtaining the signal adjusted based on the WDRC strategy, the controller can read a predetermined gain threshold, compare the signal gain of each frequency band of the signal with the gain threshold, and compress the signal gain of the signal that is greater than the gain threshold to the gain threshold. After obtaining the signal adjusted based on the gain threshold, the controller can obtain a random gain coefficient to change the signal gain of the signal.
[0097] In this embodiment, when the model output signal includes at least two channels of time-frequency domain signals, an adaptive feedback cancellation algorithm is used to eliminate feedback noise in each channel's time-frequency domain signal, obtaining feedback-cancelled signals for each channel's time-frequency domain signal. All feedback-cancelled signals are then beamformed and fused to obtain a single-channel enhanced signal. Finally, the signal gain of the single-channel enhanced signal is adjusted to obtain the desired time-frequency domain signal. In the above method, the adaptive feedback cancellation algorithm, beamforming, and signal gain adjustment achieve the shaping of the model output signal. The algorithm processing is superimposed after artificial intelligence, simultaneously suppressing audio howling and improving audio quality.
[0098] The training process of the signal processing model will be described next. In one embodiment, such as... Figure 4 As shown, the above method also includes:
[0099] S410: Obtain the noisy time-frequency domain signal of the noiseless sample audio.
[0100] The noise-free sample audio can be pure audio without noise, or audio with a noise level that meets the requirements (e.g., low noise level). Like the original audio signal, the noise-free sample audio is also a single frame audio.
[0101] Optionally, the controller can acquire audio collected by the microphone in a noise-free environment and use each frame of that audio as a noise-free sample audio. Alternatively, it can receive audio recorded by the user and similarly use each frame of that audio as a noise-free sample audio. After obtaining the noise-free sample audio, the controller can add noise to it and then perform STFT on the noise-added noise-free sample audio to obtain the noisy time-frequency domain signal of the noise-free sample audio.
[0102] S420. Determine the simulated audio based on the noisy time-frequency domain signal.
[0103] The simulated audio includes noise and acoustic feedback.
[0104] Optionally, after the controller obtains the noisy time-frequency domain signal of the noiseless sample audio, it can add acoustic feedback to the noisy time-frequency domain signal through an algorithm to simulate the acoustic feedback generated by the transmission of audio between the speaker and the microphone, and determine the simulated audio based on the noisy time-frequency domain signal after adding acoustic feedback.
[0105] S430: The signal processing model is obtained by training with noiseless sample audio and simulated audio.
[0106] Optionally, after obtaining the simulated audio, the controller can use noiseless sample audio and the corresponding simulated audio as training sample pairs, and train the initial signal processing model through multiple sets of training sample pairs until the training cutoff condition is met, thus obtaining the signal processing model.
[0107] In this embodiment, a noisy time-frequency domain signal of noiseless sample audio is acquired, and simulated audio is determined based on the noisy time-frequency domain signal. A signal processing model is then trained using the noiseless sample audio and the simulated audio. In this method, the signal processing model is trained based on the simulated audio obtained from the simulated sound feedback, enabling AI-based audio processing. This improves the adaptability of the processing to complex scenarios, enhances the audio processing effect and stability, and consequently improves the audio quality.
[0108] To obtain a noisy time-frequency domain signal, in one embodiment, such as Figure 5 As shown, the above-mentioned S410, acquiring the noisy time-frequency domain signal of the noiseless sample audio, includes:
[0109] S510: Fuse the noiseless sample audio and the preset noisy audio to obtain the noisy sample audio.
[0110] The preset noise audio is a pre-selected audio file containing scene noise, and it is also a single-frame audio file. Optionally, the preset noise audio can be any frame of scene noise audio collected in different usage scenarios.
[0111] Optionally, the controller can read pre-stored preset noise audio and fuse it with noise-free sample audio to obtain noisy sample audio. The preset noise audio to be fused can be the same or different for the noise-free sample audio from different frames.
[0112] S520. Perform a short-time Fourier transform on the noisy sample audio to obtain the noisy time-frequency domain signal.
[0113] Optionally, after obtaining the noisy sample audio, the controller can perform STFT on the noisy sample audio to obtain the noisy time-frequency domain signal of the noiseless sample audio.
[0114] In this embodiment, noisy sample audio is obtained by fusing noiseless sample audio with preset noisy audio, and a noisy time-frequency domain signal is obtained by performing a short-time Fourier transform on the noisy sample audio. In the above method, noise is added to the noiseless sample audio by fusing noiseless sample audio with preset noisy audio. The audio fusion process is simple and easy to implement in a program, thereby improving the efficiency of obtaining the noisy time-frequency domain signal and simultaneously improving the overall audio processing efficiency.
[0115] To obtain simulated audio, in one embodiment, such as Figure 6 As shown, the above-mentioned S420, determining the simulated audio based on the noisy time-frequency domain signal, includes:
[0116] S610. The noisy time-frequency domain signal is shaped to obtain the noisy time-frequency domain signal to be used.
[0117] The shaping process is used to adjust the signal waveform. For example, the shaping process includes, but is not limited to, eliminating feedback noise, beamforming fusion, and signal gain adjustment.
[0118] Optionally, the controller can perform shaping processing on the noisy time-frequency domain signal to obtain a processed signal, which can be used as the noisy time-frequency domain signal to be used. For example, taking the shaping processing as including feedback noise elimination, beamforming fusion, and signal gain adjustment, the controller can use the AFC algorithm to eliminate feedback noise from the noisy time-frequency domain signal, then perform BF fusion on the signal after feedback noise elimination, then reduce the signal gain of the noise frequency band and increase the signal gain of the speech frequency band based on the WDRC strategy for the fused signal, then compress the signal gain of the signal adjusted based on the WDRC strategy that is greater than the gain threshold to the gain threshold according to a preset gain threshold, and finally change the signal gain of the signal through random gain coefficients, and use the processed signal as the noisy time-frequency domain signal to be used.
[0119] S620 adds acoustic feedback to the noisy time-frequency domain signal to obtain simulated audio.
[0120] Optionally, after obtaining the noisy time-frequency domain signal to be used, the controller can add acoustic feedback to the noisy time-frequency domain signal to be used through an algorithm to obtain simulated audio including noise and acoustic feedback.
[0121] Adding acoustic feedback can be implemented based on pre-determined transfer functions corresponding to different usage scenarios. Therefore, in an optional embodiment, such as... Figure 7 As shown, S620 above adds acoustic feedback to the noisy time-frequency domain signal to be used, and obtains the simulated audio.
[0122] S710. Randomly select a transfer function from the set of transfer functions; the transfer function is used to characterize the conversion relationship between the output audio of the speaker and the output audio of the speaker collected by the microphone.
[0123] The transfer function set includes multiple transfer functions, each derived from the speaker output audio and the speaker output audio captured by the microphone under different usage scenarios of the hearing aid. Specifically, the transfer function can be calculated using the Wiener filtering algorithm based on the speaker output audio and the microphone-captured speaker output audio. For example, different usage scenarios for the hearing aid may include tight-fitting artificial head, loose-fitting artificial head, tight-fitting with a real person, loose-fitting with a real person, handheld hearing aid, and hearing aid suspended in the air.
[0124] Optionally, when adding acoustic feedback to the noisy time-frequency domain signal corresponding to each frame of noiseless sample audio, a transfer function can be randomly selected from the set of transfer functions. For example, when the noiseless sample audio is dual-channel audio, the transfer function also corresponds to a dual-channel transfer function, that is, each transfer function includes two sets of transfer functions.
[0125] S720 uses a transfer function to convolve the noisy time-frequency domain signal to be used, and obtains the simulated audio based on the convolved signal.
[0126] Optionally, after randomly obtaining the transfer function, the controller can use the obtained transfer function to perform convolution processing on the noisy time-frequency domain signal to be used, and then perform ISTFT on the convolutional signal to obtain the simulated audio with added acoustic feedback.
[0127] In this embodiment, a noisy time-frequency domain signal is shaped to obtain a noisy time-frequency domain signal to be used, and acoustic feedback is added to the noisy time-frequency domain signal to obtain simulated audio. Specifically, a transfer function can be randomly selected from the transfer function set to perform convolution processing on the noisy time-frequency domain signal to be used, and simulated audio is obtained from the convolutional signal. The transfer function is used to characterize the conversion relationship between the speaker's output audio and the speaker's output audio collected by the microphone. In the above method, the transfer function set includes transfer functions corresponding to different usage scenarios. Randomly selecting a transfer function from the transfer function set to add acoustic feedback results in simulated audio with added acoustic feedback. This improves the diversity and richness of the source scenarios of the simulated audio, resulting in simulated audio that is closer to the real scene. This helps to improve the adaptability of the trained model to complex scenarios, improve the robustness and adaptability of the trained model, and the random selection of the transfer function can also simulate the sudden change of the transfer function during actual audio processing, which can simultaneously improve the reliability of the trained model.
[0128] To train a signal processing model, in one embodiment, such as Figure 8 As shown, the signal processing model obtained by training S430 using noiseless sample audio and simulated audio includes:
[0129] S810: Obtain the noiseless time-frequency domain signal of the noiseless sample audio and the simulated time-frequency domain signal of the simulated audio.
[0130] Optionally, the controller performs STFT on the noiseless sample audio to obtain the noiseless time-frequency domain signal of the noiseless sample audio. Similarly, the controller performs STFT on the simulated audio to obtain the simulated time-frequency domain signal of the simulated audio.
[0131] S820. The initial signal processing model is trained based on the noiseless time-frequency domain signal and the simulated time-frequency domain signal to obtain the signal processing model.
[0132] The initial signal processing model is a neural network model with initial / default parameters. For example, this initial signal processing model can be a neural network model using a Convolutional Recurrent Network (CRN) model framework.
[0133] The CRN neural network model processes the input signal as follows:
[0134] First, a multi-layer convolutional network is applied to extract time-frequency domain features. Second, a multi-layer recurrent neural network (RNN) is used to track the temporal features of the signal. Then, a mask is obtained using the output of a fully connected layer. Finally, the amplitude spectrum of the input signal is multiplied by the mask to obtain the amplitude spectrum of the speech enhancement, which is then combined with the original phase to obtain the complex spectrum of the speech enhancement.
[0135] Optionally, after obtaining the noiseless time-frequency domain signal and the simulated time-frequency domain signal, the controller can train an initial signal processing model based on multiple frames of noiseless time-frequency domain signal and the corresponding multiple frames of simulated time-frequency domain signal until the training cutoff condition is met, thus obtaining the signal processing model.
[0136] In an alternative embodiment, such as Figure 9 As shown, S820 above trains the initial signal processing model based on the noiseless time-frequency domain signal and the simulated time-frequency domain signal to obtain the signal processing model, including:
[0137] S910: Obtain the noiseless time-frequency domain signal after shaping.
[0138] For example, during the model training phase, the shaping process includes feedback noise elimination, beamforming fusion, and gain adjustment based on the WDRC strategy.
[0139] Optionally, the controller can use the AFC algorithm to eliminate feedback noise in the noiseless time-frequency domain signal, then perform BF fusion on the signal after eliminating feedback noise, then reduce the signal gain of the noise band and increase the signal gain of the speech band based on the WDRC strategy, and then perform ISTFT on the processed signal to obtain the noiseless frequency to be used.
[0140] S920. Input the simulated time-frequency domain signal into the initial signal processing model for noise reduction and feedback noise elimination, and obtain the model output signal of the initial signal processing model for shaping and processing to obtain the simulated audio to be used.
[0141] Optionally, the controller can input the obtained simulation time-frequency domain signal into the initial signal processing model for noise reduction and feedback elimination processing to obtain the model output signal of the initial signal processing model. The AFC algorithm is then used to eliminate feedback from the model output signal of the initial signal processing model. The signal after feedback elimination is then subjected to BF fusion. The fused signal is then subjected to the WDRC strategy to reduce the signal gain of the noise frequency band and increase the signal gain of the speech frequency band. Finally, the processed signal is subjected to ISTFT to obtain the simulation audio to be used.
[0142] S930. Adjust the parameters of the initial signal processing model based on the noise-free frequency to be used and the simulated audio to be used, and obtain the signal processing model.
[0143] Optionally, after obtaining the noise-free frequency and the simulated audio to be used, the controller can adjust the parameters of the initial signal processing model based on multiple frames of noise-free frequency and the corresponding multiple frames of simulated audio to obtain the signal processing model.
[0144] In this embodiment, the noiseless time-frequency domain signal of the noiseless sample audio and the simulated time-frequency domain signal of the simulated audio are acquired. An initial signal processing model is then trained based on these signals to obtain the signal processing model. Specifically, the noiseless time-frequency domain signal is shaped to obtain a noiseless frequency signal to be used. The simulated time-frequency domain signal is input into the initial signal processing model for noise reduction and feedback noise elimination. The model output signal of the initial signal processing model is then shaped to obtain a simulated audio signal to be used. The parameters of the initial signal processing model are then adjusted based on the noiseless frequency signal and the simulated audio signal to obtain the signal processing model. In this method, the model training scenario is more closely matched to the actual link audio processing process, reducing interference between different modules in the link.
[0145] In one embodiment, such as Figure 10 As shown, in step S930 above, the parameters of the initial signal processing model are adjusted based on the noise-free frequency to be used and the simulated audio to be used, resulting in a signal processing model, including:
[0146] S1010: Add preset noise audio to the noiseless frequency to be used by using random signal-to-noise ratio parameters to obtain the target audio.
[0147] The random signal-to-noise ratio parameter is a signal-to-noise ratio parameter randomly obtained within a preset signal-to-noise ratio range. The preset noise audio is the preset noise audio obtained by fusing the noiseless sample audio in S510 to obtain the noisy sample audio.
[0148] Optionally, when adding noise to each frame of the noiseless frequency to be used, the controller can randomly obtain the signal-to-noise ratio parameter within a preset signal-to-noise ratio range to obtain a random signal-to-noise ratio parameter, and add a preset noise audio to the noiseless frequency to be used through the random signal-to-noise ratio parameter to obtain the target audio.
[0149] When adding preset noise audio to a noiseless frequency to be used based on random signal-to-noise ratio parameters, the unit of random signal-to-noise ratio parameters needs to be converted from decibels (dB) to a linear ratio first, and then the preset noise audio is added to the noiseless frequency to be used based on the linear ratio to obtain the target audio.
[0150] Random signal-to-noise ratio parameter With linear proportion The conversion relationship is as follows:
[0151]
[0152] Based on linear proportion For standby noiseless frequency Add preset noise audio , obtain the target audio The process satisfies the following equation:
[0153]
[0154] S1020. Adjust the model parameters in the initial signal processing model according to the loss function between the target audio and the simulated audio to be used until the loss function is minimized, and obtain the signal processing model.
[0155] Optionally, after obtaining the target audio and the simulated audio to be used, the controller can obtain the loss function between the target audio and the simulated audio to be used, so as to adjust the model parameters in the initial signal processing model according to the loss function, so that the loss function of the next frame is decreasing compared to the loss function of the current frame. Through the cyclic adjustment of multiple frames of target audio and multiple frames of simulated audio to be used, until the loss function is minimized, the initial signal processing model with the current parameter state is used as the final signal processing model obtained by training.
[0156] Figure 11 This is the spectrogram of the target audio. Figure 12 The spectrum of the audio signal exhibiting feedback. Figure 13The image shows the spectrogram of the processed audio obtained using the audio processing method provided in this application. It can be seen that the processed audio is very close to the target audio and achieves the effect of suppressing feedback.
[0157] In this embodiment, a preset noise audio is added to the noiseless audio to be used by means of a random signal-to-noise ratio parameter to obtain the target audio. The model parameters in the initial signal processing model are adjusted according to the loss function between the target audio and the simulated audio to be used until the loss function is minimized, thus obtaining the signal processing model. In the above method, the signal processing model is trained so that the trained signal processing model can be used for audio processing in an artificial intelligence manner, thereby improving the audio quality.
[0158] To facilitate understanding by those skilled in the art, the audio processing method provided in this application is described in detail below, such as... Figure 14 As shown, the method may include:
[0159] S1401. Fuse the noiseless sample audio and the preset noisy audio to obtain the noisy sample audio;
[0160] S1402. Perform Fourier transform on the noisy sample audio to obtain the noisy time-frequency domain signal;
[0161] S1403. The noisy time-frequency domain signal is shaped to obtain the noisy time-frequency domain signal to be used.
[0162] S1404. Randomly select a transfer function from the set of transfer functions; the transfer function is used to characterize the conversion relationship between the output audio of the speaker and the output audio of the speaker collected by the microphone;
[0163] S1405. The transfer function is used to convolve the noisy time-frequency domain signal to be applied, and the simulated audio with added acoustic feedback is determined based on the signal after convolution.
[0164] S1406. Obtain the noiseless time-frequency domain signal of the noiseless sample audio and the simulated time-frequency domain signal of the simulated audio.
[0165] S1407. Obtain the noiseless frequency domain signal after shaping;
[0166] S1408. Input the simulated time-frequency domain signal into the initial signal processing model for noise reduction and feedback noise elimination, and obtain the model output signal of the initial signal processing model for shaping and processing to obtain the simulated audio to be used.
[0167] S1409. Add preset noise audio to the noiseless frequency to be used by random signal-to-noise ratio parameters to obtain the target audio;
[0168] S1410. Adjust the model parameters in the initial signal processing model according to the loss function between the target audio and the simulated audio to be used until the loss function is minimized, and obtain the signal processing model.
[0169] S1411. Acquire the original audio signal collected by the microphone, input the time-frequency domain signal obtained by short-time Fourier transform of the original audio signal into the signal processing model for noise reduction and feedback noise elimination, and obtain the model output signal.
[0170] S1412. The model output signal includes at least two channels of time-frequency domain signals; an adaptive feedback cancellation algorithm is used to eliminate feedback noise in each channel of time-frequency domain signals to obtain the feedback cancellation signal of each channel of time-frequency domain signals;
[0171] S1413. Beamforming and fusing all feedback cancellation signals to obtain a single-channel enhanced signal;
[0172] S1414. Adjust the signal gain of the single-channel enhancement signal to obtain the time-frequency domain signal to be used.
[0173] S1415. Perform inverse short-time Fourier transform on the frequency domain signal to be used to obtain the audio to be played.
[0174] It should be noted that the descriptions in S1401-S1415 above can be found in the relevant descriptions in the above embodiments, and their effects are similar, so they will not be repeated here.
[0175] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0176] Based on the same inventive concept, this application also provides an audio processing apparatus for implementing the audio processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio processing apparatus embodiments provided below can be found in the limitations of the audio processing method described above, and will not be repeated here.
[0177] In one embodiment, such as Figure 15 As shown, an audio processing device is provided, including: a signal acquisition module 1501, a model processing module 1502, and an audio determination module 1503, wherein:
[0178] The signal acquisition module 1501 is used to acquire the raw audio signal collected by the microphone;
[0179] The model processing module 1502 performs noise reduction and feedback noise elimination on the original audio signal according to the signal processing model to obtain the model output signal; the signal processing model is obtained by training with simulated audio, which is obtained through closed-loop simulation and includes noise and feedback noise;
[0180] The audio determination module 1503 is used to determine the audio to be played based on the model output signal.
[0181] Each module in the aforementioned audio processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0182] In one embodiment, a hearing aid device is provided, including a controller, a microphone, and a speaker, wherein the controller is used to implement the steps of any of the above-described audio processing methods.
[0183] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described audio processing methods.
[0184] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of any of the above-described audio processing methods.
[0185] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0186] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0187] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An audio processing method, characterized in that, The method includes: Acquire the raw audio signal captured by the microphone; The original audio signal is denoised and feedback is eliminated using a signal processing model to obtain the model output signal. The signal processing model is trained using simulated audio, which is obtained through closed-loop simulation and includes noise and feedback. The audio to be played is determined based on the output signal of the model.
2. The method according to claim 1, characterized in that, The step of determining the audio to be played based on the model output signal includes: The output signal of the model is shaped to obtain the time-frequency domain signal to be used; The inverse short-time Fourier transform is performed on the time-frequency domain signal to be used to obtain the audio to be played.
3. The method according to claim 2, characterized in that, The model output signal includes at least two channels of time-frequency domain signals; the shaping process of the model output signal to obtain the time-frequency domain signal to be used includes: An adaptive feedback cancellation algorithm is used to eliminate feedback noise in the time-frequency domain signals of each channel, resulting in feedback cancellation signals for each channel's time-frequency domain signals. Beamforming fusion is performed on all feedback cancellation signals to obtain a single-channel enhanced signal; The signal gain of the single-channel enhanced signal is adjusted to obtain the time-frequency domain signal to be used.
4. The method according to claim 3, characterized in that, The signal gain adjustment includes at least one of the following processes: Reduce the signal gain of the noise band in the single-channel enhanced signal and increase the signal gain of the speech band in the single-channel enhanced signal; The gain of the signal in the single-channel enhanced signal that is greater than the gain threshold is compressed to the gain threshold. The signal gain of the single-channel enhanced signal is changed by a random gain coefficient.
5. The method according to any one of claims 1-4, characterized in that, The training process of the signal processing model includes: Obtain the noisy time-frequency domain signal of the noiseless sample audio; The simulated audio is determined based on the noisy time-frequency domain signal; The signal processing model is obtained by training the noise-free sample audio and the simulated audio.
6. The method according to claim 5, characterized in that, The acquisition of the noisy time-frequency domain signal of the noiseless sample audio includes: Noise-free sample audio and preset noisy audio are fused to obtain noisy sample audio; The noisy sample audio is subjected to Fourier transform to obtain the noisy time-frequency domain signal.
7. The method according to claim 5, characterized in that, Determining the simulated audio based on the noisy time-frequency domain signal includes: The noisy time-frequency domain signal is shaped to obtain the noisy time-frequency domain signal to be used; Acoustic feedback is added to the noisy time-frequency domain signal to be used to obtain the simulated audio.
8. The method according to claim 7, characterized in that, The step of adding acoustic feedback to the noisy time-frequency domain signal to obtain the simulated audio includes: A transfer function is randomly selected from the set of transfer functions; the transfer function is used to characterize the conversion relationship between the output audio of the speaker and the output audio of the speaker collected by the microphone; The transfer function is used to convolve the noisy time-frequency domain signal to be used, and the simulated audio is obtained from the signal after convolution.
9. The method according to claim 5, characterized in that, The process of training the signal processing model using the noiseless sample audio and the simulated audio includes: Obtain the noiseless time-frequency domain signal of the noiseless sample audio and the simulated time-frequency domain signal of the simulated audio; The initial signal processing model is trained based on the noiseless time-frequency domain signal and the simulated time-frequency domain signal to obtain the signal processing model.
10. The method according to claim 9, characterized in that, The step of training the initial signal processing model based on the noiseless time-frequency domain signal and the simulated time-frequency domain signal to obtain the signal processing model includes: Obtain the noiseless time-frequency domain signal after shaping and processing to obtain the noiseless frequency signal to be used; The simulated time-frequency domain signal is input into the initial signal processing model for noise reduction and feedback noise elimination, and the model output signal of the initial signal processing model is obtained and shaped to obtain the simulated audio to be used. The initial signal processing model is adjusted by adjusting the parameters of the noise-free frequency to be used and the simulated audio to be used, thereby obtaining the signal processing model.
11. The method according to claim 10, characterized in that, The step of adjusting the parameters of the initial signal processing model based on the noise-free frequency to be used and the simulated audio to be used, to obtain the signal processing model, includes: The target audio is obtained by adding preset noise audio to the noiseless frequency to be used by random signal-to-noise ratio parameters; The model parameters in the initial signal processing model are adjusted according to the loss function between the target audio and the simulated audio to be used until the loss function is minimized, thereby obtaining the signal processing model.
12. A hearing aid device, the hearing aid device comprising a controller, a microphone, and a speaker, characterized in that, The controller is used to implement the audio processing method according to any one of claims 1-11.