Method for reducing audio signal noise and hearing aid
The method identifies environmental scenes and user intent to adjust noise reduction processing based on personalized signal-to-noise ratios, addressing the limitations of existing hearing aid noise reduction by enhancing speech clarity in diverse environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-04-03
AI Technical Summary
Existing noise reduction methods in hearing aids fail to adapt effectively to varying environmental scenes, leading to inadequate noise reduction in complex audio environments, and do not account for user-specific noise tolerance and sensitivity.
A method that identifies the environmental scene using a microphone, determines personalized signal-to-noise ratios for maximum speech intelligibility and speech intelligibility threshold, and adjusts noise reduction processing accordingly, incorporating user intent detection through bone conduction transducers.
Enables personalized noise reduction that effectively removes noise without weakening useful signals, adapting to different environments and improving user experience by enhancing speech clarity.
Smart Images

Figure 2026058339000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio processing, and particularly to a method for reducing noise in an audio signal and a hearing aid.
Background Art
[0002] Hearing aids mainly serve to improve the hearing of hearing-impaired people, enabling them to hear and understand sounds better. With the increase in ambient noise in the listening environment, when facing a complex audio environment, hearing aids need to perform noise reduction processing on the ambient noise in order to improve the user's auditory experience.
[0003] [[ID=1s]] Currently, in the noise reduction method for hearing aids, usually, the level of noise reduction is set, and the fitter adjusts according to the user's situation, or the hearing aid uses a corresponding noise reduction solution for a specific frequency or a specific scene to perform noise reduction processing. However, the above noise reduction method is difficult to meet the user's desire for noise reduction.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Based on this, it is necessary to provide a method for reducing noise in an audio signal and a hearing aid with a better noise reduction effect for the above technical problems.
Means for Solving the Problems
[0005] According to a first aspect, this application identifies an environmental scene corresponding to the audio signal based on the audio signal collected by a microphone, determines a first signal-to-noise ratio, which is the signal-to-noise ratio when reaching the highest speech intelligibility after matching the environmental scene and compensating for the user's hearing, and a second signal-to-noise ratio, which is the signal-to-noise ratio when reaching the speech intelligibility threshold after matching the environmental scene and compensating for the user's hearing. The present invention provides a method for reducing noise in an audio signal, which includes performing noise reduction processing on the audio signal based on the first signal-to-noise ratio or the second signal-to-noise ratio.
[0006] According to a second aspect, the present invention further provides a hearing aid comprising a microphone, a speaker, a bone conduction transducer, and a processor, wherein the processor is connected to the microphone, the speaker, and the bone conduction transducer, the bone conduction transducer identifies a user intent based on a captured vibration signal and transmits the identified user intent to the processor, the microphone collects an audio signal and transmits the audio signal to the processor, the processor performs the steps of the audio signal noise reduction method, performs noise reduction processing on the audio signal, and transmits the noise-reduced audio signal to the speaker.
[0007] The above-described audio signal noise reduction method and hearing aid identify the environmental scene corresponding to the audio signal based on the audio signal collected by the microphone, then determine the signal-to-noise ratio (i.e., first signal-to-noise ratio) corresponding to the user's maximum speech intelligibility in that scene (i.e., second signal-to-noise ratio) based on the environmental scene, quantify the user's tolerance and sensitivity to noise, and then perform noise reduction processing based on the first or second signal-to-noise ratio. This allows for effective removal of noise during the noise reduction process without excessively weakening useful signals, achieving personalized noise reduction according to the user's tolerance to noise and environmental conditions, enabling the user to hear clearer speech in various environments. Furthermore, the above solution can automatically adjust the noise reduction policy according to different environmental scenes, making it applicable to various noisy environments and improving adaptability and flexibility.
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings necessary for describing the embodiments of this application or related technologies are briefly described below. Clearly, the drawings described below are only a few embodiments of this application, and those skilled in the art can obtain other relevant drawings based on these drawings without any creative work. [Brief explanation of the drawing]
[0009] [Figure 1] This is a flowchart of a method for reducing audio signal noise in one embodiment. [Figure 2] This is a flowchart of the steps for identifying an environmental scene in one embodiment. [Figure 3] This is a flowchart of the steps for identifying an environmental scene in another embodiment. [Figure 4] This is a flowchart of the steps for noise reduction processing in one embodiment. [Figure 5] This is a flowchart of the steps for noise reduction processing in another embodiment. [Figure 6] This is a structural block diagram of a call voice signal noise reduction device according to one embodiment. [Figure 7] This is a structural block diagram of a call audio signal noise reduction device according to another embodiment. [Figure 8] This is a structural block diagram of a hearing aid according to one embodiment. [Figure 9] This is an internal configuration diagram of a computer device in one embodiment. [Modes for carrying out the invention]
[0010] To further clarify the purpose, technical solution, and advantages of this application, the application will be described in more detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and do not limit the application.
[0011] In one embodiment, as shown in Figure 1, a method for reducing audio signal noise is provided, and although this embodiment has described the method as being applied to an auditory aid such as a hearing aid, it can be understood that the method may also be applied to a server, or to a system including a hearing aid and a server, and may be implemented through interaction between the hearing aid and the server. In this embodiment, the method includes steps S200 to S600.
[0012] The S200 identifies the environmental scene corresponding to the audio signal based on the audio signal collected by the microphone.
[0013] In this embodiment, environmental scenes include, but are not limited to, a variety of scenes such as quiet environments, meetings, street traffic, noisy restaurants, shopping malls, mahjong parlors, and gyms. As hearing aids are everyday items for users with moderate to severe hearing loss, users typically wear them for long periods, and users wear them in various life situations; therefore, it is necessary to satisfy the scene identification requirements for a variety of environmental scenes.
[0014] In a concrete implementation, taking a hearing aid as an example, ambient sound signals may be continuously collected by a microphone built into the hearing aid, and these sound signals may include target speech and background noise. Subsequently, the collected sound signals are preprocessed to extract important environmental and sound features, which may include spectral characteristics, energy distribution, and time-domain and frequency-domain statistics. Then, by analyzing the extracted environmental and sound features using a pre-trained classifier (e.g., a support vector machine, neural network, etc.), the current environmental scene can be identified. As can be understood, since the user may encounter different scenes during the actual wearing and use process, the environmental scene identification process is a continuous process, and if the scene changes, the noise reduction policy can be automatically adjusted according to the scene change to meet the user's noise reduction needs.
[0015] The S400 determines a first signal-to-noise ratio, which is the signal-to-noise ratio at which the highest speech intelligibility is achieved after matching the environmental scene and compensating for the user's hearing, and a second signal-to-noise ratio, which is the signal-to-noise ratio of noise corresponding to the user's speech intelligibility threshold, after matching the environmental scene.
[0016] The signal-to-noise ratio (SNR) is the ratio of signal power to noise power, usually expressed in decibels (dB). The SNR is an indicator used to judge signal strength relative to background noise intensity and directly affects signal quality and intelligibility. Speech intelligibility is usually expressed as a percentage, i.e., the ratio of the number of speech elements correctly identified by the subject to the total number of test elements. Speech intelligibility is an important indicator for evaluating an individual's speech comprehension ability. Maximum speech intelligibility refers to the highest speech intelligibility a user can achieve under the most ideal conditions, such as a quiet environment. The speech recognition threshold (SRT) refers to the lowest sound level required for a user to correctly identify half (50%) of the speech elements, i.e., the signal-to-noise ratio at which a user can identify 50% of the speech elements.
[0017] In this embodiment, the first signal-to-noise ratio refers to the signal-to-noise ratio at which the user achieves the highest level of speech intelligibility, and the first signal-to-noise ratio may include the signal-to-noise ratio at which the user achieves the highest level of speech intelligibility in a quiet environment or in a noisy environment. The second signal-to-noise ratio refers to the signal-to-noise ratio at which the user can identify 50% of the speech material, and similarly, the second signal-to-noise ratio may include the signal-to-noise ratio at which the user can identify 50% of the speech material in a quiet environment or in a noisy environment.
[0018] In actual applications, even if the user fits the noise reduction parameters in a quiet environment and the speech intelligibility threshold reaches the desired target value, when the user enters a noisy environment, the noise is too loud to hear clearly, the noise reduction effect does not adapt in each scene, and there may be a situation where repeating the adjustment by oneself does not improve the effect either. To achieve personalized noise reduction, a large number of environmental scenes are pre-recorded, K typical environmental scenes are clustered, and then the speech intelligibility thresholds of different environmental scenes are tested for the hearing aid users to determine the speech intelligibility thresholds of the users in different environmental scenes and the signal-to-noise ratios required to achieve different speech intelligibilities, so as to quantify the user's tolerance and sensitivity to noise.
[0019] Specifically, in the fitting room, after the user wears the hearing aid, in different environmental scenes, the signal-to-noise ratio (referred to as the first signal-to-noise ratio SNR1) when the user reaches the highest speech intelligibility X% after compensating for the user's hearing, and the signal-to-noise ratio (referred to as the second signal-to-noise ratio SNR2) when the user reaches 50% speech intelligibility (i.e., the speech intelligibility threshold) may be tested. Then, the first signal-to-noise ratio and the second signal-to-noise ratio corresponding to different environmental scenes are memorized.
[0020] In a specific implementation, after identifying the environmental scene where the user is currently located, the first signal-to-noise ratio and the second signal-to-noise ratio that match the identified environmental scene may be searched from the correspondence between the pre-memorized environmental scenes and the first signal-to-noise ratio and the second signal-to-noise ratio. For example, if the identified environmental scene is a meeting scene, the first signal-to-noise ratio and the second signal-to-noise ratio that match the meeting scene are obtained. In some other embodiments, the AI (Artificial Intelligence) processing model may directly process the audio signal to identify the corresponding environmental scene and determine the first signal-to-noise ratio and the second signal-to-noise ratio of the environmental scene.
[0021] In the S600, noise reduction processing is performed on the audio signal based on either the first signal-to-noise ratio or the second signal-to-noise ratio.
[0022] After determining a first signal-to-noise ratio and a second signal-to-noise ratio that match the environmental scene, a noise reduction target is determined based on the first or second signal-to-noise ratio, signal processing parameters are adjusted, the processed audio signal is further compressed and gain-adjusted, and finally output to the user via the hearing aid speaker. For example, if it is necessary to adjust the noise reduction target based on the first signal-to-noise ratio SNR1, the signal-to-noise ratio of the output audio signal may be brought as close to SNR1 as possible. In some other embodiments, the noise reduction policy may be adjusted based on the first or second signal-to-noise ratio to match the user's intention whether or not they want to speak.
[0023] Furthermore, the hearing aid further includes a user feedback mechanism that allows the user to adjust the noise reduction level or other parameters, for example, by pressing a button on the hearing aid or through a compatible smartphone application program. The hearing aid further adjusts parameters to optimize performance based on user feedback and preferences. For example, if the user provides feedback that a particular scene is difficult to hear due to the set signal-to-noise ratio parameters, the hearing aid may automatically learn and optimize the noise reduction policy for that scene.
[0024] In the above-described audio signal noise reduction method and hearing aid, a signal-to-noise ratio (first signal-to-noise ratio) corresponding to the user's highest speech intelligibility in different environmental scenes and a signal-to-noise ratio (second signal-to-noise ratio) corresponding to when the speech intelligibility threshold is reached are predetermined to quantify the user's tolerance and sensitivity to different noise levels. In actual application, the environmental scene corresponding to the audio signal is identified based on the audio signal collected by the microphone, then the first and second signal-to-noise ratios that match the environmental scene are obtained, and then noise reduction processing is performed based on the first or second signal-to-noise ratio. This allows noise to be effectively removed by the noise reduction processing without excessively weakening useful signals, achieving the effect of personalized noise reduction according to the user's tolerance to noise and environmental conditions, enabling the user to hear clearer speech in various environments. Furthermore, the above solution can automatically adjust the noise reduction policy according to different environmental scenes, making it applicable to various noisy environments and improving adaptability and flexibility.
[0025] In some exemplary embodiments, the method further includes, before determining a first signal-to-noise ratio and a second signal-to-noise ratio that match an environmental scene, performing a speech intelligibility threshold test under information masking on a user at different signal-to-noise ratio levels in multiple environmental scenes and obtaining test results, and determining the first signal-to-noise ratio and the second signal-to-noise ratio for the user in different environmental scenes based on the test results and the user's hearing data.
[0026] The speech intelligibility threshold test under information masking aims to evaluate an individual's speech intelligibility ability in the presence of informative noise. Informative noise refers to noise that resembles the speech signal spectrum. User hearing data includes, but is not limited to, an audiogram and type of hearing loss. Since users may encounter various environmental scenes in their lives, audio signals in N different environmental scenes may be collected and recorded. Subsequently, clustering algorithms, such as the K-means algorithm or an artificial intelligence algorithm, may be used to project the N environmental scenes into a high-dimensional space, and K typical environmental scenes, including meeting scenes, various transportation scenes, restaurant scenes, shopping mall scenes, and home scenes, may be selected.
[0027] Subsequently, standardized speech test material is prepared, including speech signals and informational noise, such as speech signals and informational noise including two-syllable words. Based on the user's audiogram and type of hearing loss, a speech intelligibility test is performed on the user in a quiet environment. The signal-to-noise ratio at which the user achieves maximum speech intelligibility of X% in a quiet environment is determined. Subsequently, a hearing aid is fitted and relevant noise reduction parameters are set based on this signal-to-noise ratio parameter.
[0028] Next, an environmental scene N1 is selected from K typical scenes, a fitted hearing aid is placed on user A1, and the noise reduction function is turned off. Then, to ensure that the user can easily identify speech material, the test is started with a high signal-to-noise ratio, e.g., 25 dB, and an audio signal is provided. The user is given speech test material (e.g., 10 two-syllable words) and background noise is played simultaneously. The user is asked to repeat the words they heard, and the user's feedback data is collected and statistically analyzed to obtain speech intelligibility at the current signal-to-noise ratio. Then, the signal-to-noise ratio is gradually reduced, for example, by 5 dB at a time (e.g., reduced from 25 dB to 20 dB, and then further reduced to 15 dB), and the above steps are repeated until it becomes difficult for the user to accurately identify more than half of the speech test material. If a user can accurately identify x% of the test material at a given signal-to-noise ratio, this ratio is recorded as A1_N1_SRTx, and A1_N1_SRTx includes the corresponding signal-to-noise ratio (i.e., the first signal-to-noise ratio) at which the user reaches their highest speech intelligibility. If a user can accurately identify 50% of the speech test material at a given signal-to-noise ratio, this ratio is recorded as A1_N1_SRT50 (i.e., the second signal-to-noise ratio). To illustrate with an example, assuming a user in a quiet environment (considered scene N1) with a highest speech intelligibility of x% (e.g., 90%), the test may begin with a signal-to-noise ratio of 25 dB and gradually decrease. Suppose the user can accurately identify 90% of the test material at a signal-to-noise ratio of y1 dB. In this case, A1_N1_SRTx = y1 dB. Subsequently, the signal-to-noise ratio is continued to decrease until the user can only accurately identify 50% of the test material. Assume that the user can accurately identify 50% of the test material with a signal-to-noise ratio of y2dB. Therefore, A1_N1_SRT50 = y2dB.
[0029] Finally, the above test process is repeated for each subsequently selected environment scene (e.g., N2, N3, ..., Nk) to obtain the A1_Nk_SRTx and A1_Nk_SRT50 for each environment scene of the user. Similarly, the A1_Nk_SRTx and A1_Nk_SRT50 obtained above for each environment scene may be initial values, and the user may fine-tune and calibrate the A1_Nk_SRTx and A1_Nk_SRT50 for each environment scene using a feedback mechanism to determine the final A1_Nk_SRTx and A1_Nk_SRT50 for each environment scene.
[0030] Furthermore, if the user is unable to reach the fitting room to undergo a speech intelligibility threshold test, the user may perform the test themselves. The user wears and uses a hearing aid, and a microphone built into the hearing aid collects an external audio signal. Subsequently, the audio features are extracted from the audio signal, and a scene recognition algorithm may be used to identify the scene in which the user is currently located. Alternatively, the audio features may be clustered with audio features from K typical scenes in the test environment, and the most similar typical scene may be matched. Subsequently, noise reduction processing is performed based on the SRTx and SRT50 of the matched most similar typical scene, and the processed audio signal is output to the user. The user is then prompted to confirm whether or not they need to adjust the strength of the noise reduction. If the user needs to adjust the noise reduction level or other parameters, the hearing aid may further adjust the noise reduction parameters based on the user's feedback and preferences. In this way, the user's SRTx and SRT50 for different environmental scenes are obtained.
[0031] In this embodiment, by combining standardized test materials with the user's hearing data in different environmental scenarios and performing a speech intelligibility threshold test on the user, it is possible to quantify the user's tolerance and sensitivity to different noise levels, and further provide a basis for personalized adjustment of hearing aids, thereby improving the user's listening experience.
[0032] There are various methods for identifying environmental scenes. As shown in Figure 2, in some exemplary embodiments, identifying an environmental scene corresponding to an audio signal includes steps S220 and S240.
[0033] S220 extracts the first environmental features of the audio signal.
[0034] In S240, the first environmental feature is matched to the second environmental feature of a different environmental scene to identify the environmental scene corresponding to the audio signal.
[0035] Environmental features refer to characteristics that describe the attributes of an acoustic environment and are used to identify and distinguish different environmental scenes. Environmental features include, but are not limited to, spectral features, time-domain features, modulation features, and statistical properties such as the mean and variance of a signal.
[0036] In actual applications, a classifier that predicts environmental scenes based on environmental features may be pre-trained based on training data that includes historical environmental features (called second environmental features) in various environmental scenes. Alternatively, an audio signal from the current scene may be collected using a microphone, the collected audio signal may be pre-processed to extract environmental features from the audio signal, and then the extracted environmental features may be input to a trained classifier. The classifier may then output a predicted environmental scene, thereby obtaining the environmental scene corresponding to the audio signal.
[0037] In other embodiments, second environmental features in different environmental scenes may be predetermined. In actual application, after extracting the first environmental features of the audio signal collected by the microphone, similarity matching may be performed between the extracted first environmental features and second environmental features in different environmental scenes to find the second environmental feature most similar to the first environmental feature, and the environmental scene corresponding to the second environmental feature may be determined as the environmental scene corresponding to the audio signal.
[0038] In this embodiment, the accuracy of environmental scene classification can be improved by predicting scenes based on the environmental characteristics of multiple environmental scenes.
[0039] Environmental feature matching may be achieved by clustering. As shown in Figure 3, in some exemplary embodiments, S240 includes S242. In S242, second environmental features in different environmental scenes are clustered to obtain clustering results, the distance between each cluster center in the clustering results and the first environmental feature is determined, and the environmental scene characterized by the cluster center with the closest distance is determined as the environmental scene corresponding to the audio signal.
[0040] There may be a large number of second environmental features in different environmental scenes, and there may be correlations between these second environmental features. Therefore, following the previous embodiment, after extracting the environmental features of the audio signal, a clustering algorithm may be used to cluster the second environmental features in different environmental scenes, an appropriate number of clusters may be determined according to the elbow method, clustering may be performed on the extracted second environmental features based on the number of clusters, and a cluster center may be determined in each cluster to obtain the clustering result. This cluster center characterizes the typical features of the cluster. Subsequently, for each cluster center, the distance to the first environmental feature, for example, the Euclidean distance, may be calculated, and the environmental scene characterized by the cluster center with the closest distance may be determined as the environmental scene most similar to the environment in which the user is currently located, and the environmental scene characterized by this cluster center may be determined as the environmental scene corresponding to the audio signal.
[0041] In this embodiment, by performing clustering analysis on the environmental features of different environmental scenes, the user's current location in the environmental scene can be identified more accurately. Furthermore, the clustering analysis can automatically adjust the cluster centers to adapt to matching new environmental scenes.
[0042] Considering personalized noise reduction, it is most preferable to combine it with the user's subjective intentions in order to achieve a personalized noise reduction effect. In one exemplary embodiment, as shown in Figure 4, S600 includes steps S620 to S660.
[0043] The S620 identifies user intent.
[0044] In S640, the signal-to-noise ratio target value is determined based on the identified user intent, using either the first or second signal-to-noise ratio.
[0045] In the S660, noise reduction processing is performed on the audio signal based on a target signal-to-noise ratio.
[0046] In this embodiment, user intent primarily refers to the subjective intention of the user to want to contact the outside world, and includes, but is not limited to, the intention to speak and the intention to remain silent. The intention to speak refers to the user's desire to pay attention to communication with others and to be able to develop good communication with them. The intention to remain silent refers to the user's unwillingness to speak and unwillingness to communicate with others.
[0047] In practical implementation, the hearing aid may analyze the audio signal collected by a pre-trained classifier to identify the user's intent. For example, if it is identified that there is a clear conversational component, the hearing aid may determine that the user's intent is a speech intent. Subsequently, based on the identified user intent, the hearing aid may appropriately select a first or second signal-to-noise ratio, determine a target signal-to-noise ratio, and further adjust the noise reduction policy.
[0048] In this embodiment, a better personalized noise reduction effect can be achieved by identifying the user intent and adjusting the noise reduction level by appropriately selecting either the first signal-to-noise ratio or the second signal-to-noise ratio in combination with the user intent.
[0049] As shown in Figure 5, in some exemplary embodiments, S620 includes S622. In S622, the vibration signal of the bone conduction transducer is acquired, and the user intent is identified based on the vibration signal.
[0050] Bone conduction transducers can also be called bone voiceprint sensors. In practical applications, bone conduction transducers are built into hearing aids, and when a user speaks, sound is generated primarily by the vibration of the vocal cords, or when airflow passes through the vocal cords, causing them to vibrate and generate sound. As sound passes through resonant cavities such as the oral cavity, nasal cavity, and pharynx, it causes minute vibrations and resonances in the bones and soft tissues of these areas, enhancing and altering the characteristics of the sound. Bone conduction transducers can capture the vibration signals of the vocal cords and the minute vibration signals of the resonant cavities to determine whether or not the user is speaking.
[0051] In practical implementation, user intent may be identified by a bone conduction transducer built into the hearing aid. The bone conduction transducer may be mounted inside the hearing aid case to be close to the bones of the user's head and face and to effectively capture vibration signals generated when the user speaks. After capturing vibration signals from the user's vocal cords and resonant cavity, the bone conduction transducer performs preprocessing such as filtering or enhancement on the captured vibration signals, and then extracts important features related to the vibration mode during speech, such as vibration frequency and amplitude, from the preprocessed vibration signals. Subsequently, the user intent is determined by identifying whether or not the user is speaking based on the extracted important features. Exemplarily, by setting a threshold, it may be determined that the user is speaking if the feature values of the captured vibration signals exceed the threshold, and otherwise, it may be determined that the user is not speaking.
[0052] In this embodiment, by capturing vibration signals using bone conduction transducers and identifying user intent, it is possible to accurately capture user vibration signals and accurately identify user intent even in noisy environments.
[0053] In some other exemplary embodiments, identifying user intent based on vibration signals includes determining vibration modes based on vibration signals and identifying user intent based on vibration modes.
[0054] In this embodiment, the vibration modes mainly include a first vibration mode during speech and a second vibration mode during non-speech states, and the second vibration mode includes, but is not limited to, modes such as chewing, swallowing, and coughing.
[0055] Specifically, a pre-configured vibration mode identification algorithm may be used to analyze features extracted from the vibration signal and determine whether the vibration mode is the first vibration mode in a speech state or the second vibration mode in a non-speech state. If the vibration mode is the first vibration mode, the user intention is determined to be a speech intention; if the vibration mode is the second vibration mode, the user intention is determined to be a silence intention.
[0056] Alternatively, vibration signals from a large number of users during speech may be collected, a first vibration feature may be extracted from these vibration signals, vibration signals from users during non-speech states may be collected, a second vibration feature may be extracted from these vibration signals, vibration mode labels may be added to the first and second vibration features, and then a vibration mode discrimination model may be trained based on the labeled first and second vibration features. In actual application, vibration signals output from a bone conduction transducer are collected, feature data is extracted from them, and the feature data is input into the trained vibration mode discrimination model to determine the vibration mode. If the vibration mode is the first vibration mode, it is determined that the user intention is to speak, and if the vibration mode is the second vibration mode, it is determined that the user intention is to remain silent. Similarly, a feedback mechanism may be used to confirm whether the user intention has been accurately determined, and if the hearing aid incorrectly determines that the user is speaking, it may be used to prompt the user to confirm whether they are speaking or not through voice prompts or haptic feedback.
[0057] In this embodiment, by analyzing specific vibration modes and identifying user intentions based on those modes, it is possible to more accurately distinguish between the user's actual speech intentions and other oral activities (e.g., chewing, swallowing, etc.), thereby improving the user experience.
[0058] Exemplaryly, as shown in Figure 5, in some other embodiments, S640 includes S642 and S644.
[0059] In S642, if the identified user intent is a speech intent, the first signal-to-noise ratio is determined as the target signal-to-noise ratio value.
[0060] In S644, if the identified user intent is a silent intent, the target signal-to-noise ratio is determined based on the second signal-to-noise ratio.
[0061] Specifically, if the user intent is a speech intent, the system may characterize whether the user wants to speak or is speaking, set the goal of noise reduction to maximize speech intelligibility, and characterize the need for more noise reduction. In this case, the first signal-to-noise ratio at which the user achieves the highest speech intelligibility in a noisy environment may be determined as the signal-to-noise ratio target value, and noise reduction processing may be performed on the speech signal based on this signal-to-noise ratio target value.
[0062] If the user intent is silence intent, the noise reduction goal may be to maintain the richness of ambient sound by exposing the user to richer sounds, thus characterizing the user's unwillingness to speak and ensuring the user cannot tolerate the noise. In some other embodiments, the signal-to-noise ratio target may be set directly to the second signal-to-noise ratio, or the signal-to-noise ratio target may be determined by combining user needs.
[0063] In this embodiment, by identifying the user's intent, the target value for noise reduction processing is determined, and based on this, noise reduction processing is performed on the audio signal, thereby significantly improving the user's listening experience and providing a more personalized and efficient hearing aid experience for the user.
[0064] In some other embodiments, determining the target signal-to-noise ratio based on the second signal-to-noise ratio is possible. Based on the second signal-to-noise ratio, a recommended signal-to-noise ratio value greater than the second signal-to-noise ratio is determined, The system pushes signal-to-noise ratio (SNOR) information, including recommended SNOR values, and allows users to adjust based on these recommended values to obtain a target SNOR value. This includes receiving a target signal-to-noise ratio value as feedback from the user.
[0065] In this embodiment, the recommended signal-to-noise ratio (SNR) may be understood as the SNR value initially determined and confirmed by the user, or it may be considered a recommended SNR reference value. As described in the previous embodiment, in order to expose the user to richer sounds and maintain the richness of ambient sounds, the recommended signal-to-noise ratio determined based on the second SNR is slightly higher than the second SNR. For example, if the second SNR is 5 dB, the determined recommended signal-to-noise ratio may be 7 dB.
[0066] The signal-to-noise ratio (SNR) target value is a SNR determined by the user based on their needs and a recommended SNR, and characterizes the user's desired SNR level. In this embodiment, the SNR target value may be determined by the user receiving SNR information, including a recommended SNR, and then deciding whether or not to use this recommended SNR, depending on their needs and actual circumstances, or by setting a SNR that meets their expectations based on the recommended SNR.
[0067] In practical implementation, to ensure that the noise reduction effect meets user expectations, signal-to-noise ratio (SNOR) information, including recommended SNOR values, may be pushed. This SNOR information may be pushed in the form of speech tones, allowing users to check whether they need to adjust the recommended SNOR values and provide feedback on a target SNOR value that matches their preferences.
[0068] In some other embodiments, to achieve personalized noise reduction, the audio signal is first subjected to noise reduction processing according to a recommended signal-to-noise ratio (SNR) and output to the user. Subsequently, SNR information, including the recommended SNR, is pushed to the user to confirm whether the user needs to adjust the SNR. If adjustment is necessary, the user can fine-tune and calibrate the current noise reduction level using a button on the hearing aid or a suitable smartphone application program, determine a SNR target value that matches their preference, and feed this SNR target value back to the hearing aid.
[0069] Specifically, if a user believes that the recommended signal-to-noise ratio (SNR) value meets their expectations and that there is no need to adjust the current SNR level, they may determine the SNR target value using a button on the hearing aid or a suitable smartphone application program, and then feed this target value back to the hearing aid. If a user believes that the recommended SNR value does not meet their expectations and that there is a need to adjust the current SNR level, they may adjust the SNR value using a button on the hearing aid or a suitable smartphone application program (for example, by increasing or decreasing the SNR) to determine a SNR target value that meets their expectations, and then feed this target value back to the hearing aid. After receiving the SNR target value fed back from the user, the hearing aid processes the subsequent audio signal according to the SNR target value to meet the user's hearing needs.
[0070] In this embodiment, by pushing signal-to-noise ratio information to the user, user participation can be increased, and personalized noise reduction can be achieved according to the user's needs.
[0071] In some other embodiments, the method further includes reducing the low-frequency gain and / or pausing playback of media data if the identified user intent is a speech intent.
[0072] In practical applications, when a user speaks, those sounds are transmitted through the bone to the inner ear, causing the user to perceive their own voice, especially the low-frequency range, as brighter than the actual speech. This phenomenon is called the "occlusion effect." To mitigate the impact of the occlusion effect on the user, reducing the gain in the low-frequency range can reduce the perceived muddiness of the sound and improve the user experience.
[0073] Therefore, in practical implementation, the hearing aid can detect the user's intent in real time using a bone-voiceprint sensor and determine whether the user is speaking or not. If the hearing aid identifies that the user is speaking, it can reduce the low-frequency gain to reduce the sensation of occlusion.
[0074] If a user is listening to external media data (e.g., music, TV programs) while speaking, these media sounds can interfere with the user's communication. Therefore, by automatically pausing the playback of media content when the hearing aid detects that the user is speaking, the user can focus more on the conversation while the hearing aid is speaking, and the impact of background noise can be reduced. To understand this, the hearing aid can reduce the low-frequency gain while simultaneously pausing the playback of media data in order to maximize the user's listening experience.
[0075] In this embodiment, by reducing the low-frequency gain and pausing the playback of media data, the listening experience of the user during speech can be significantly improved.
[0076] Considering that hearing aids can be equipped with AI (Artificial Intelligence) processing modules, speech noise reduction can be performed using AI models. In some exemplary embodiments, the method is as follows: The process further includes extracting environmental features of an audio signal collected by a microphone, calling a trained audio signal processing model using environmental features and user hearing feature data as input to identify the environmental scene corresponding to the audio signal, determining a first signal-to-noise ratio and a second signal-to-noise ratio that match the environmental scene, performing noise reduction processing on the audio signal based on the first or second signal-to-noise ratio, and outputting the audio signal after noise reduction. The audio signal processing model is obtained by training it based on hearing data from different users and the first and second signal-to-noise ratios in different environmental scenes.
[0077] In this embodiment, the speech recognition model may be a deep learning model or an end-to-end AI model. The end-to-end AI model can learn directly from the original input data to the final output and does not require explicit manual feature engineering or artificial intervention in intermediate steps.
[0078] Specifically, the training process for a speech recognition model may be as follows: Collect test results from speech intelligibility threshold tests under information masking in different environmental scenes for different users, including Nk_SRTx and Nk_SRT50 in different environmental scenes and hearing data for different users, to construct an original dataset. Then, extract environmental feature data and user hearing feature data from the original dataset, assign environmental scene labels to each data point, and mark the first signal-to-noise ratio and second signal-to-noise ratio corresponding to each environmental scene. Subsequently, select an appropriate machine learning model or deep learning model, such as a convolutional neural network, to construct an initial speech signal processing model. By inputting the environmental feature data, hearing feature data, and corresponding environmental scene labels into the initial speech signal processing model and training it, the model can predict different environmental scenes and their corresponding first and second signal-to-noise ratios.
[0079] In actual application, the system uses a model to extract environmental features of the audio signal collected by a microphone, and uses the environmental features and user hearing feature data as input to call a trained audio signal processing model. This model predicts the environmental scene corresponding to the audio signal and the first and second signal-to-noise ratios that match the environmental scene. Based on the first or second signal-to-noise ratio, a target signal-to-noise ratio is determined, noise reduction parameters are adjusted based on the target signal-to-noise ratio, and the audio signal with reduced noise is output to the user.
[0080] In this embodiment, by performing noise reduction processing using a trained speech signal processing model, it is possible to more accurately identify the current environmental scene, automatically perform personalized noise reduction processing based on the user's hearing data, and the model can be applied to more scene identification, effectively improving the user experience.
[0081] To more clearly explain the audio signal noise reduction processing method according to the present invention, a specific embodiment will be described below with reference to the following steps S100 to S112.
[0082] The S100 acquires audio signals collected by the microphone.
[0083] In S102, the first environmental features of the audio signal are extracted.
[0084] In S104, the second environmental features in different environmental scenes are clustered to obtain the clustering results. The distance between each cluster center in the clustering results and the first environmental feature is determined, and the environmental scene characterized by the cluster center with the closest distance is determined as the environmental scene corresponding to the audio signal.
[0085] In S106, a first signal-to-noise ratio is determined, which is the signal-to-noise ratio at which the highest speech intelligibility is achieved after matching the environmental scene and compensating for the user's hearing. A second signal-to-noise ratio is determined, which is the signal-to-noise ratio at which the speech intelligibility threshold is reached after matching the environmental scene and compensating for the user's hearing.
[0086] In S108, the vibration signal from the bone conduction transducer is acquired, the vibration mode is determined based on the captured vibration signal, and the user's intent is identified based on the vibration mode.
[0087] In S110, if the user intent is identified as a silent intent, a recommended signal-to-noise ratio (SNOR) value greater than the second SNOR is determined based on the second SNOR, the SNOR information including the recommended SNOR is pushed, the user adjusts it based on the recommended SNOR to obtain a target SNOR, and the target SNOR is received as feedback from the user.
[0088] In S112, if the identified user intent is a speech intent, the first signal-to-noise ratio is determined as the noise reduction target, and the low-frequency gain is reduced.
[0089] To ensure clarity, the steps in the flowcharts for each of the above embodiments are indicated sequentially by arrows, but these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise explicitly stated herein, the execution of these steps is not limited to a strict order, and these steps may be performed in other orders. Furthermore, at least some of the steps in the flowcharts for each of the above embodiments may include multiple steps or stages, and these steps or stages are not necessarily performed at the same time, but may be performed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed sequentially or alternately with other steps or at least some of the steps or stages within other steps.
[0090] Based on a similar inventive concept, embodiments of the present application further provide an audio signal noise reduction device that realizes the above-described audio signal noise reduction method. Since the means for solving the problems related to the device are similar to the means for solving the problems described in the above-described method, specific limitations in one or more embodiments of the audio signal noise reduction device described below can be made by referring to the limitations for the above-described audio signal noise reduction method, and therefore the explanation is omitted here.
[0091] In one exemplary embodiment, as shown in Figure 6, an audio signal noise reduction device 600 is provided, which includes a scene identification module 610, a data determination module 620, and a noise reduction processing module 630. The scene identification module 610 identifies the environmental scene corresponding to the audio signal based on the audio signal collected by the microphone.
[0092] The data determination module 620 determines a first signal-to-noise ratio, which is the signal-to-noise ratio at which the highest speech intelligibility is achieved after matching the environmental scene and compensating for the user's hearing, and a second signal-to-noise ratio, which is the signal-to-noise ratio at which the speech intelligibility threshold is reached after matching the environmental scene and compensating for the user's hearing.
[0093] The noise reduction processing module 630 performs noise reduction processing on the audio signal based on a first signal-to-noise ratio or a second signal-to-noise ratio.
[0094] The above-described audio signal noise reduction device identifies the environmental scene corresponding to the audio signal based on the audio signal collected by the microphone, then determines the signal-to-noise ratio (first signal-to-noise ratio) corresponding to the user's maximum speech intelligibility in that scene and the signal-to-noise ratio (second signal-to-noise ratio) corresponding to the user's speech intelligibility threshold based on the environmental scene, thereby quantifying the user's tolerance and sensitivity to noise, and then performs noise reduction processing based on the first or second signal-to-noise ratio. This allows the noise reduction process to effectively remove noise without excessively weakening useful signals, achieving personalized noise reduction according to the user's tolerance to noise and environmental conditions, enabling users to hear clearer speech in various environments. Furthermore, the device can automatically adjust the noise reduction policy according to different environmental scenes, making it applicable to various noisy environments and improving adaptability and flexibility.
[0095] As shown in Figure 7, in another exemplary embodiment, the device further includes an intent recognition module 622 that identifies user intent, and a noise reduction processing module 630 that further determines a signal-to-noise ratio target value based on a first signal-to-noise ratio or a second signal-to-noise ratio for the identified user intent, and performs noise reduction processing on the audio signal based on the signal-to-noise ratio target value.
[0096] In another exemplary embodiment, the noise reduction processing module 630 further determines a first signal-to-noise ratio as a target signal-to-noise ratio if the identified user intent is a speech intent, and determines a target signal-to-noise ratio based on a second signal-to-noise ratio if the identified user intent is a silence intent.
[0097] In another exemplary embodiment, the noise reduction processing module 630 further determines a recommended signal-to-noise ratio value greater than the second signal-to-noise ratio based on the second signal-to-noise ratio, pushes signal-to-noise ratio information including the recommended signal-to-noise ratio value, adjusts it based on the recommended signal-to-noise ratio value by the user to obtain a target signal-to-noise ratio value, and receives the target signal-to-noise ratio value as feedback from the user.
[0098] In another exemplary embodiment, the intent identification module 622 acquires vibration signals from a bone conduction transducer and identifies user intent based on the vibration signals.
[0099] In another exemplary embodiment, the intent identification module 622 determines a vibration mode based on the vibration signal and identifies the user intent based on the vibration mode.
[0100] In another exemplary embodiment, the scene identification module 610 further extracts a first environmental feature of the audio signal, matches the first environmental feature to a second environmental feature of a different environmental scene, and identifies the environmental scene corresponding to the audio signal.
[0101] In another exemplary embodiment, the scene identification module 610 further clusters the second environmental features in different environmental scenes to obtain clustering results, determines the distance between each cluster center in the clustering results and the first environmental feature, and determines the environmental scene characterized by the cluster center with the closest distance as the environmental scene corresponding to the audio signal.
[0102] As shown in Figure 7, in another exemplary embodiment, the device further includes an AI processing module 640 that extracts environmental features of an audio signal collected by a microphone, takes environmental features and user hearing feature data as input, invokes a trained audio signal processing model to identify the environmental scene corresponding to the audio signal, determines a first signal-to-noise ratio and a second signal-to-noise ratio that match the environmental scene, performs noise reduction processing on the audio signal based on the first signal-to-noise ratio or the second signal-to-noise ratio, and outputs the audio signal after noise reduction. The audio signal processing model is obtained by training on different user hearing data and the first signal-to-noise ratio and the second signal-to-noise ratio in different environmental scenes.
[0103] As shown in Figure 7, in another exemplary embodiment, the apparatus further includes a test module 602 that performs a speech intelligibility threshold test under information masking on a user at different signal-to-noise ratio levels in multiple environmental scenes, obtains test results, and determines a first signal-to-noise ratio and a second signal-to-noise ratio for the user in different environmental scenes based on the test results and the user's hearing data.
[0104] In another exemplary embodiment, the device further includes a signal optimization module 650 that reduces the low-frequency gain and / or pauses playback of media data if the identified user intent is a speech intent.
[0105] All or part of each module in the above-described audio signal noise reduction device may be implemented by software, hardware, or a combination thereof. Each of the above modules may be built into the processor in the computer device in hardware form or independent of this processor, or may be stored in the memory of the computer device in software form, in order to facilitate the processor calling and executing the operations corresponding to each of the above modules.
[0106] As shown in Figure 8, in one exemplary embodiment, the present application further provides a hearing aid 800 comprising a microphone 810, a bone conduction transducer 820, a processor 830, and a speaker 840, wherein the processor 830 is connected to the microphone 810, the speaker 840, and the bone conduction transducer 820. The bone conduction transducer 820 identifies the user's intent based on the captured vibration signal and transmits the identified user's intent to the processor 830.
[0107] The microphone 810 collects the audio signal and transmits the audio signal to the processor 830. The processor 830 performs the steps in the above audio signal noise reduction method, performs noise reduction processing on the audio signal, and transmits the noise-reduced audio signal to the speaker 840.
[0108] As those skilled in the art will understand, the structure of the hearing aid described above is only a part of the structure related to the solution of the present invention, and does not limit the hearing aids to which the solution of the present invention applies. A specific hearing aid may contain more or fewer components than those shown, may combine several components, or may have a different arrangement of components.
[0109] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal configuration diagram may be as shown in Figure 9. The computer device includes a processor, memory, an input / output interface (abbreviated as I / O), and a communication interface. The processor, memory, and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device provides computation and control functions. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the execution of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores signal-to-noise ratio level data and user hearing data in different environmental scenes, etc. The input / output interface of the computer device exchanges information between the processor and external devices. The communication interface of the computer device is connected to communicate with external terminals via a network. When the computer program is executed by the processor, a method for reducing audio signal noise is realized.
[0110] As those skilled in the art will understand, the structure shown in Figure 9 is merely a block diagram of some of the structures related to the solution of the present invention, and does not limit the computer equipment to which the solution of the present invention is applied. Specific computer equipment may include more or fewer components than those shown, may combine several components, or may have different component arrangements.
[0111] In one exemplary embodiment, a computer device is further provided that includes a memory in which a computer program is stored, and a processor that, when the computer program is executed, performs the steps in the embodiment of the audio signal noise reduction method described in any one of the above paragraphs.
[0112] In one embodiment, a computer-readable storage medium is provided on which a computer program is stored, and when the computer program is executed by a processor, the steps in the embodiment of the audio signal noise reduction method described in any one of the above paragraphs are realized.
[0113] In one embodiment, a computer program product is provided which includes a computer program, when executed by a processor, enables the implementation of the steps in the embodiment of the audio signal noise reduction method described in any one of the above paragraphs.
[0114] Furthermore, user information (including, but not limited to, user device information and user hearing information) and data (including, but not limited to, data for analysis, stored data, and displayed data) relating to this application are all approved by the user or fully approved by each party, and the collection, use, and processing of related data must comply with the relevant regulations.
[0115] As those skilled in the art will understand, all or part of the flows in the methods of the above embodiments can be implemented by a computer program instructing the relevant hardware, and the computer program may be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it may include the flows of the embodiments of each of the above embodiments. Any reference to memory, database or other medium used in each embodiment of the present application may include at least one of non-volatile memory and / or volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. Rather than being limited, RAM may take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database in the embodiment of this application may include at least one of relational databases and non-relational databases. Non-relational databases may include, but are not limited to, blockchain-based distributed databases.The processor in each embodiment of this application may be, but is not limited to, a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, programmable logic, a data processing logic based on quantum computing, or artificial intelligence (AI).
[0116] The technical features of the above embodiments can be combined in any way, and for the sake of explanation, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in these combinations of technical features, they should fall within the scope described in this application.
[0117] The embodiments described above merely illustrate some embodiments of the present application, and although the descriptions are specific and detailed, they should not be understood as limiting the scope of the claims of the present application. Furthermore, a person skilled in the art could make several modifications and improvements without departing from the concept of the present application, and these also fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the claims.
Claims
1. Based on the audio signal collected by the microphone, the system identifies the environmental scene corresponding to the audio signal. The process involves determining a first signal-to-noise ratio, which is the signal-to-noise ratio at which the user's hearing is compensated for and the user reaches the highest speech intelligibility level, and a second signal-to-noise ratio, which is the signal-to-noise ratio at which the user's hearing is compensated for and the user reaches the speech intelligibility threshold, A method for reducing noise in an audio signal, characterized by comprising performing noise reduction processing on the audio signal based on the first signal-to-noise ratio or the second signal-to-noise ratio.
2. Performing noise reduction processing on the audio signal based on the first signal-to-noise ratio or the second signal-to-noise ratio is: Identifying user intent, Based on the identified user intent, a target signal-to-noise ratio value is determined based on the first signal-to-noise ratio or the second signal-to-noise ratio. The method according to claim 1, characterized in that it includes performing noise reduction processing on the audio signal based on the signal-to-noise ratio target value.
3. Determining a target signal-to-noise ratio based on the first signal-to-noise ratio or the second signal-to-noise ratio for an identified user intent is: If the identified user intent is a speech intent, the first signal-to-noise ratio is determined as the target signal-to-noise ratio value. The method according to the second, characterized in that, if the identified user intent is a silence intent, a target signal-to-noise ratio value is determined based on the second signal-to-noise ratio.
4. Determining the target signal-to-noise ratio based on the second signal-to-noise ratio means that Based on the second signal-to-noise ratio, a recommended signal-to-noise ratio greater than the second signal-to-noise ratio is determined, The signal-to-noise ratio information, including the recommended signal-to-noise ratio value, is pushed, and the user adjusts based on the recommended signal-to-noise ratio value to obtain a target signal-to-noise ratio value. The method according to the previous version, characterized in that it includes receiving a target signal-to-noise ratio value fed back from the user.
5. Identifying the environmental scene corresponding to the aforementioned audio signal is: Extracting the first environmental features of the aforementioned audio signal, The method according to claim 1, characterized in that it includes matching the first environmental feature with a second environmental feature of a different environmental scene to identify the environmental scene corresponding to the audio signal.
6. Matching the aforementioned first environmental characteristics to the second environmental characteristics of a different environmental scene is, The second environmental features in different environmental scenes are clustered, and the clustering results are obtained. Determining the distance between each cluster center in the clustering result and the first environmental feature, The method according to claim 5, characterized by comprising determining an environmental scene characterized by the nearest cluster center as the environmental scene corresponding to the audio signal.
7. The process further includes extracting environmental features of the audio signal collected by the microphone, calling a trained audio signal processing model using the environmental features and user hearing feature data as input to identify the environmental scene corresponding to the audio signal, determining a first signal-to-noise ratio and a second signal-to-noise ratio that match the environmental scene, performing noise reduction processing on the audio signal based on the first signal-to-noise ratio or the second signal-to-noise ratio, and outputting the audio signal after noise reduction. The method according to any one of claims 1 to 6, characterized in that the audio signal processing model is obtained by training it based on hearing data of different users and a first signal-to-noise ratio and a second signal-to-noise ratio in different environmental scenes.
8. Before determining the first signal-to-noise ratio and the second signal-to-noise ratio that match the aforementioned environmental scene, In multiple environmental scenarios, we will conduct speech intelligibility threshold tests under information masking using different signal-to-noise ratio levels for the user, and obtain the test results. The method according to any one of claims 1 to 6, further comprising determining a first signal-to-noise ratio and a second signal-to-noise ratio for different environmental scenes of the user based on the test results and the user's hearing data.
9. Identifying user intent is Acquiring vibration signals from bone conduction transducers, The method according to claim 2, characterized in that it includes identifying the user's intent based on the vibration signal.
10. Identifying user intent based on the aforementioned vibration signal is Determining the vibration mode based on the aforementioned vibration signal, The method according to 9, characterized in that it includes identifying user intent based on the vibration mode.
11. The method according to 2, further comprising reducing the low-frequency gain and / or pausing playback of media data if the identified user intent is a speech intent.
12. A hearing aid comprising a microphone, a speaker, a bone conduction transducer, and a processor, wherein the processor is connected to the microphone, the speaker, and the bone conduction transducer, A hearing aid characterized in that the bone conduction transducer identifies a user intent based on the captured vibration signal and transmits the identified user intent to the processor; the microphone collects an audio signal and transmits the audio signal to the processor; and the processor performs the steps of the audio signal noise reduction method described in any one of claims 1 to 11, performs noise reduction processing on the audio signal, and transmits the noise-reduced audio signal to the speaker.