Method and system for optimizing speaker output of PSAP and hearing aid
By using bone conduction sensors and air conduction microphones in PSAP or hearing aids to detect one's own speech and ambient sounds, and adjusting the speaker output based on the transfer function, the problem of wearers being dissatisfied with their own speech is solved, and a more natural speech experience is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-12
- Publication Date
- 2026-03-31
AI Technical Summary
When wearing PSAP or hearing aids, users are dissatisfied with their own voice, which is described as being too loud or unnatural, and current technology has not been able to effectively solve the problems of occlusion effect and bone conduction sound enhancement.
By using bone conduction sensors and air conduction microphones to detect one's own voice and ambient sounds, speaker output is automatically adjusted to optimize the voice experience based on a pre-estimated transfer function, including compensation function and operator adjustment.
It improves the wearer's auditory experience of their own voice, solves the problems of occlusion effect and bone conduction sound enhancement, and makes the voice sound more natural.
Smart Images

Figure CN121773631A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to techniques for improving the experience of hearing a person's own voice. Specifically, this disclosure relates to a method and system for optimizing the speaker output of a personal voice amplification product (PSAP) and a hearing aid to improve the wearer's auditory experience of their own voice by using bone sensors. Background Technology
[0002] Currently, there are two types of products on the market that amplify sound for users: personal sound amplification devices (PSAPs) and hearing aids. These two products have different intended uses. A PSAP is an electronic device designed for users with normal hearing to amplify sound in certain environments. For example, a PSAP can amplify received sounds and play them through a speaker that is typically inserted into the user's ear. In contrast, hearing aids can also amplify received sounds, but are designed to compensate for the wearer's hearing loss.
[0003] When wearing a PSAP or hearing aid, users often experience dissatisfaction with their own voice, even after professional PSAP or hearing aid fitting. This problem is typically described as the wearer hearing their own voice too loud, or ambient sounds or other people's voices not loud enough. Furthermore, the wearer's own voice may be perceived as unnatural, having excessive low frequencies, being muffled, booming, or boxy. The causes of these problems include the following:
[0004] First, there is bone conduction of the wearer's own voice. In addition to sound conducted through the air, there is also bone conduction sound that travels directly to the ear canal and cochlea, while other environmental sounds, including other people's voices, do not travel through the bone conduction pathway.
[0005] Secondly, the occlusion effect is a major contributing factor. Because PSAPs and hearing aids use closed earmolds that block the ear canal, the low-frequency sound pressure in the ear canal increases due to the increased acoustic impedance, resulting in a greater perceived sound, especially in the low-frequency range. As frequency increases, the energy of the speech signal transmitted to the tympanic membrane via bone conduction (BC) decreases significantly, while users with hearing loss typically have normal sensitivity in lower frequency bands but weaker sensitivity in higher frequency bands.
[0006] Third, air-conducted self-speech (i.e., air-conducted (AC) self-speech) can be louder. When an air-conducted microphone picks up the wearer's own speech, it is generally louder than other speech due to the shorter distance and less sound energy loss during air conduction.
[0007] Additionally, some PSAPs may amplify the wearer's own voice more than others due to beamforming gain of the wearer's own voice. For example, differential beamforming (designed for telephone calls) with the maximum response angle near the wearer's mouth is commonly used in PSAPs such as TWS (True Wireless Stereo) earbuds, which may result in further amplification of the wearer's own voice.
[0008] For PSAPs and hearing aids with occlusion structures, no perfect solution has yet been found for the aforementioned problems, especially regarding the occlusion effect. Wearers may eventually have to get used to it.
[0009] As PSAP functionality becomes increasingly popular in TWS products, and more and more individuals with hearing loss use hearing aids, it is both necessary and highly beneficial to develop a way to overcome these issues and provide users of such devices with a better auditory experience of their own voice. Summary of the Invention
[0010] According to one or more embodiments of this disclosure, a method is provided for improving the auditory experience of a wearer's own speech by optimizing the speaker output of a PSAP or hearing aid. The method may include: determining the presence of a user's own speech based on bone conduction signals received by a bone conduction sensor; determining the presence of ambient sound based on an air conduction microphone signal received by an in-air microphone in response to determining the presence of the user's own speech; and performing a first adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and a transfer function in response to determining the absence of ambient sound.
[0011] According to another aspect of this disclosure, a system is provided for improving the auditory experience of a wearer's own speech by optimizing the speaker output of a PSAP or hearing aid. The system may include: at least one bone conduction sensor configured to receive bone conduction signals; at least one air conduction microphone configured to receive air conduction microphone signals; a memory storing a pre-estimated transfer function; and at least one processor coupled to the at least one bone conduction sensor, the at least one air conduction microphone, and the memory. The at least one processor may be configured to: determine the presence of a user's own speech based on the bone conduction signals; determine the presence of ambient sound based on the air conduction microphone signals in response to determining the presence of the user's own speech; and perform a first adjustment operation on the speaker output based on the bone conduction signals, the air conduction microphone signals, and the transfer function in response to determining the absence of ambient sound.
[0012] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium including computer-executable instructions that, when executed by a computer, cause the computer to perform the methods disclosed herein. Attached Figure Description
[0013] Figure 1 It shows two different pathways through which the human ear and sound can reach the cochlea.
[0014] Figure 2 Examples of devices or systems equipped with both an air conduction microphone and a bone conduction sensor according to one or more embodiments of the present disclosure are shown.
[0015] Figure 3 An example is shown of estimating the transfer function between sound pressure measurements at the ear reference point (ERP) and the tympanic membrane reference point (DRP) with the ear open.
[0016] Figure 4 An example is shown of estimating the transfer function between sound pressure measurements at ERP and DRP with the ear closed.
[0017] Figure 5 An example is shown of estimating the transfer function between the self-speech signal received by a bone conduction sensor (such as a speech pickup (VPU) bone sensor) and the self-speech signal received at the DRP with the ear closed.
[0018] Figure 6 An example is shown of estimating the transfer function between the self-speech signal received by a bone conduction sensor (such as a VPU bone sensor) and the self-speech signal received at the DRP with the ear open.
[0019] Figure 7 A flowchart is shown of a method for optimizing speaker output of a PSAP (such as a TWS earbud) or hearing aid according to one or more embodiments of the present disclosure.
[0020] It is conceivable that an element disclosed in one embodiment may be advantageously used in another embodiment without special specification. The accompanying drawings referred to herein should not be construed as being drawn to scale unless otherwise specified. Furthermore, for clarity of illustration and explanation, the drawings are generally simplified and details or components are omitted. The drawings and discussion are used to explain the principles discussed below, wherein the same reference numerals denote the same elements. Detailed Implementation
[0021] Examples will be provided below for illustration. The descriptions of the various examples are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0022] To address the dissatisfaction with the auditory experience of self-voice from PSAPs (such as TWS earbuds with this function) or hearing aids, this disclosure proposes a novel method and system for optimizing the speaker output of a PSAP or hearing aid to improve the auditory experience of the wearer's self-voice. Specifically, the method can detect the wearer's self-voice activity (OVA) using signals captured by at least one air conduction microphone and a bone conduction sensor. The method can pre-estimate some transfer functions. Based on the estimated transfer functions and the signals received by at least one air conduction microphone and bone conduction sensor, the speaker output of the PSAP or hearing aid is automatically adjusted. The adjustment of the speaker output may include switching operators (which are process functions) based on different conditions. The adjustment of the speaker output may further include performing a compensation function on the speaker output based on the transfer functions, particularly a compensation function on the signals captured by the bone conduction sensor. References will follow below. Figures 1 to 7 The proposed methods and systems will be explained in detail.
[0023] Human auditory perception is related to the vibration of the cochlea. Vibration can be caused by sound. Sound can reach the cochlea in the inner ear by: 1) air conduction, which uses the ear canal as a propagation path, 2) bone conduction, which passes through the skull and the bones of the tissues, and 3) a combination of air and bone conduction. Figure 1 The human ear is shown, and two different propagation paths through which sound can reach the cochlea are illustrated. For example, bone conduction signal 102 can reach the cochlea through bones and tissues, while air conduction signal 104 can reach the cochlea through the ear canal. Figure 1 The approximate locations of the ear reference point (ERP) 106 and the tympanic membrane reference point (DRP) 108 are also shown, which will be described further later.
[0024] A PSAP, or hearing aid, typically adds gain to the signal received by an air-conducting microphone, and the gain at different frequency bands can vary depending on the loudness of the received signal. This can be achieved, for example, through a wide dynamic range compressor (WDRC) and other modules. While PSAPs and hearing aids operate differently, such as for different purposes and target uses, they share some similar electronic components, such as at least one air-conducting microphone and a loudspeaker. The device or system (e.g., a PSAP or hearing aid, or a component thereof) can be partially inserted into a user's ear. As used herein, the term "user" can be interchangeably referred to as "wearer," and the term "device" can be interchangeably referred to as "system," "PSAP," "hearing aid," or "part thereof."
[0025] Figure 2 Examples of devices or systems equipped with both air conduction microphones and bone conduction sensors according to one or more embodiments are shown. Figure 2 As shown, the device or system may include a bone conduction sensor 202, an air conduction microphone 204, a processor 206, and a speaker 208. It is understood that... Figure 2 This is only for illustrating the principles of the proposed methods and systems, and therefore only basic components are shown for clarity. Those skilled in the art will understand that the device or system may include other necessary components for implementing PSAP or hearing aid functions. For example, the device or system may further include a memory (not shown) for storing information or executable code or instructions used by a processor to perform the methods proposed in this disclosure. Note that although... Figure 2 Only one bone conduction sensor 202, one air conduction microphone 204, one processor 206, and one speaker 208 are shown, but the device or system may include more bone conduction sensors, air conduction microphones, processors, and speakers.
[0026] According to one or more embodiments, processor 206 may perform some processing (such as methods or algorithms described later) based on signals received by air conduction microphone 204 or signals from the air conduction microphone array and further processed by modules (such as beamformers) and signals received by bone conduction sensor 202, and provide the processed signals to speaker 208 for playback, thereby providing a better auditory experience of the user's own voice.
[0027] Directly measuring cochlear vibrations is difficult, if not impossible, to achieve. However, because the ear canal is close to the cochlea, it is assumed that the same vibrations of the cochlea also reach the ear canal, making it relatively easy to measure the ear canal sound pressure level (ECSP). Although there may be some differences between the vibrations in the cochlea and the ear canal, ECSP measurement is likely the best way to estimate cochlear vibrations. ECSP can be measured using a probe microphone inserted into the ear canal near the tympanic membrane.
[0028] The transfer function between the self-speech signal received by the air conduction microphone and the bone conduction sensor (such as the VPU bone sensor) can be easily estimated. For example, a human subject wears the device in an anechoic chamber and speaks and reads. The mathematical relationship between the recorded signals from the air conduction microphone and the VPU bone sensor is estimated as follows: ,in The frequency depends on the performance of the VPU and is typically below 1500 Hz. There may be some variation among different subjects due to changes in the skull and tissues, but the mean transfer function can be estimated and adaptively updated for the user when speaking in a quiet environment.
[0029] The transfer function between sound pressure measurements at the ear reference point (ERP) and the tympanic membrane reference point (DRP) with the ear open can be estimated. Figure 3 An example is shown of estimating the transfer function between sound pressure measurements at ERP 306 and DRP 308 with the ear open. An air-conducting microphone 302 is placed near the subject's ear. For the same ear, a probe microphone 304 is inserted into the subject's ear canal and close to the tympanic membrane; a tube 310 is required for safety reasons. Figure 3 As shown. Then, sound was played in an anechoic chamber, and signals from both microphones were recorded while the ears remained open. The mathematical relationship between the recorded signals from the air-conducting microphone 302 and the probe microphone 304 was estimated as follows: ,in This is the symbol for frequency. This experiment can also be performed using a Head and Torso Simulator (HATS), and some reference functions can be obtained from the IEEE standard. Note that in practice, slight differences may exist between the signals received at the air-conducting microphone and the ERP on the device due to the physical dimensions of the microphone and apparatus, but this will not significantly affect performance.
[0030] The transfer function between sound pressure measurements at ERP and DRP can be estimated when the ear is closed. Figure 4An example of estimating the transfer function between sound pressure measurements at ERP 406 and DRP 408 with the ear closed is shown. For cases where the ear canal is well isolated, two methods can be used to measure the signals at ERP 406 and DRP 408. The first method involves inserting an earplug into the ear, playing sound using its own speaker, and recording the signals from two microphones near ERP 406 and DRP 408 separately, similar to a reference method. Figure 3 The method described. Alternatively, a second method involves placing noise-isolating earmuffs 410 that provide high-quality sound on the ears, then playing sound using their own speakers and recording signals from two microphones 402 and 404 near ERP 406 and DRP 408, respectively. Both methods should be performed in an anechoic chamber to avoid any additional noise. Figure 4 The second method is shown. The mathematical relationship between the recorded signals from the air-conducting microphone 402 and the probe microphone 404 is estimated as follows: ,in It refers to the frequency. This experiment can also be performed using a Head and Trunk Simulator (HATS), and some reference functions can be obtained from the IEEE standard. Note that the method of isolating the ear is not limited to the two methods listed. Figure 4 As shown in the diagram.
[0031] The transfer function between the self-speech signal received by the bone conduction sensor when the ear is closed and the self-speech signal received at the tympanic membrane reference point (DRP) can be estimated. Figure 5 An example is shown for estimating the transfer function between the self-speech signal received by bone conduction sensor 502 (such as a VPU bone sensor) and the self-speech signal received at DRP 506 with the ear closed. (This is similar to the function used above for estimating...) Similar to the measurements, human subjects wore the device and spoke and read in an anechoic chamber. The mathematical relationship between the recorded signals from the bone conduction sensor 502 (such as a VPU bone sensor) and the probe microphone 504 was estimated as follows: ,in It's about frequency. The 508 tube is needed for safety reasons, and it's also convenient for controlling the open and closed states in the ear canal. For example, some soft silicone can be filled in for the closed state, and there will be no sound leakage. Silicon is... Figure 5 The shaded area is drawn in the middle. Note that the isolation method can be different from... Figure 4 The method is the same, but no sound is played.
[0032] The transfer function between the self-speech signal received by the bone conduction sensor when the ear is open and the self-speech signal received at the tympanic membrane reference point (DRP) can be estimated. Figure 6 An example is shown for estimating the transfer function between the self-speech signal received by bone conduction sensor 602 (such as a VPU bone sensor) and the self-speech signal received at DRP 606 with the ear open. (This is similar to the function used above for estimating...) Similar to the measurements, human subjects wore the device and spoke and read in an anechoic chamber, but a flexible tube 608 was added to allow sound to leak out of the ear canal. The long tube was positioned away from the subject's mouth and in a different direction to prevent air-conducted speech from entering the ear canal. The mathematical relationship between the recorded signals captured by the bone conduction sensor 602 and the probe microphone 604 was estimated as follows: ,in It is frequency. Through careful design, such as using different diameters (calculated) on both sides of the tube and placing the tube in the opposite direction to the mouth, the ear canal is in a state of almost openness, in which the bone conduction signal of one's own speech can leak out, but very little air conduction signal enters the ear canal.
[0033] The five transfer functions are listed together below.
[0034]
[0035] The method proposed in this disclosure aims to process signals received by an air conduction microphone and a bone conduction sensor, such that the processed signal, when played back by a speaker in a wearable device, sounds closer to the intended target sound. In the modeling and discussion of the transfer function described above, the transfer function for both open and closed ear states has been considered. It is understood in this disclosure that the speech heard by the ear in the closed state is the actual speech, while the speech heard by the ear in the open state is the intended target speech. Therefore, for ease of explanation, this disclosure will below provide some definitions of signals and some equations representing the relationships between signals under different conditions to clearly explain the basic principles of the invention in conjunction with the aforementioned transfer function.
[0036] When the user speaks, their own voice signal The speech signal reaches the air-conducting microphone with very little loss (i.e., the air-conducting microphone near the ERP), and the speech signal received by the air-conducting microphone (which can be interchangeably referred to as the air-conducting microphone signal received by the air-conducting microphone when only the user speaks) is represented as The signal then travels through the ear canal and reaches the tympanic membrane (DRP), and the arriving signal is represented as... The speaker's own speech signal travels through the skull and tissues to the bone conduction sensor, and the speaker's own speech signal received by the bone conduction sensor (which can be interchangeably referred to as the bone conduction signal received by the bone conduction sensor) is represented as... The speaker's own speech signal also reaches the tympanic membrane via bone conduction, and the arriving signal is represented as... Ambient sounds (which can be speech signals from other people or non-speech signals from the environment) reach the air-conducting microphone (i.e., reach the ERP), and the ambient sounds received by the air-conducting microphone are represented as... Then the ambient sound travels through the ear canal and reaches the tympanic membrane (DRP), and the arriving ambient sound is represented as... For simplicity, symbols are omitted in the following description. Although transfer functions and signals are represented in the frequency domain, it will be clear to those skilled in the art that transfer functions and signals can also be handled in the time domain. Some equations and relationships can be expressed as follows:
[0037]
[0038] When the ears are open
[0039]
[0040] When the ears are closed, and the speaker plays the same signal as received by the air-conducting microphone, then
[0041]
[0042] PSAP or hearing aid processing functions are represented as operators. Such as WDRC. Note that here... It represents a series of operations, rather than a simple matrix multiplication operator. It processes signals received by an air-conducting microphone. And generate output accordingly. The speaker in the device or system is shown below.
[0043]
[0044]
[0045] When the ear is open, the sound measurement of the speaker output of the device at DRP is expressed as
[0046]
[0047] Additionally, when a user speaks, their own speech travels through the bone conduction path and reaches the DRP, and the mixed signal at the DRP is represented as...
[0048]
[0049] When the ear is closed, the sound output of the speaker at the DRP is expressed as:
[0050]
[0051] In addition, the signal at the DRP, which is mixed with the bone conduction of the speech signal at the DRP, is represented as
[0052]
[0053] As mentioned earlier, additional bone conduction (BC) signals and occlusion effects cause the body's own speech signals to act as... The DRP, with its stronger low-frequency response, sounds loud, deep, and unnatural. This disclosure addresses this issue to improve the wearer's auditory experience of their own voice.
[0054] This disclosure aims to automatically change the operator in two situations. This refers to situations where only the user is speaking, or where another voice is being heard simultaneously from someone else. Both of these situations can be detected using bone conduction sensors and air conduction microphones. Currently, speech activity detection (VAD) algorithms based on air conduction signals are well-developed, but these algorithms struggle to distinguish whether the signal originates from the wearer's own voice or from nearby other people or media (such as TVs and radios). Generally, one's own voice is louder than other sounds, but this is not always the case. Bone conduction-based speech detection offers the following advantages in recognizing the wearer's own voice signal: The wearer's own voice signal captured by a bone conduction sensor typically has a high signal-to-noise ratio (SNR) and a distinct harmonic structure. Furthermore, voice signals from nearby other people or played in media will not be captured by the bone conduction sensor.
[0055] The following is for reference Figure 7 The description will go on to explain how the output of the speaker in a wearable device (e.g., a PSAP or hearing aid or a component thereof) can be automatically optimized and adjusted using air conduction and bone conduction microphones, based on a pre-estimated transfer function and depending on the different usage scenarios of the worn device (e.g., only the user (i.e., the wearer) speaks or there is another voice from someone else while the user is speaking).
[0056] Figure 7 A flowchart of a method for optimizing speaker output according to one or more embodiments of the present disclosure is shown. At S702, the method may determine the presence of a user's own voice based on bone conduction signals received by a bone conduction sensor. As mentioned above, this determination may be based on any well-developed voice activity detection (VAD) algorithm.
[0057] In some examples, if it is determined that the user's own voice is not present, the method proceeds to S704. At S704, the default operator is used, and no additional operations are performed. It is understood that the described device or system (e.g., a PSAP or hearing aid or a portion thereof) can often be pre-designed or directed to hear ambient sounds well in the fit, thereby obtaining a default operator tailored to the individual. The method is well-configured regardless of the user's own speech characteristics. When the user is not speaking, the method will not perform any additional operations, but will only use default operators representing a set of default operations or processing functions (such as WDRC functions) predefined in the PSAP or hearing aid.
[0058] In some implementations, if the presence of the user's own voice is determined, the method proceeds to S706. At S706, the method may determine the presence of any ambient sound based on the air conduction microphone signal received by the air conduction microphone. In some examples, this determination may be based on the correlation between the bone conduction signal received by the bone conduction sensor and the air conduction microphone signal received by the air conduction microphone. In a preferred example, this determination may be based on the correlation between the bone conduction signal and the convolution of the air conduction microphone signal and a first transfer function.
[0059] In some examples, when the user speaks and no one else speaks (here) This method detects self-voice activity (OVA) in both the bone conduction signal received by the bone conduction sensor and the air conduction microphone signal received by the air conduction microphone. In this case, there is a very high correlation between the two signals. More specifically, in the bone conduction signal... and There is a very high correlation between the air-conducting microphone signal and the convolution of the first transfer function.
[0060] In some implementations, if it is determined that no ambient sound is present, the method proceeds to S708. At S708, the method may perform a first adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and the transfer function. In some examples, the first adjustment operation includes a first change to the default operator and compensation for the bone conduction signal.
[0061] In some examples, based on the above equations and corresponding descriptions, it can be seen that the mixed signal at the DRP with the default operator can be described by the following equation:
[0062]
[0063] The target sound signal, which sounds as natural as sound heard with the ears open, is given by the following equation.
[0064]
[0065] To achieve the above objectives, a new operator described by the following equation is used. To perform the first operation,
[0066]
[0067] The air conduction compensation function used for ear canal permeability is, i.e., It can be used to adjust the default operator. And the bone conduction compensation function for ear canal permeability, i.e. It can be used to compensate for bone conduction signals.
[0068] The optimized speaker output is then described by the following equation:
[0069]
[0070] Therefore, it is expected that by using the new operator The mixed signal at the DRP is close to the target sound signal as described by equation (14) above.
[0071] Due to the default Typically used to hear ambient sounds, the self-speech signal captured by an air conduction microphone usually has a higher loudness. Bone conduction of self-speech also exists in the ear canal, so without further processing, self-speech would be too loud. To improve the loudness of self-speech, further adjustments are performed, which can be described by the following equation:
[0072]
[0073] in This represents the loudness compensation coefficient vector, and The elements are in the range (0, 1). The value is pre-tuned for each model of the device or system, and The two sets of values are used for beamforming on and off, respectively. Attention It is frequency The function is given, and it varies at each frequency range (bin). In practice, it is applied... The process can be an equalizer filter. Here This is used to compensate for situations where one's own speech is louder than ambient sound due to shorter distances and additional bone conduction, and also covers situations where beamforming techniques increase loudness. Transfer functions (such as...) They are used together to compensate for the occlusion effect.
[0074] In some implementations, if ambient sound is determined to be present, the method proceeds to S710. At S710, the method may perform a second adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and the transfer function. In some examples, the second adjustment operation includes a second change to the default operator and compensation for the bone conduction signal.
[0075] In some examples, where there is another voice from someone else while the user is speaking (in other words, ambient sound is present), the VAD module uses the signal captured from the bone conduction sensor (or a bone conduction sensor and an air conduction microphone together), and therefore can accurately detect its own speech. Based on the VAD results, the portions of the two signals (captured by the bone conduction sensor and the air conduction microphone) containing the user's own speech signal can be located. In each of these portions, the [unclear text - possibly a function or calculation] can be calculated. and Correlation between This is used to estimate the ratio of the power of its own speech signal to the power of the air-conducting microphone signal received by the air-conducting microphone (e.g., this ratio is described as...). Because they believed in their own voice With ambient sound Irrelevant.
[0076] In some examples, based on the above equations and corresponding descriptions, it can be seen that the mixed signal at the DRP with the default operator can be described by the following equation:
[0077]
[0078] The target sound signal, which sounds as natural as sound heard with the ears open, is given by the following equation:
[0079]
[0080] To achieve the above objectives, a new operator described by the following equation is used. To perform the second operation,
[0081]
[0082] The air conduction compensation function used for ear canal permeability is, i.e., It can be used to adjust the default operator. And the bone conduction compensation function for ear canal permeability, i.e. It can be used to compensate for bone conduction signals.
[0083] The optimized speaker output is then described by the following equation (16).
[0084] To balance the loudness of one's own speech with that of the surrounding sounds, further adjustments are made, which can be described by the following equation:
[0085]
[0086] in The values of the elements in the equation are the same as those in equation (17). Here and This is used to compensate for situations where one's own speech is louder than ambient sound due to shorter distances and additional bone conduction, and also covers situations where beamforming techniques increase loudness. Transfer functions (such as...) They are used together to compensate for the occlusion effect.
[0087] Note that due to the vibration frequencies of bones and tissues, bone conduction sensors can only detect low-frequency vibrations, therefore... The modifications are primarily in the low-frequency range, but they will still address the issue and provide a better auditory experience for users to hear their own voices when wearing the product.
[0088] In cases where bone conduction sensors may not be able to detect frequencies high enough to match the frequency range of bone conduction sound in the ear canal, spectral spreading with F0 estimation can be performed to compensate for the lack of signal components in slightly higher frequency ranges.
[0089] As described above, this disclosure describes the proposed method and a system that can implement the proposed method in various aspects. To overcome the problems of occlusion effects due to device insertion, the inclusion of additional bone conduction speech in the speaker's own voice, and higher loudness of the speaker's own voice due to shorter distances, the method pre-estimates five transfer functions between signals captured or measured using at least one air conduction sensor, a bone conduction sensor, and an in-ear probe microphone, respectively, with the ear open or closed. Based on the five transfer functions, the loudness compensation coefficient vector, and the signal-to-noise ratio estimated by the correlation between the bone conduction signal and the air conduction microphone signal, a PSAP or hearing aid processing function (i.e., operator) is calculated and adaptively applied accordingly to scenarios with only ambient sound, only speaker's own voice, or mixed sound.
[0090] It will be appreciated that the methods discussed above can be implemented by at least one processor included in a system for a personal sound amplification product (PSAP) or hearing aid. For example, the system may include a memory and at least one processor. The memory may be configured to store computer-readable instructions or code for causing at least one processor to implement the aspects described above in this disclosure. The processor may be any technically feasible hardware unit configured to process data and execute software applications, including but not limited to a central processing unit (CPU), microcontroller unit (MCU), application-specific integrated circuit (ASIC), digital signal processor (DSP) chip, etc.
[0091] As can be seen in the detailed description above, different features are grouped together in the examples. This manner of disclosure should not be construed as an intention to have more features than those expressly mentioned in each clause. Rather, aspects of this disclosure may include fewer features than all the features of the individual example clauses disclosed. Therefore, the following clauses should be regarded accordingly as incorporated into the specification, whereby each clause may serve as a separate example. Although each dependent clause may refer in the clause to a particular combination with one of the other clauses, the aspect of that dependent clause is not limited to that particular combination. It should be understood that other example clauses may also include combinations of aspects of a dependent clause with the subject matter of any other dependent or independent clause, or any feature combined with other dependent and independent clauses. The aspects disclosed herein expressly include these combinations unless expressly stated or readily inferred that no particular combination is intended to be used (e.g., contradictory aspects, such as defining an element as both an electrical insulator and an electrical conductor). Furthermore, it is intended that aspects of a clause may be included in any other independent clause, even if that clause does not directly depend on that independent clause.
[0092] Examples of implementation methods are described in the following numbered clauses:
[0093] Clause 1. A method for optimizing speaker output of a personal sound amplification product (PSAP) or hearing aid, comprising: determining the presence of a user's own voice based on a bone conduction signal received by a bone conduction sensor; determining the presence of ambient sound based on an air conduction microphone signal received by an air conduction microphone in response to determining the presence of the user's own voice; and performing a first adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and a transfer function in response to determining the absence of ambient sound.
[0094] Clause 2. The method according to Clause 1 further includes: in response to determining the presence of ambient sound, performing a second adjustment operation on the speaker output based on bone conduction signals, air conduction microphone signals, and transfer functions.
[0095] Clause 3. The method according to any one of Clauses 1 to 2, wherein the first adjustment operation includes a first change of the default operator and compensation for the bone conduction signal, wherein the default operator includes a series of processing functions predefined in the PSAP or hearing aid.
[0096] Clause 4. The method according to any one of Clauses 1 to 3, wherein the second adjustment operation includes a second change of the default operator and compensation for the bone conduction signal.
[0097] Clause 5. The method according to any one of Clauses 1 to 4, wherein the transfer function is pre-estimated and comprises: a first transfer function, the first transfer function being a transfer function between the self-speech signal received by the air conduction microphone and the bone conduction sensor; a second transfer function, the second transfer function being a transfer function between sound pressure measurements at the ear reference point (ERP) and the tympanic membrane reference point (DRP) with the ear open; a third transfer function, the third transfer function being a transfer function between sound pressure measurements at the ERP and the DRP with the ear closed; a fourth transfer function, the fourth transfer function being a transfer function between the self-speech signal received by the bone conduction sensor with the ear closed and the self-speech signal received at the DRP; and a fifth transfer function, the fifth transfer function being a transfer function between the self-speech signal received by the bone conduction sensor with the ear open and the self-speech signal received at the tympanic membrane reference point (DRP).
[0098] Clause 6. The method according to any one of Clauses 1 to 5, wherein the first change of the default operator is calculated based on the loudness compensation coefficient vector, the air conduction compensation function for ear canal permeability, and the air conduction microphone signal; wherein the air conduction microphone signal includes only the user's own speech; and wherein the air conduction compensation function is based on the second transfer function and the third transfer function.
[0099] Clause 7. The method according to any one of Clauses 1 to 6, wherein the second change of the default operator is calculated based on the correlation between the loudness compensation coefficient vector, the bone conduction signal and the air conduction microphone signal and the first transfer function, the air conduction compensation function for ear canal permeability, and the air conduction microphone signal; and wherein the air conduction microphone signal includes the user's own speech and ambient sound.
[0100] Clause 8. The method according to any one of Clauses 1 to 7, wherein compensation for bone conduction signals is calculated based on bone conduction signals, a third transfer function, a fourth transfer function, and a fifth transfer function.
[0101] Clause 9. The method according to any one of Clauses 1 to 8, wherein determining the presence of ambient sound comprises: calculating the correlation between the bone conduction signal and the convolution of the air conduction microphone signal and a first transfer function; and estimating the ratio of the power of the self-speech signal contained in the bone conduction signal and the air conduction microphone signal to the power of the air conduction microphone signal.
[0102] Clause 10. The method according to any one of Clauses 1 to 9 further includes using a default operator in response to determining that the user's own voice does not exist.
[0103] Clause 11. A system for optimizing speaker output of a personal sound amplification product (PSAP) or hearing aid, comprising: at least one bone conduction sensor configured to receive a bone conduction signal; at least one air conduction microphone configured to receive an air conduction microphone signal; a memory storing a pre-estimated transfer function; and at least one processor coupled to the at least one bone conduction sensor, the at least one air conduction microphone, and the memory, and configured to: determine the presence of a user's own voice based on the bone conduction signal; determine the presence of ambient sound based on the air conduction microphone signal in response to determining the presence of the user's own voice; and perform a first adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and the transfer function in response to determining the absence of ambient sound.
[0104] Clause 12. The system according to Clause 11, wherein at least one processor is further configured to: in response to determining the presence of ambient sound, perform a second adjustment operation on the speaker output based on bone conduction signals, air conduction microphone signals, and a transfer function.
[0105] Clause 13. The system according to any one of Clauses 11 to 12, wherein the first adjustment operation includes a first change of the default operator and compensation for the bone conduction signal, wherein the default operator includes a series of processing functions predefined in the PSAP or hearing aid.
[0106] Clause 14. The system according to any one of Clauses 11 to 13, wherein the second adjustment operation includes a second change of the default operator and compensation for bone conduction signals.
[0107] Clause 15. The system according to any one of Clauses 11 to 14, wherein the pre-estimated transfer function comprises: a first transfer function, the first transfer function being a transfer function between the self-speech signal received by the air conduction microphone and the bone conduction sensor; a second transfer function, the second transfer function being a transfer function between sound pressure measurements at the ear reference point (ERP) and the tympanic membrane reference point (DRP) with the ear open; a third transfer function, the third transfer function being a transfer function between sound pressure measurements at the ERP and the DRP with the ear closed; a fourth transfer function, the fourth transfer function being a transfer function between the self-speech signal received by the bone conduction sensor with the ear closed and the self-speech signal received at the DRP; and a fifth transfer function, the fifth transfer function being a transfer function between the self-speech signal received by the bone conduction sensor with the ear open and the self-speech signal received at the tympanic membrane reference point (DRP).
[0108] Clause 16. The system according to any one of Clauses 11 to 15, wherein the first change of the default operator is calculated based on the loudness compensation coefficient vector, the air conduction compensation function for ear canal permeability, and the air conduction microphone signal; wherein the air conduction microphone signal includes only the user's own speech; and wherein the air conduction compensation function is based on the second transfer function and the third transfer function.
[0109] Clause 17. The system according to any one of Clauses 11 to 16, wherein the second change of the default operator is calculated based on the correlation between the loudness compensation coefficient vector, the bone conduction signal and the air conduction microphone signal and the first transfer function, the air conduction compensation function for ear canal permeability, and the air conduction microphone signal; and wherein the air conduction microphone signal includes the user's own speech and ambient sound.
[0110] Clause 18. The system according to any one of Clauses 11 to 17, wherein compensation for bone conduction signals is calculated based on bone conduction signals, a third transfer function, a fourth transfer function, and a fifth transfer function.
[0111] Clause 19. The system according to any one of Clauses 11 to 18, wherein at least one processor is further configured to: calculate the correlation between the convolution of the bone conduction signal and the air conduction microphone signal and the first transfer function; and estimate the ratio of the power of the self-speech signal contained in the bone conduction signal and the air conduction microphone signal to the power of the air conduction microphone signal.
[0112] Clause 20. A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by a computer, cause the computer to perform a method pursuant to any one of Clauses 1 to 10.
[0113] Various embodiments have been described for illustrative purposes; however, these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, the practical application of technology found in the market or improvements to the technology, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0114] Reference has been made to the embodiments presented in this disclosure. However, the scope of this disclosure is not limited to the specific described embodiments. Rather, any combination of the foregoing features and elements, whether or not different embodiments are involved, is contemplated for implementing and practicing the contemplated embodiments. Furthermore, while the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, the scope of this disclosure is not limited regardless of whether a given embodiment achieves a specific advantage. Therefore, the foregoing aspects, features, embodiments, and advantages are merely illustrative and should not be considered as elements or limitations of the appended claims unless expressly stated in the claims.
[0115] The various aspects of this disclosure may take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or a combination of software and hardware implementations, all of which may be referred to herein as “circuit,” “module,” “unit,” or “system.”
[0116] This disclosure can be a system, method, and / or computer program product. A computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to perform aspects of this disclosure.
[0117] A computer-readable storage medium can be a tangible means capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punched cards or recessed protrusions on which instructions are recorded), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0118] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0119] This document describes aspects of the disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0120] These computer-readable program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which are executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in the flowchart and / or one or more block diagram boxes.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the figures. For example, depending on the function involved, two blocks shown consecutively may be executed substantially simultaneously, or sometimes the blocks may be executed in reverse order. It should also be noted that each block in the block diagram and / or flowchart illustrations, and combinations of blocks in the block diagram and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0122] While the foregoing relates to embodiments of this disclosure, other and additional embodiments of this disclosure may be contemplated without departing from the basic scope of this disclosure, the scope of which is determined by the appended claims.
Claims
1. A method for optimizing a speaker output of a personal sound amplification product (PSAP) or a hearing aid, comprising: determining whether own voice of a user is present based on a bone conduction signal received by a bone conduction sensor; in response to determining that own voice of the user is present, determining whether environmental sound is present based on an air conduction microphone signal received by an air conduction microphone; and in response to determining that environmental sound is not present, performing a first adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and a transfer function.
2. The method of claim 1, further comprising: in response to determining that the environmental sound is present, performing a second adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and a transfer function.
3. The method of any one of claims 1-2, wherein the first adjustment operation comprises a first change of a default operator and a compensation for the bone conduction signal, wherein the default operator comprises a pre-defined series of processing functions in the PSAP or hearing aid.
4. The method of any one of claims 1-3, wherein the second adjustment operation comprises a second change of the default operator and the compensation for the bone conduction signal.
5. The method of any one of claims 1-4, wherein the transfer function is pre-estimated and comprises: a first transfer function that is a transfer function between own voice signals received by the air conduction microphone and the bone conduction sensor; a second transfer function that is a transfer function between sound pressure measurements at an ear reference point (ERP) and a drum reference point (DRP) in an ear open state; a third transfer function that is a transfer function between sound pressure measurements at the ERP and the DRP in an ear closed state; a fourth transfer function that is a transfer function between own voice signals received by the bone conduction sensor and own voice signals received at the DRP in the ear closed state; and a fifth transfer function that is a transfer function between own voice signals received by the bone conduction sensor and own voice signals received at the drum reference point (DRP) in the ear open state.
6. The method of any one of claims 1-5, wherein the first change of the default operator is calculated based on a loudness compensation coefficient vector, an air conduction compensation function for ear canal venting, and the air conduction microphone signal; wherein the air conduction microphone signal includes only own voice of the user; and wherein the air conduction compensation function is based on the second transfer function and the third transfer function.
7. The method of any one of claims 1-6, wherein the second change of the default operator is calculated based on the loudness compensation coefficient vector, a correlation between the bone conduction signal and a convolution of the air conduction microphone signal and the first transfer function, the air conduction compensation function for ear canal venting, and the air conduction microphone signal; and wherein the air conduction microphone signal comprises own voice of the user and the environmental sound.
8. The method of any one of claims 1 to 7, wherein the compensation of the bone conduction signal is calculated based on the bone conduction signal, the third transfer function, the fourth transfer function, and the fifth transfer function.
9. The method of any one of claims 1 to 8, wherein determining whether the environmental sound is present comprises: calculating the correlation between the bone conduction signal and a convolution of the air conduction microphone signal and the first transfer function; and estimating a ratio of a power of an own voice signal contained in the bone conduction signal and the air conduction microphone signal to a power of the air conduction microphone signal.
10. The method of any one of claims 1 to 9, further comprising using the default operator in response to determining that own voice of the user is not present.
11. A system for optimizing a speaker output of a personal sound amplification product (PSAP) or a hearing aid, comprising: at least one bone conduction sensor configured to receive a bone conduction signal; at least one air conduction microphone configured to receive an air conduction microphone signal; a memory storing a pre-estimated transfer function; and at least one processor coupled to the at least one bone conduction sensor, the at least one air conduction microphone, and the memory, and configured to determine whether own voice of a user is present based on the bone conduction signal; in response to determining that own voice of the user is present, determine whether an environmental sound is present based on the air conduction microphone signal; and in response to determining that the environmental sound is not present, perform a first adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and the transfer function.
12. The system of claim 11, wherein the at least one processor is further configured to, in response to determining that the environmental sound is present, perform a second adjustment operation on the speaker output based on the bone conduction signal, the air conduction microphone signal, and the transfer function.
13. The system of any one of claims 11 to 12, wherein the first adjustment operation comprises a first change of a default operator and a compensation of the bone conduction signal, wherein the default operator comprises a pre-defined series of processing functions in the PSAP or hearing aid.
14. The system of any one of claims 11 to 13, wherein the second adjustment operation comprises a second change of the default operator and the compensation of the bone conduction signal. 15. The system of any one of claims 11 to 14, wherein the pre-estimated transfer functions comprise: a first transfer function that is a transfer function between a self voice signal received by the air conduction microphone and the bone conduction sensor; a second transfer function that is a transfer function between sound pressure measurements at an ear reference point (ERP) and a drum reference point (DRP) in an ear open state; a third transfer function that is a transfer function between sound pressure measurements at the ERP and the DRP in an ear closed state; a fourth transfer function that is a transfer function between a self voice signal received by the bone conduction sensor and the self voice signal received at the DRP in the ear closed state; and a fifth transfer function that is a transfer function between the self voice signal received by the bone conduction sensor and the self voice signal received at the drum reference point (DRP) in the ear open state.
16. The system of any one of claims 11 to 15, wherein the first change to the default operator is calculated based on a loudness compensation coefficient vector, an air conduction compensation function for ear canal venting, and the air conduction microphone signal; wherein the air conduction microphone signal includes only self voice of the user; and wherein the air conduction compensation function is based on the second transfer function and the third transfer function.
17. The system of any one of claims 11 to 16, wherein the second change to the default operator is calculated based on the loudness compensation coefficient vector, a correlation between the bone conduction signal and a convolution of the air conduction microphone signal and the first transfer function, the air conduction compensation function for ear canal venting, and the air conduction microphone signal; and wherein the air conduction microphone signal includes self voice of the user and the ambient sound.
18. The system of any one of claims 11 to 17, wherein the compensation of the bone conduction signal is calculated based on the bone conduction signal, the third transfer function, the fourth transfer function, and the fifth transfer function.
19. The system of any one of claims 11 to 18, wherein the at least one processor is further configured to: calculate the correlation between the bone conduction signal and a convolution of the air conduction microphone signal and the first transfer function; and estimate a ratio of a power of a self voice signal contained in the bone conduction signal and the air conduction microphone signal to a power of the air conduction microphone signal.
20. A non-transitory computer-readable storage medium comprising computer- executable instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 10.