System and method for detecting one's own voice in a listening system
The system effectively detects self-voice content in audio signals by analyzing SPLs and symmetry levels between ipsilateral and contralateral audio signals, addressing the inefficiencies of conventional methods and enhancing performance with machine learning.
Patent Information
- Application Number
- JP2023533208
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-11-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-11-30
AI Technical Summary
Conventional methods for detecting self-voice content in audio signals require dedicated components and complex processing, which are inefficient and resource-intensive.
The proposed system uses a listening device with ipsilateral and contralateral microphones to detect self-voice content by analyzing the sound pressure levels (SPLs) of different spectral portions of the audio signal and determining the symmetry level between the ipsilateral and contralateral audio signals.
This approach allows for accurate detection of self-voice content without additional components, improving performance and reducing computational power, while also enabling dynamic adjustments through machine learning algorithms.
Smart Images

Figure 0007692483000001 
Figure 0007692483000002 
Figure 0007692483000003
Abstract
Description
Technical Field
[0001] Background Technical Information The listening device can be configured to provide a processed version of the audio content to enhance the user's listening. However, when the audio content includes the user's own voice (self-voice content), amplifying and / or processing such content in the same way as another detected audio content may produce an output that does not sound natural or is not beneficial to the user. Also, different processing strategies may be required or preferred for another audio content or self-voice content. Therefore, identifying self-voice content in audio content is important for the optimal performance of the listening device.
[0002] U.S. Patent Application Publication No. 20080189107 describes a method of attempting to identify self-voice content using the ratio of signal energy between the direct sound portion and the reverberant sound portion.
[0003] U.S. Patent No. 10025668 describes a listening system with left and right listening devices, each including an ear-hook microphone, an in-ear microphone, and an adaptive filter that attempts to detect the voice of the wearer of the listening device.
[0004] U.S. Patent No. 10616694 describes a listening device that analyzes sound with respect to a match with a self-sound type. The sound is identified as self-voice depending on the strength of the match between the sound and the self-voice.
[0005] Each of these conventional approaches for detecting self-voice content has the disadvantage that dedicated components and / or complex processing are required.
[0006] The accompanying drawings illustrate various embodiments and are part of this specification. The illustrated embodiments are merely examples and do not limit the scope of the present disclosure. Throughout the drawings, the same or similar reference numerals indicate the same or similar elements.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0008] Detailed Description Described herein are exemplary systems and methods for self-voice detection in a listening system. For example, the listening system may include an ipsilateral microphone associated with (e.g., disposed near) the user's ipsilateral ear and configured to detect an ipsilateral audio signal representing audio content, a contralateral microphone associated with (e.g., disposed near) the user's contralateral ear and configured to detect a contralateral audio signal representing audio content, and a listening device associated with the ipsilateral ear (e.g., configured to provide a processed version of the audio content). The listening device may be configured to identify a first sound pressure level (SPL) of a first spectral portion of the ipsilateral audio signal, identify a second SPL of a second spectral portion of the ipsilateral audio signal, identify that the first SPL is greater than the second SPL by at least a threshold SPL amount, and identify that a symmetry level between the ipsilateral audio signal and the contralateral audio signal is at least a threshold symmetry level. Based on identifying that the first SPL is greater than the second SPL by at least the threshold SPL amount and that the symmetry level is at least the threshold symmetry level, the listening device may be configured to identify that the audio content has own voice content representing the user's voice.
[0009] The systems and methods described herein can advantageously provide many benefits to a user of a listening device. For example, the listening devices described herein can provide an audio signal that more accurately reproduces audio content, including self-voice content, perceived by normal listening than conventional listening systems. Further, the systems and methods described herein can more accurately detect self-voice content without requiring additional components as compared to conventional listening systems. Additionally, the systems and methods described herein can detect self-voice content more reliably and quickly as compared to conventional listening systems, while at the same time having a lower computational power used. Further, the systems and methods described herein can, in some implementations, use machine learning algorithms to dynamically adjust one or more of the thresholds described herein, thereby improving the self-voice detection function of the systems and methods described herein over time. For at least these reasons, the systems and methods described herein can advantageously provide additional functionality and / or features to a user of a listening device as compared to conventional listening systems. These and other benefits of the systems and methods described herein will become apparent herein.
[0010] FIG. 1 shows an exemplary listening system 100 that can be used to convey sound to a user. The listening system 100 includes a first listening device 102-1 and a second listening device 102-2 (collectively referred to as "listening devices 102"). As represented by the positioning of the listening devices 102 in FIG. 1 with respect to the dashed line 104 and from the perspective of the listening device 102-1, the listening device 102-1 is associated with the user's ipsilateral ear, and the listening device 102-2 is associated with the user's contralateral ear. For example, if the listening device 102-1 is associated with the user's left ear, the listening device 102-2 is associated with the user's right ear. Alternatively, if the listening device 102-1 is associated with the user's right ear, the listening device 102-2 is associated with the user's left ear. As used herein, a listening device is "associated with" a particular ear by being configured to be worn on or in a particular ear and / or by providing listening capabilities to a particular ear.
[0011] The plurality of listening devices 102 can communicate with each other via a communication link 106 that may be wired or wireless and that can be useful for a particular implementation.
[0012] Each listening device 102 can be implemented by any type of listening device configured to provide or enhance the listening of a user of the listening system 100. For example, each listening device 102 can be implemented by a hearing aid configured to apply amplified audio content to the user, an acoustic processor included in a cochlear implant system configured to apply electrical stimulation representing audio content to the user, an acoustic processor included in an electroacoustic stimulation system configured to apply electroacoustic stimulation to the user, a head-mounted headset, ear-mounted earphones, earbuds, smart headphones, or any other suitable listening device. In some embodiments, listening device 102-1 is of a different type than listening device 102-2. For example, listening device 102-1 may be a hearing aid, and listening device 102-2 may be an acoustic processor included in a cochlear implant system. As another example, listening device 102-1 may be a unilateral hearing aid, and listening device 102-2 may be a CROS (contralateral routing of signals) hearing aid.
[0013] As shown, listening device 102-1 may include a processor 108-1, a memory 110-1, a microphone 112-1, and an output transducer 114-1. Similarly, listening device 102-2 may include a processor 108-2, a memory 110-2, a microphone 112-2, and an output transducer 114-2. Listening device 102 may include additional or alternative components that may be useful for a particular implementation.
[0014] The processor 108 (e.g., processor 108-1 and processor 108-2) is configured to execute various processing operations, such as processing of audio content received by the listening device 102 and mutual transmission of data. Each processor 108 can be implemented by any suitable combination of hardware and software. Any reference herein to an operation performed by a listening device (e.g., listening device 102-1) can be understood to be performed by the processor of the listening device (e.g., processor 108-1).
[0015] The memory 110 (e.g., memory 110-1 and memory 110-2) can be implemented by any suitable type of non-transitory computer-readable storage medium and can hold (e.g., store) data used by the processor 108. For example, the memory 110 may store data representing an operation program that specifies how each processor 108 processes audio content and sends it to the user. For the sake of explanation, if the listening device 102-1 is a hearing aid, the memory 110-1 can hold data representing an operation program that specifies an audio amplification method (e.g., amplification level, etc.) used by the processor 108-1 to send acoustic content to the user. As another example, if the listening device 102-1 is an acoustic processor included in a cochlear implant system, the memory 110-1 can hold data representing an operation program that specifies a stimulation method used by the listening device 102-1 to apply an electrical stimulus representing acoustic content to the user by instructing the cochlear implant.
[0016] The microphone 112 (e.g., microphones 112-1 and 112-2) can be implemented by any suitable audio detection device and is configured to detect an audio signal presented to a user of the listening device 102. As shown in FIG. 1, the microphone 112 can be included in the listening device 102 (e.g., embedded inside or on the surface of the listening device 102 or otherwise disposed on the listening device 102). One or both microphones 112 may alternatively be separate from and communicatively connected to their respective listening devices 102. For example, the microphone 112-1 can be removably attached to the listening device 102-1.
[0017] In this specification, the microphone 112-1 may be referred to as the ipsilateral microphone associated with the user's ipsilateral ear. Similarly, in this specification, the microphone 112-2 may be referred to as the contralateral microphone associated with the user's contralateral ear. By being disposed relatively close to a particular ear such that an audio signal presented to the particular ear is detected by the microphone, the microphone can be "associated with" the particular ear. For example, the microphone 112-1 may be configured to detect an audio signal presented to the ipsilateral ear (and for that reason, this audio signal may also be referred to as the "ipsilateral audio signal" in this specification). Similarly, the microphone 112-2 may be configured to detect an audio signal presented to the opposite ear (and for that reason, this audio signal may also be referred to as the "contralateral audio signal" in this specification). The ipsilateral audio signal and the contralateral audio signal may represent the same audio content (e.g., music, voice, noise, self-voice content, etc.), but may have different characteristics due to the different positions of the microphones 112.
[0018] The output transducer 114 can be implemented by any suitable audio output device. For example, the output transducer 114 may be implemented by a speaker of the listening device (also referred to as a receiver) or one or more electrodes of a cochlear implant system.
[0019] FIG. 2 shows an alternative listening system 200 that can be used in accordance with the systems and methods described herein. The listening system 200 is similar to the listening system 100 in that a listening device 102-1 associated with the user's ipsilateral ear is included in the listening system 200. However, as shown, the listening system 200 does not include a second listening device associated with the user's contralateral ear. Instead, the listening system 200 includes a contralateral microphone 202 that is associated with the user's contralateral ear and is communicatively connected to the listening device 102-1 via a communication link 204 that may be wired or wireless and that may be useful in certain implementations.
[0020] As described herein, a listening device (e.g., listening device 102-1 and / or listening device 102-2) can be configured to identify when the audio content represented by the ipsilateral audio signal and the contralateral audio signal detected by the ipsilateral microphone and the contralateral microphone, respectively, includes self-voice content. As will be described, this can be performed, at least in part, based on a comparison of the SPLs of different spectral portions of the ipsilateral audio signal.
[0021] FIG. 3 shows an exemplary graph 300 depicting the SPL for an audio signal including self-voice content. Graph 300 includes a y-axis 302 representing relative SPL with respect to an x-axis 304 representing relative distance. For example, the x-axis 304 shows two positions, namely a position 306 representing a position at the user's mouth and a position 308 representing a position at the user's ear. The solid line 310 depicts the SPL of the first spectral portion of the audio signal, and the dashed line 312 depicts the SPL of the second spectral portion of the audio signal. The first spectral portion corresponds to frequencies including the low-frequency range of the audio signal, whereas the second spectral portion corresponds to frequencies including the high-frequency range of the audio signal.
[0022] The frequency range for the first spectral portion may be any suitable frequency range lower than the remaining frequency range of the audio signal. For example, the low-frequency range may be a frequency band of any suitable width (e.g., 10 Hz to 2 kHz) centered at any suitable relatively low audio frequency (e.g., 500 Hertz (Hz) to 2 kHz). For example, the low-frequency range may be 800 Hz to 1200 Hz, 975 Hz to 1025 Hz, or any other suitable region. The SPL of the spectral portion may be any suitable SPL associated with the spectral portion, such as average SPL, median SPL, maximum SPL, minimum SPL, etc. The frequency range for the second spectral portion may be any suitable frequency range higher than the low-frequency range of the audio signal. For example, the high-frequency range may be a frequency band of any suitable width (e.g., 10 Hz to 2 kHz) centered at any suitable relatively high audio frequency (e.g., 4 kilohertz (kHz) to 10 kHz). For example, the high-frequency range may be 4 kHz to 7 kHz, 5 kHz to 6 kHz, or any other suitable region.
[0023] When the audio content includes self-voice content, as shown at position 306, the audio signal can originate from the user's mouth with a relatively similar SPL for both the low-frequency range and the high-frequency range. However, the low-frequency range and the high-frequency range may reach the ear via different acoustic paths. The low-frequency range of the audio signal (or at least a part of the low-frequency range of the audio signal) may be transmitted from the mouth to the ear via direct conduction through the user's head. However, the high-frequency range of the audio signal may not be able to conduct through the head and instead may be transmitted via an indirect path between the mouth and the ear (this includes paths via reflections from other surfaces). As a result, as shown at position 308, when the audio signal travels from the mouth to the ear, the SPL of the low-frequency range may be less attenuated than the SPL of the high-frequency range.
[0024] FIG. 4 shows an exemplary graph 400 further illustrating the SPL for an audio signal representing audio content that includes self-voice content. Graph 400 includes a y-axis 402 representing SPL with respect to an x-axis 404 representing frequency. The dashed line 406 represents the audio signal at the source of the audio signal. In this example, the audio signal may have the same SPL across the spectrum of the audio signal at the source of the audio signal, and thus, the dotted line 406 has the same SPL value for all values of frequency.
[0025] The solid line 408 represents the transmission of an audio signal to the user's ear when the audio signal represents self-voice content. As described with respect to graph 300, when the audio signal represents self-voice content, the low-frequency range of the audio signal is attenuated less than the high-frequency range of the audio signal. In contrast, the dashed line 410 represents the transmission (over the same distance) of an audio signal to the user's ear when the audio signal represents audio content that does not include self-voice content. As shown, when the audio content does not include self-voice content, the low-frequency range of the audio signal is attenuated by a relatively equal amount compared to the high-frequency range. This is because both frequency ranges travel the same acoustic path from the source of the audio content to the user's ear.
[0026] The contrast between the audio signal having self-voice content and the audio signal not having it is emphasized by arrow 412 and arrow 414. What these arrows show in the low-frequency range is that for the audio signal without self-voice content (arrow 414), the drop in SPL is relatively large compared to the audio signal with self-voice content (arrow 412). In contrast, arrow 416 and arrow 418 show that in the high-frequency range, the difference in the drop of SPL between the audio signal without self-voice content (arrow 416) and the audio signal with self-voice content (arrow 418) is relatively small. Rather, as shown in this embodiment, at a given frequency level, the audio signal having self-voice content may be attenuated more than the audio signal not having self-voice content. These differences between the SPLs for the spectral portions of the audio signal detected near the user's ear (e.g., nearby) can be factors that the listening device (e.g., listening device 102-1) can consider when identifying whether the audio signal represents audio content that includes self-voice content.
[0027] FIG. 5 shows an exemplary configuration 500 of a listening device 102 that may represent either the listening device 102-1 or the listening device 102-2 described herein. As shown, the listening device 102 receives a same-side audio signal 502-1 and a counter-side audio signal 502-2 (collectively referred to as the audio signal 502). As described, the same-side audio signal 502-1 may be detected by a same-side microphone (e.g., microphone 112-1), and the counter-side audio signal 502-2 may be detected by a counter-side microphone (e.g., microphone 112-2 or microphone 202).
[0028] The listening device 102 can perform various operations on the audio signal 502 to identify whether the audio signal 502 contains self-voice content, as represented by analysis functions 504-510. For example, as shown, the listening device 102 can perform a spectrum SPL analysis 504, a direction analysis 506, an overall SPL analysis 508, and / or a voice content analysis 510 on the same-side audio signal 502-1 and / or the counter-side audio signal 502-2 to identify whether these audio signals contain self-voice content. The listening device 102 can use any combination of one or more of these analysis functions 504-510 that may be useful in a particular implementation. For example, in some cases, based on the spectrum SPL analysis 504, the direction analysis 506 alone or in combination with the overall SPL analysis 508, and / or the voice content analysis 510, the listening device 102 may identify that the same-side audio signal 502-1 and / or the counter-side audio signal 502-2 contains self-voice content. Based on the processing of the same-side audio signal 502-1 and / or the counter-side audio signal 502-2 by one or more of the analysis functions 504-510, the listening device 102 can output self-voice identification data 512 indicating whether the audio signal 502 contains self-voice content. Each of the analysis functions 504-510 is described herein.
[0029] The listening device 102 can perform spectrum SPL analysis 504 in any suitable manner. For example, the listening device 102 can identify the first SPL of the first spectral portion of the ipsilateral audio signal. The first spectral portion may have frequencies included in a first frequency range. The listening device 102 can further identify the second SPL of the second spectral portion of the ipsilateral audio signal. The second spectral portion may have frequencies included in a second frequency range higher than the first frequency range. The listening device 102 can further identify whether the first SPL is greater than the second SPL by at least a threshold SPL amount. The threshold amount can be any suitable threshold SPL amount. For example, for audio content that does not include self-voice content, the average first SPL for the ipsilateral audio signal may be about 10 decibels (dB) higher than the second SPL. In contrast, for audio content that includes self-voice content, the average first SPL for the ipsilateral audio signal may be about 30 dB higher than the second SPL. Thus, the threshold SPL amount can be set for values between average difference values (e.g., 15 dB, 20 dB, 25 dB, etc.).
[0030] Additionally or alternatively, the listening device 102 can identify the ratio of the first SPL to the second SPL and determine whether this ratio is greater than a threshold ratio associated with the threshold SPL amount to determine whether the first SPL is greater than the second SPL by at least the threshold amount. The threshold SPL ratio can be any suitable SPL ratio indicating that the attenuation of the first spectral portion is less than that of the second spectral portion by the threshold SPL amount. Thus, the threshold SPL ratio can indicate that the SPL of the first spectral portion is greater than the SPL of the second spectral portion by at least the threshold SPL amount. For example, the threshold ratio can be 25 - 35 (e.g., 28 - 32, set to 30, or any other threshold within 25 - 35) or any other suitable ratio.
[0031] The listening device 102 can perform the direction analysis 506 in any suitable manner. For example, the listening device 102 (e.g., the direction / space classifier of the listening device 102) can identify the symmetry level between the ipsilateral audio signal 502-1 and the contralateral audio signal 502-2, and can compare this symmetry level with a threshold symmetry level. The listening device 102 can further use the head-related transfer function to identify the direction from which the audio signal 502 arrives with respect to the user. Since the mouth is in front of the user's ears, the audio signal generated by the mouth may appear to arrive from in front of the user (and / or may actually arrive from in front of the user after being reflected by an object).
[0032] When the audio signal from in front of the user is detected by the left and right ears, it can be relatively symmetric. Therefore, the listening device 102 can identify the symmetry level between the ipsilateral audio signal 502-1 and the contralateral audio signal 502-2. This symmetry level can be identified in any suitable manner, such as by comparing the SPL of the audio signal 502, the waveform shape of the audio signal 502, etc. The listening device 102 can identify whether the symmetry level is at least the threshold symmetry level. The threshold symmetry level can be any suitable threshold symmetry level. Additionally, the listening device 102 can further identify whether the relatively symmetric audio signal appears to arrive from in front of the user or behind the user. This is because the audio signal from behind the user can also be relatively symmetric. Such identification can be performed in any suitable manner, for example, by using the head-related transfer function, etc.
[0033] The listening device 102 can perform the overall SPL analysis 508 in any suitable manner. For example, the listening device 102 can identify the SPL of the ipsilateral audio signal 502 (e.g., over most or all frequencies). The SPL for an audio signal with self-voice content is generally higher overall than the SPL for an audio signal without self-voice content because the source of the self-voice content is the user's mouth and thus the distance from the user's ear is fixed. On the other hand, an audio signal without self-voice content is generally from a source that is farther from the user's ear than the user's mouth and thus generally has a relatively low overall SPL. The listening device 102 can compare the overall SPL with a threshold SPL to determine whether the audio content may include self-voice content. The threshold SPL can be any suitable SPL.
[0034] The listening device 102 can perform the voice content analysis 510 in any suitable manner. Self-voice content may generally include speech content, and thus the detection of speech content may be another factor used to determine whether the audio content includes self-voice content. Additionally, the fact that the overall SPL is generally relatively high may particularly apply to an audio signal representing audio content that includes speech content.
[0035] Based on one or more of these analyses, the listening device 102 can provide an output 512 for self-voice identification. In some embodiments, a machine learning algorithm can be used to optimize the detection of self-voice content based on these factors and other factors. In some embodiments, the self-voice identification may further be based on the self-voice identification of the contralateral listening device. Based on the analysis functions 504 - 510, both the ipsilateral and contralateral listening devices should arrive at the same determination as to whether the audio signal includes self-voice content. Thus, each device can further perform its respective self-voice identification based on the self-voice identification of the other listening device.
[0036] FIG. 6 shows an exemplary configuration 600 of a listening device 102 configured to implement such a machine learning algorithm, including a machine learning module 602. Configuration 600 shows a listening device 102 similar to configuration 500 with a machine learning module 602 added. The machine learning module 602 may be implemented using any suitable machine learning algorithm, such as neural networks (e.g., artificial neural networks (ANN), convolutional neural networks (CNN), deep neural networks (DNN) and / or recurrent neural networks (RNN), etc.), reinforcement learning, linear regression, etc. The machine learning module 602 can identify optimal parameters, weightings, etc. for various characteristics of the audio signal 502 analyzed by the listening device 102. For example, the machine learning module 602 can identify an optimal threshold for spectral SPL analysis 504, an optimal frequency range for spectral SPL analysis 504, a threshold for the symmetry level for direction analysis 506, a threshold for overall SPL analysis 508, etc. The machine learning module 602 may be trained in any suitable manner. For example, the machine learning module 602 may be configured to update the threshold based on identifying whether the audio signal 502 contains self-voice content. Such optimization is described herein. Additionally or alternatively, the machine learning module 602 can be trained using a supervised approach, e.g., using an initial dataset of audio signals labeled according to whether they contain self-voice content and / or receiving input from a user when the audio signal contains (or does not contain) self-voice content.
[0037] Configuration 600 shows the machine learning module 602 included in the listening device 102. Alternatively, the machine learning module may be implemented remotely and communicatively connected to the listening device 102 (e.g., in a smartphone, server, etc.). Additionally or alternatively, any of the analysis functions 504-510 can be executed in a remote device communicatively connected to the listening device 102.
[0038] FIG. 7 shows an exemplary graph 700 showing the SPL ratio for audio signals representing audio content with and without self-voice content. Graph 700 includes a y-axis 702 representing the SPL ratio and an x-axis 704 representing eight objects for which the sample SPL ratio is specified. For each object, S1-S8, the SPL was measured for the frequency range and the SPL ratio specified based on the SPL for an audio signal with self-voice content and an audio signal without self-voice content.
[0039] The solid line 706 shows the SPL ratio for the audio signal with self-voice content for objects S1-S8, and the dashed line 708 shows the SPL ratio for the audio signal without self-voice content for objects S1-S8. For example, the solid line 706-1 shows an SPL ratio of about 36 between the SPL for the high-frequency range and the SPL for the low-frequency range for the audio signal with self-voice content for object S1. The dashed line 708-1 shows an SPL ratio of about 27 between the SPL for the high-frequency range and the SPL for the low-frequency range for the audio signal without self-voice content for object S1.
[0040] There is a dashed line 710 between the solid line 706 and the dashed line 708, and this dashed line 710 may be an exemplary threshold between the SPL ratios for an audio signal having self-voice content and an audio signal having no self-voice content. For example, the dashed line 710-1 shows an SPL ratio of about 31 that can be used as the threshold SPL ratio for the object S1. Additionally or alternatively, the dashed line 712 shows an average threshold SPL ratio (e.g., an SPL ratio of about 30) specified based on the threshold SPL ratios for the objects S1 to S8. The average threshold SPL ratio can be used as a default threshold SPL ratio (e.g., an SPL ratio of 28 to 32), and this default threshold SPL ratio can then be adjusted based on the individual SPL ratios described herein.
[0041] FIG. 8 shows an exemplary flowchart 800 for identifying self-voice content by a listening device (e.g., the listening device 102). The listening device 102 can receive the ipsilateral audio signal and the contralateral audio signal, and in operation 802, it can identify the SPL for a first spectral portion including the low-frequency range of the ipsilateral audio signal. The SPL can be identified in any suitable manner. In operation 804, the listening device 102 can identify the SPL for a second spectral portion including the high-frequency range of the ipsilateral audio signal.
[0042] In operation 806, the listening device 102 can identify the SPL ratio between the SPL of the low-frequency range and the SPL of the high-frequency range. The SPL ratio can be identified in any of the ways described herein. For example, the SPL of the low-frequency range can be divided by the SPL of the high-frequency range. Additionally or alternatively, in the frequency domain, the SPL of the high-frequency range may be subtracted from the SPL of the low-frequency range. Additionally or alternatively, based on the SPL difference and the difference in the frequency range, the slope can be determined.
[0043] In operation 808, the listening device 102 can identify the symmetry level between the ipsilateral audio signal and the contralateral audio signal. The symmetry level can be identified in any of the ways described herein.
[0044] In operation 810, the listening device 102 can determine whether the audio signal appears to be arriving from in front of the user based on the symmetry level. In some embodiments, as described herein, this determination may further be based on the head-related transfer function. If the listening device 102 determines that the audio signal does not appear to be arriving from in front of the user (no in operation 810), in operation 812, the listening device 102 can determine that the audio content represented by this audio signal does not include self-voice content.
[0045] In some embodiments, also in operation 812, the listening device 102 can update the analysis parameters. For example, the listening device 102 can use the characteristics of the audio signal to identify and / or adjust a threshold value, whereby additional audio signals are compared to this threshold value to identify self-voice content. For example, the characteristics of the audio signal may include the overall SPL, SPL ratio, SPL for different spectral portions of the audio signal (e.g., for adjusting the frequency range for spectral portions), etc. Based on such characteristics, the listening device 102 can adjust the threshold SPL amount, overall SPL threshold, frequency ranges for the first and second spectral portions, threshold symmetry level, and / or any other threshold for detecting self-voice content. As described in connection with FIG. 7, in some embodiments, the machine learning module 602 can be used to make these adjustments. Additionally or alternatively, any other suitable process can be used to make these adjustments.
[0046] In some embodiments, the listening device 102 can also determine, in operation 814, whether the audio content includes voice content. The listening device 102 can analyze the audio signal in any suitable manner to detect voice content. If the listening device 102 determines that the audio content does not include voice content (No in operation 814), the listening device 102 can execute operation 812 based on the characteristics of the audio signal, thereby determining that the audio signal does not represent self-voice content and updating the analysis parameters accordingly. If the listening device 102 determines that the audio content includes voice content (Yes in operation 814), the listening device 102 can execute operation 816.
[0047] In operation 810, if the listening device 102 determines that the audio signal appears to be arriving from in front of the user (Yes in operation 810), in operation 816, the listening device 102 can determine whether the SPL ratio identified in operation 806 is at least the threshold SPL ratio. If the listening device 102 determines that the SPL ratio is less than the threshold SPL ratio (No in operation 816), the listening device 102 can execute operation 812, thereby determining, based on the characteristics of the audio signal, that the audio signal does not represent self-voice content and updating the analysis parameters accordingly. Thus, even for ipsilateral and contralateral audio signals having at least a threshold symmetry level, the listening device 102 can determine that the ipsilateral audio signal does not include self-voice content based on the ipsilateral audio signal not meeting the threshold SPL ratio. Conversely, even if having at least the threshold SPL ratio, the listening device 102 can determine that the ipsilateral audio signal does not include self-voice content based on the ipsilateral and contralateral audio signals not meeting the threshold symmetry level.
[0048] When the listening device 102 determines that the SPL ratio is at least the threshold SPL ratio (yes in operation 816), in operation 818, the listening device 102 can determine that the audio signal represents self-voice content. Therefore, the determination that the audio signal represents self-voice content is based on both the determination that the SPL ratio is at least the SPL ratio (yes in operation 816) and the determination that the audio signal appears to be arriving from in front of the user (yes in operation 810).
[0049] The listening device 102 can use this determination of self-voice content in any suitable way. For example, an audio signal containing self-voice content can be processed differently from an audio signal that does not contain self-voice content. Such processing may be configured to supply the user's own voice to the user in a way that sounds more natural to the user, thereby improving keyword detection, occlusion control, etc. For example, the listening device 102 may include various acoustic processing programs, some of which can be configured to process self-voice content. Such programs may be selected and / or adjusted based on the determination that the audio signal contains self-voice content. Additionally or alternatively, the self-voice content can be used in any suitable way, for example, supplied to a telephone for transmission, mixing of sidetone for a telephone, etc.
[0050] Furthermore, based on the identification that the audio signal represents self-voice content from the listening device 102, it is also possible to update the analysis parameters using the characteristics of the audio signal. For example, the attenuation in the low-frequency range compared to the high-frequency range may generally follow a recognizable pattern, but this pattern may vary based on each specific user. Furthermore, even for each specific user, the characteristics (and thus the optimal threshold) may vary based on the content of the voice, as well as the user's emotions, volume, health status, activities, acoustic environment, etc. Therefore, the listening device 102 can further update the analysis parameters based on the identification that the audio signal represents self-voice content. Any suitable machine learning algorithm can be used in the same way. In some embodiments, the analysis parameter values for the listening device 102 may first be programmed and / or trained using a machine algorithm based on the profile, characteristics, model, and / or voice samples of a specific user.
[0051] FIG. 9 shows an exemplary computing device 900 that may be specially configured to perform one or more processes described herein. Any of the systems, units, computing devices, and / or other components described herein may be implemented by the computing device 900.
[0052] As shown in FIG. 9, the computing device 900 may include a communication interface 902, a processor 904, a storage device 906, and an input / output ("I / O") module 908 that are communicatively connected to each other via a communication infrastructure 910. Although an exemplary computing device 900 is shown in FIG. 9, the components shown in FIG. 9 are not intended to be limiting. In other embodiments, additional or alternative components may be used. Next, the components of the computing device 900 shown in FIG. 9 will be described in more detail.
[0053] The communication interface 902 may be configured to communicate with one or more computing devices. Examples of the communication interface 902 include, but are not limited to, a wired network interface (such as a network interface card), a wireless network interface (such as a wireless network interface card), a modem, an audio / video connection, and any other suitable interface.
[0054] The processor 904 generally represents any type or form of processing unit that can process the data described herein and / or interpret, execute, and / or cause the execution of one or more instructions, processes, and / or operations. The processor 904 can perform operations by executing computer-executable instructions 912 (such as applications, software, code, and / or other executable data instances) stored in the storage device 906.
[0055] The storage device 906 may include one or more non-transitory computer-readable data storage media, devices, or configurations, and the storage device 906 can use any type, form, and combination of data storage media and / or devices. For example, the storage device 906 may include, but is not limited to, any combination of non-volatile media and / or volatile media described herein. Electronic data including the data described herein may be stored temporarily and / or permanently in the storage device 906. For example, data representing computer-executable instructions 912 configured to cause the processor 904 to perform any of the operations described herein may be stored within the storage device 906. In some embodiments, the data may be arranged in one or more databases resident within the storage device 906.
[0056] The I / O module 908 may include one or more I / O modules configured to receive user input and supply output to the user. The I / O module 908 may include any hardware, firmware, software, or combination thereof that supports input and output capabilities. For example, the I / O module 908 may include hardware and / or software for capturing user input, including, but not limited to, a keyboard or keypad, a touch screen component (e.g., a touch screen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.
[0057] The I / O module 908 may include one or more devices for presenting output to the user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O module 908 is configured to supply graphical data to a display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be useful in a particular implementation.
[0058] FIG. 10 shows an exemplary method 1000. One or more of the operations shown in FIG. 10 may be performed by any of the listening devices described herein. FIG. 10 shows exemplary operations according to one embodiment, but in other embodiments, any of the operations shown in FIG. 10 may be omitted, added, rearranged, and / or changed. Each of the operations shown in FIG. 10 may be performed in any of the ways described herein.
[0059] In operation 1002, the listening device associated with the ipsilateral ear identifies the first sound pressure level (SPL) of the first spectral portion of the ipsilateral audio signal, which has frequencies included in the first frequency range.
[0060] In operation 1004, the listening device identifies the second SPL of the second spectral portion of the ipsilateral audio signal, which has frequencies included in a second frequency range higher than the first frequency range.
[0061] In operation 1006, the listening device identifies that the first SPL is greater than the second SPL by at least a threshold SPL amount.
[0062] In operation 1008, the listening device identifies that the symmetry level between the ipsilateral audio signal and the contralateral audio signal is at least a threshold symmetry level.
[0063] In operation 1010, the listening device identifies that the audio content includes voice content.
[0064] In operation 1012, based on the identification that the first SPL is greater than the second SPL by at least a threshold SPL amount, the identification that the symmetry level is at least a threshold symmetry level, and the identification that the audio content includes voice content, the listening device identifies that the audio content has self-voice content representing the user's voice.
[0065] In the foregoing description, various exemplary embodiments have been described with reference to the accompanying drawings. However, it will be apparent that various modifications and changes can be made to the above embodiments without departing from the scope of the invention as set forth in the following claims, and additional embodiments can be implemented. For example, the particular features of one embodiment described herein can be combined with or replaced by the features of another embodiment described herein. Accordingly, the present specification and drawings should be regarded in an illustrative rather than a restrictive sense.
Claims
1. A listening system, wherein the listening system comprises: A left microphone (112-1) associated with the user's left ear and configured to detect a left audio signal (502-1) representing audio content; A right microphone (112-2, 202) associated with the user's right ear and configured to detect a right audio signal (502-2) representing the audio content; A listening device (102, 102-1) associated with the left ear; And having, the listening device is configured to: Identify a first sound pressure level of a first spectral portion of the left audio signal, the first spectral portion having frequencies included in a first frequency range; Identify a second sound pressure level of a second spectral portion of the left audio signal, the second spectral portion having frequencies included in a second frequency range higher than the first frequency range; Identify that the first sound pressure level is greater than the second sound pressure level by at least a threshold sound pressure level amount; Identify that a symmetry level between the left audio signal and the right audio signal is equal to or greater than a threshold symmetry level, the symmetry level being determined based on a difference in sound pressure levels between the left audio signal and the right audio signal; Based on the identification that the first sound pressure level is greater than the second sound pressure level by at least the threshold sound pressure level amount and the identification that the symmetry level is equal to or greater than the threshold symmetry level, identify that the audio content has self-voice content representing the user's voice. A listening system configured as described above.
2. The identification that the first sound pressure level is greater than the second sound pressure level by at least the threshold sound pressure level amount includes: Identifying a ratio between the first sound pressure level and the second sound pressure level; Identifying that the ratio is greater than a threshold ratio associated with the threshold sound pressure level amount. The listening system according to Claim 1, comprising the above.
3. The listening system according to Claim 2, wherein the first frequency range is 800 Hz to 1200 Hz, the second frequency range is 4 kHz to 7 kHz, and the threshold ratio is 25 to 35.
4. The listening device (102, 102-1) is further configured to identify an overall sound pressure level of the left audio signal. The identification that the audio content has self-voice content is further based on the overall sound pressure level of the ipsilateral audio signal, The listening device is configured to compare the overall sound pressure level with a threshold sound pressure level, and identify that the audio content has self-voice content when the overall sound pressure level is higher than the threshold sound pressure level. The listening system according to claim 1.
5. The listening device (102, 102-1) is further configured to identify that the audio content includes voice content based on the ipsilateral audio signal, The identification that the audio content has self-voice content is further based on the identification that the audio content includes voice content. The listening system according to claim 1.
6. The ipsilateral microphone (112-1) is configured to detect an additional ipsilateral audio signal representing additional audio content, The contralateral microphone (112-2, 202-1) is configured to detect an additional contralateral audio signal representing the additional audio content, The listening device further identifies that an additional symmetry level between the additional ipsilateral audio signal and the additional contralateral audio signal is smaller than the threshold symmetry level, and based on the identification that the additional symmetry level is smaller than the threshold symmetry level, identifies that the additional audio content does not have the self-voice content. The listening system according to claim 1, which is configured as described above.
7. The listening device (102, 102-1) further identifies a third sound pressure level of the first spectral portion of the additional ipsilateral audio signal, identifies a fourth sound pressure level of the second spectral portion of the additional ipsilateral audio signal, Based on the identification that the additional audio content does not have the self-voice content, adjusts the threshold sound pressure level amount based on the difference between the third sound pressure level and the fourth sound pressure level. The listening system according to claim 6, which is configured as described above.
8. The adjustment of the threshold sound pressure level amount includes using a machine learning algorithm. The listening system according to claim 7.
9. The ipsilateral microphone (112-1) is configured to detect an additional ipsilateral audio signal representing additional audio content, The contralateral microphones (112-2, 202) are configured to detect an additional contralateral audio signal representing the additional audio content, The listening device further comprises identifying a third sound pressure level of the first spectral portion of the additional ipsilateral audio signal, identifying a fourth sound pressure level of the second spectral portion of the additional ipsilateral audio signal, identifying that the third sound pressure level is greater than the fourth sound pressure level by less than the threshold sound pressure level amount, identifying that the additional audio content does not have the self-voice content based on the identifying that the third sound pressure level is greater than the fourth sound pressure level by less than the threshold sound pressure level amount, The listening system according to claim 1, which is configured as described above.
10. The listening device (102, 102-1) is further configured to identify that an additional symmetry level between the additional ipsilateral audio signal and the additional contralateral audio signal is equal to or greater than the threshold symmetry level, The identifying that the additional audio content does not have the self-voice content has no relation to the identifying that the additional symmetry level is equal to or greater than the threshold symmetry level, The listening system according to claim 9.
11. The listening device (102, 102-1) is further configured to adjust the threshold sound pressure level amount based on a difference between the third sound pressure level and the fourth sound pressure level, according to the listening system of claim 9.
12. The adjustment of the threshold sound pressure level amount includes using a machine learning algorithm, according to the listening system of claim 11.
13. The listening device (102, 102-1) has the ipsilateral microphone (112-1), according to the listening system of claim 1.
14. Furthermore, it has an additional listening device (102-2) associated with the contralateral ear and has the contralateral microphone (112-2), and the listening device and the additional listening device are configured to communicate with each other via a communication link (106), according to the listening system of claim 1.
15. A step of identifying a first sound pressure level of a first spectral portion of a same-side audio signal (502-1) representing audio content, the first spectral portion having frequencies included in a first frequency range, by a listening device (102, 102-1) associated with the user's same-side ear; A step of identifying a second sound pressure level of a second spectral portion of the same-side audio signal, the second spectral portion having frequencies included in a second frequency range higher than the first frequency range, by the listening device; A step of identifying, by the listening device, that the first sound pressure level is greater than the second sound pressure level by at least a threshold sound pressure level amount; A step of identifying, by the listening device, that a symmetry level between the same-side audio signal representing the audio content and a contralateral audio signal (502-2) is equal to or greater than a threshold symmetry level, wherein the symmetry level is obtained based on a difference in sound pressure levels between the same-side audio signal and the contralateral audio signal; A step of identifying, based on the identification that the first sound pressure level is greater than the second sound pressure level by at least the threshold sound pressure level amount and the identification that the symmetry level is equal to or greater than the threshold symmetry level, that the audio content has self-voice content representing the user's voice; A method comprising the above steps.
Citation Information
Patent Citations
Environment detection and adaptation in hearing assistance devices
EP1835785A2
A hearing device comprising an own voice detector
EP3328097A1
A hearing device comprising an acoustic event detector
EP3588981A1
Method and apparatus for rapid detection of one's own voice
JP2017535204A
Hearing device with own-voice detection and related method
JP2020102836A