Systems and methods for adaptive sound equalization in personal listening devices
By integrating passive sound suppression and adaptive sound equalization technology, incoming audio signals are processed in real time, solving the problems of loudness and spectrum balance in noisy environments, achieving better listening experience and hearing protection.
Patent Information
- Application Number
- CN202080050041.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-07
- Filing Date
- 2020-06-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-06-07
AI Technical Summary
The prior art is difficult to effectively reduce loudness and maintain spectrum balance of audio signals in noisy listening environments, resulting in poor listening experience and disadvantageous to hearing health.
Using technology that integrates passive sound suppression and adaptive sound equalization, we actively monitor and adjust the spectrum balance, dynamics and loudness characteristics of the audio signal to match the target performance criteria by processing incoming audio signals in real time.
It achieves effective reduction of loudness in noisy environments while maintaining optimized spectrum balance of audio signals, thereby improving listening experience and protecting hearing health.
Smart Images

Figure CN114128307B_ABST
Abstract
Description
[0001] Related Applications and Priority Claims
[0002] This application is related to and claims priority from U.S. Provisional Application No. 62 / 858,469, filed on Jun. 7, 2019, titled "ACTIVE AND PASSIVE NOISE CANCELATION WITH ACTIVE EQUALIZATION", which is hereby incorporated by reference in its entirety. BACKGROUND OF THE INVENTION
[0003] Some listening experiences, such as at a concert, can be sub - optimal perceptually. For example, a pop concert in a large venue may be rendered at extremely loud volumes, which may be neither to the listener's taste nor beneficial to the listener's hearing. Additionally, the music may be rendered with an unfavorable spectral imbalance, such as too much bass. Such problems also occur in other scenarios and are common obstacles to the listening experience. For example, a listener may wish to hear a conversation more clearly in a noisy environment (such as an airplane) or when using public transportation. In another example, a cyclist on a city street may want their listening device to be sound - permeable for safety reasons while still limiting the harmful effects of excessive noise pollution from passing vehicles or road works. Preferences regarding level and equalization can vary from person to person and may depend on the listening environment. Thus, listeners desire the ability to control and customize both the loudness and spectral equalization of the sound scene they experience based on the listener's personal preferences.
[0004] A common method for reducing loudness and protecting hearing in noisy listening environments such as concerts is to use foam earplugs. The purpose of using foam earplugs is to attenuate the sound reaching the eardrum by physically blocking the ear canal. While this does reduce the sound reaching the eardrum, such earplugs disproportionately reduce high - frequency content more than low - frequency content. This has a negative impact on the spectral balance of the listener's sound. Current solutions available to consumer listeners do not address the spectral balance problem. SUMMARY OF THE INVENTION
[0005] The Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. The Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0006] Listening experiences at live concerts and other scenarios can be sub - optimal perceptually. For example, they may be rendered at levels that are too high for the listener's comfort or auditory health and have a spectral balance that does not match the listener's preferences. Embodiments of the present invention address both the loudness problem and the spectral balance problem to improve the listening experience.
[0007] Embodiments of the systems and methods disclosed herein use novel techniques to integrate passive sound suppression and adaptive sound equalization to process incoming audio signals in real time in personal listening devices. Embodiments of the systems and methods described herein include new techniques for actively monitoring and adjusting the spectral balance, dynamics, and loudness characteristics of "real-world" auditory events such as musical performances to adaptively match target performance criteria. This has not been attempted by any other product or application and addresses common problems associated with the combination of effective hearing protection and improved sound quality in scenarios such as live acoustic performances.
[0008] Embodiments of the present invention use the earbud form factor for level attenuation when using earbuds. In some embodiments, the present invention uses acoustically sealed headset receivers. In some embodiments, the listener uses a pair of earbuds, one for each ear. In other embodiments, each earbud includes an external microphone for receiving incoming sound, an internal signal processing unit for processing the incoming sound, and a transducer for presenting the processed sound to the listener. In some embodiments, each earbud also includes an internal microphone to monitor the sound rendered to the listener. Other embodiments of the present invention use a headset form factor with acoustically sealed earcups. In some embodiments, the headset includes an external microphone for receiving incoming sound at each ear. In some embodiments, an internal microphone is also included at each ear to monitor the sound rendered to the listener.
[0009] A listener's preferences can vary depending on the listening environment. For example, in a purely noisy environment, the listener's preference may be to hear no sound from the environment. In other cases, the listener's preference may be to hear sound from the environment but at an attenuated level, such as maintaining some awareness of the environment. In other cases, the listener's preference may be to hear the environmental sound but with an improved spectral balance. In other cases, the listener's preference may be to selectively hear aspects of the environmental sound. In some cases, the environmental sound can change over time, such that continuously meeting the listener's preferences may require some adaptation to the environment. Thus, adaptively controlling the level and spectral balance of incoming sound to improve the listening experience is of interest. In some embodiments, the adaptive processing of incoming sound is configured to achieve a target level. In some embodiments, the adaptive processing of incoming sound is configured to achieve a target spectral balance.
[0010] Some embodiments of the present invention establish a target loudness for rendering sound, at least in part, based on user preferences. Some embodiments of the present invention establish a target loudness for rendering sound, at least in part, based on mandatory hearing protection guidelines. Some embodiments establish a target spectral balance for rendering incoming sound, at least in part, based on user settings. Some embodiments establish a target spectral balance by analyzing the incoming sound to determine the spectral balance based on the analysis results. For example, the analysis may determine that the incoming sound is jazz music, such that the user's preferred equalization for jazz music should inform the selection of the target balance.
[0011] Embodiments of systems and methods for adaptive sound equalization in personal listening devices are disclosed. In some embodiments, a method for processing an incoming audio signal applies active equalization to the incoming audio signal to obtain an equalized audio signal. This equalized audio signal is then tuned to a target equalization to obtain an output audio signal. The output audio signal is rendered for playback to a listener. The output audio signal is a perceptually improved version of the incoming audio signal, at least for a particular listener for which the target equalization is tuned.
[0012] In some embodiments, the method includes determining a target equalization based on knowledge of an artist's recording. Typically, this is the music artist whose recording is included in the incoming audio signal. In some embodiments, one or more machine learning techniques are used to analyze a database containing the artist's recordings. This allows embodiments of the system and method to determine a target equalization based on the database of the artist's recordings. In some embodiments, active filters are used to tune the equalized audio signal to the target equalization.
[0013] Embodiments of the present invention also include a method for processing an incoming audio signal, including actively monitoring the audio characteristics of the incoming audio signal and providing target data that includes target performance criteria. In some cases, this target data includes the listener's audiogram or measured hearing loss curve. This allows the incoming audio to be tuned or adapted such that certain frequencies where the listener may have a hearing impairment can be amplified. In some embodiments, the method includes emphasizing the conversational band to obtain an adapted audio characteristic. This allows the listener to hear the dialogue in a television program or movie that the listener might otherwise have difficulty hearing.
[0014] Embodiments of the method also include adapting the audio characteristics of the incoming audio signal to the target performance criteria to obtain an adapted audio characteristic of the incoming audio signal. In some embodiments, the adaptation of the audio characteristics is achieved by updating an adaptive filter based on the target performance criteria. The adapted audio characteristic is rendered in the output signal. This provides a better listening experience for the listener from the output signal compared to the incoming audio signal. In some embodiments, the adapted audio characteristic in the audio signal is rendered on a personal listening device.
[0015] In some embodiments, a microphone in a personal listening device receives sound from a listener's environment. The sound is then analyzed to determine one or more desired targets, such as a loudness level or a spectral balance. The determined targets are then used to control the adaptive processing of the sound received by the microphone to generate a perceptually improved sound for rendering to the listener.
[0016] Embodiments of the method also include identifying a song in an incoming audio signal to obtain the identified song and determining target data based on the identified song. Other embodiments include identifying the genre of a song in an incoming audio signal to obtain the identified genre and determining target data based on the identified genre. Other embodiments include identifying the audio characteristics of an unwanted sound in an incoming audio signal and determining target data based on the audio characteristics of the unwanted sound.
[0017] Embodiments also include an audible device for processing an incoming audio signal. The audible device includes a processor and a memory storing instructions. The instructions, when executed by the processor, configure the audible device to actively monitor the audio characteristics of the incoming audio signal. The listening device is also configured to provide target data including target performance criteria and to adapt the audio characteristics of the incoming audio signal to the target performance criteria to obtain adapted audio characteristics of the incoming audio signal. These adapted audio characteristics are rendered in an output signal for an improved auditory experience for the listener.
[0018] For purposes of summarizing the disclosure, certain aspects, advantages, and novel features of the invention have been described herein. It is to be understood that not necessarily all such advantages can be achieved in accordance with any particular embodiment of the invention disclosed herein. Accordingly, the invention disclosed herein can be implemented or carried out in a manner that realizes or optimizes one advantage or a group of advantages as taught herein, without necessarily realizing other advantages as taught or suggested herein.
[0019] It should be noted that alternative embodiments are possible, and the steps and elements discussed herein can be changed, added, or eliminated according to a particular embodiment. Without departing from the scope of the invention, these alternative embodiments include alternative steps and alternative elements that can be used, as well as structural changes that can be made. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Throughout the drawings, reference numerals are reused to indicate correspondence between reference elements. The drawings are provided to illustrate embodiments of the invention described herein and not to limit its scope.
[0021] Figure 1 A first exemplary embodiment of a personal listening device in an earbud form factor in accordance with an embodiment of the invention is illustrated.
[0022] Figure 2 Illustrates a second exemplary embodiment of a personal listening device in an earbud form factor in accordance with an embodiment of the present invention.
[0023] Figure 3 is in accordance with an embodiment of the present invention Figure 1 system block diagram of the audio processor shown in
[0024] Figure 4 is in accordance with an embodiment of the present invention Figure 3 block diagram of the central processor shown in
[0025] Figure 5 is in accordance with an embodiment of the present invention Figure 4 block diagram of the target determination module shown in
[0026] Figure 6 is in accordance with a first embodiment of the present invention Figure 4 block diagram of the adaptive processing module shown in
[0027] Figure 7 is in accordance with a second embodiment of the present invention Figure 4 block diagram of the adaptive processing module shown in
[0028] Figure 8 is a general flowchart of a method illustrating a first set of embodiments of the present invention.
[0029] Figure 9 is a general flowchart of a method illustrating a second set of embodiments of the present invention. DETAILED DESCRIPTION
[0030] As described in the Background and Summary above, the real-world listening experience may be perceptually degraded due to extreme loudness, poor spectral balance, or other factors. Thus, a useful solution will address the two goals of reducing loudness and maintaining a preferred spectral balance. With 32 million people attending at least one live music festival in 2014 (according to Nielsen), it is clear that there is a large pool of potential users who could benefit from an improved live music listening experience.
[0031] An example of a degraded listening experience is a large stage concert. Listeners at a typical rock concert may be exposed to sound levels averaging 120 dB SPL, which can cause short-term hearing problems such as tinnitus and long-term hearing problems such as permanent hearing loss. As a further degradation of the listening experience, large subwoofers and low-frequency acoustic modes in the venue can cause significant spectral imbalance and loss of performance clarity. Such degradation can depend on the seat location. Another example of a degraded listening experience is a blockbuster movie, where the sound level in the theater can be loud enough to be dangerous. In scenarios such as concerts or movies, listeners may prefer to limit the loudness to a desired level. Listeners may prefer to render the sound with a desired spectral balance. In some cases, the listener can select a spectral balance target to compensate for the listener's audiogram or to emphasize the spectral balance of dialogue. In some cases, the listener may prefer to render the sound with the spectral balance of the original program material without incurring the degraded spectral effects of the listening environment. Thus, embodiments of the present invention provide adaptive sound equalization and level adjustment in a personal listening device to address such preferences.
[0032] Listeners may prefer to adaptively reproduce external sounds in a personal listening device in other listening scenarios in addition to the above scenarios. For example, when having a conversation in a noisy environment such as an airplane or public transportation, the listener may prefer to attenuate background noise while enhancing the voice. As another example, a listener using a personal listening device (e.g., earbuds or headphones) to listen to music may not be able to hear potentially important sounds from the external environment, such as the sound of vehicle traffic or a public announcement. Such external sounds can be rendered to the listener while limiting the listener's exposure to overly loud sound levels that may be objectionable to the listener or harmful to the listener's hearing. Thus, embodiments of the present invention provide adaptive sound equalization and level adjustment in a personal listening device to enable improved rendering of external sounds of interest to the listener.
[0033] Embodiments of the present invention also include personal listening devices having an earbud form factor. Due to the physical occlusion of the ear canal, the earbud form factor provides some attenuation (or passive suppression) of incoming sounds from the user's environment. Embodiments of the present invention also include personal listening devices having a headphone form factor, which can similarly provide passive attenuation of sounds from the user's environment.
[0034] Passive loudness attenuation using earplugs is common at concerts. Some active personal listening devices are configured to apply a fixed gain or fixed equalization to sounds from the listener's environment. Embodiments of the present invention improve upon prior methods by adapting the sound rendered to the listener to a specific loudness level, which may be specified by the listener or specified according to guidelines for a safe listening level. Embodiments of the present invention improve upon prior methods by adapting the sound rendered to the listener to achieve a specified spectral balance, which may be specified by the listener or specified according to data associated with the sound to be rendered.
[0035] According to some embodiments of the present invention, Figure 1 FIG. illustrates a first exemplary embodiment of the present invention, as a personal listening device 100, having a form factor of an earplug or earbud that is inserted into the listener's ear canal. The depicted device includes an external microphone for receiving incoming sound, an audio processor for processing the received sound, and an internal transducer for rendering sound to the listener. It should be noted that an "external microphone" refers to a microphone that is open to the external listening environment. In contrast, an "internal microphone" is a microphone that is open to the listener's ear. Similarly, an "internal speaker" is a speaker that is open to the listener's ear. Typically, there is some type of physical acoustic barrier between the external microphone and the internal microphone. Although the figure shows a single element, in a typical embodiment, such a device is used in each ear of the listener.
[0036] Referring Figure 1 , the shape of the personal listening device 100 is designed to provide occlusion of the listener's ear canal, such that external sounds are physically blocked from reaching the ear canal. The mounted microphone 102 is open to the external environment to receive incoming acoustic signals. In other words, sound acoustically propagates from the listener's environment to the listener. The mounted microphone 102 transduces the received acoustic signal into an electrical form. The transduced electrical signal is provided to the audio processor 104, which in turn provides the processed signal to the speaker 106 that is inserted into the ear canal. The speaker 106 transduces the processed signal into an acoustic signal for transmission through the user's ear canal to the user's eardrum.
[0037] According to some embodiments of the present invention, Figure 2FIG. illustrates a second exemplary embodiment of the present invention, as a personal listening device 200, having a form factor of an earplug or earbud that is inserted into the ear canal of a listener. Although the figure shows a single element, in a typical embodiment, such a device is used in each ear of the listener. The personal listening device 200 includes an external microphone for receiving incoming sound, an audio processor for processing the received sound, an internal transducer for rendering sound to the listener, and an internal microphone for monitoring the sound rendered to the listener. The signal received by the internal microphone is also provided to the processor.
[0038] Referring Figure 2 , the shape of the personal listening device 200 is designed to provide occlusion of the ear canal such that external sound is physically blocked from reaching the ear canal. The mounted microphone 202 is open to the external environment to receive incoming acoustic signals. The microphone transduces the received acoustic signals into an electrical form. The transduced electrical signals are provided to the audio processor 204, which in turn provides the processed signals to the speaker 206, which is housed in the ear canal. The speaker 206 transduces the processed signals into acoustic signals for transmission through the user's ear canal to the user's eardrum.
[0039] The mounted microphone 208 is open to the ear canal to transduce internal acoustic signals into an electrical form. The signals in electrical form are provided to the audio processor 204 to monitor the acoustic signals in the ear canal, which are a combination of the acoustic signals rendered by the speaker 206 and any external sound that physically propagates into the interior. For example, external sound may not be completely blocked by the ear canal occlusion of the personal listening device 200 and may thus partially propagate into the ear canal.
[0040] Figure 3 is a block diagram of the audio processor 104 shown in Figure 1 some embodiments of the present invention. The audio processor 104 receives the output signal from a microphone (not shown) as an input on line 302. The analog / digital converter 304 converts the signal 302 from an analog signal to a digital signal, which is provided as an input to the processor 306.
[0041] The processor 306 outputs a processed digital signal, which is converted from a digital signal to an analog signal by the digital / analog converter 308. The output of the digital / analog converter is provided on line 310 for subsequent transduction by a speaker (not shown). Those of ordinary skill in the art will recognize that Figure 2 the audio processor 204 (or an audio processor for other embodiments of the present invention) shown in
[0042] In Figure 3 the block diagram shown, central processor 306 is connected to memory unit 312. In some embodiments, central processor 306 stores information in and retrieves information from memory unit 312. In some embodiments, memory unit 312 includes control logic 314. In some embodiments, memory unit 312 includes data unit 316. Data unit 316 stores data derived by central processor 306, such as energy measures of multiple frequency bands of an input audio signal. In some embodiments, data unit 316 stores data to be used by central processor 306. In some embodiments, data unit 316 stores the target loudness level of the sound to be rendered to a listener. In some embodiments, the target loudness is used by central processor 306 to at least partially determine the required processing of the processor input signal to determine the processor output signal. In some embodiments, data unit 316 stores the target spectral balance for rendering sound to a listener, sometimes referred to as spectral equalization. In some embodiments, the target spectral balance is used by central processor 306 to at least partially determine the required processing of the processor input signal to determine the processor output signal.
[0043] As described above, in some embodiments, data unit 316 stores target data for central processor 316 to use to determine the required processing of an input signal to generate an output signal. The target data can include the target average loudness level of the output signal. The target data can include the target peak loudness level of the output signal. The target data can include the target spectral balance or equalization of the output signal. The target data can include information other than the target loudness level or target spectral balance. In some embodiments of the present invention, the target data corresponds to a default setting. In some embodiments, the target data is at least partially based on a preset selected by the user. In some embodiments, the target data can be fixed. In certain cases, the target data can be time-varying. In some embodiments, the target data is at least partially based on user input. For example, the user selects a target loudness level or establishes a target spectral balance. In some embodiments, the target data is at least partially based on the analysis of signals transduced by an external microphone (such as Figure 1 the microphone 102 shown in Figure 2 or the microphone 202 shown in
[0044] For example, in a live concert scenario, an input signal or certain characteristics of the input signal can be streamed to a music recognition application that can identify a song and return target data associated with that song, such as a target loudness or a target spectral balance. This data can be used to at least partially determine the target data used by the central processor 316. In other cases, such as when such data is not available as part of a music recognition service, such song-specific target data can be pre-computed locally offline using the user's own content library, e.g., by analyzing the spectral balance and loudness characteristics of relevant portions of the library, such as songs by a particular artist or of a particular genre.
[0045] Figure 4 is a processing block diagram of the central processor 306 according to an embodiment of the present invention. Referring Figure 3 to Figure 4 , the central processor 306 includes an adaptive processing module 404 and a target determination module 408 that provides data to at least partially control the operation of the adaptive processing module 404. The central processor 306 receives a digital audio signal as an input on line 402. The digital audio signal is processed by the adaptive processing module 404, which provides the digital audio signal as an output on line 406. The digital audio signal on line 402 is further provided as an input to the target determination module 408. In some embodiments, the target determination module 408 determines the target data to be provided to the adaptive processing module 404 on line 410 at least partially based on the input digital audio signal on line 402. In some embodiments, the adaptive processing module 404 provides data to the target determination module 408 on line 412. In some embodiments, the target determination module 408 uses the data on line 412 to at least partially determine the target data to be provided to the adaptive processing module 404 on line 410.
[0046] In some embodiments, the adaptive module 404 is configured not to provide data to the target determination module 408. In some embodiments, the target determination module 408 is configured not to receive the digital audio signal on line 402. Those of ordinary skill in the art will understand and recognize that variations can be made to the configuration of the central processor 306. For example, some embodiments of the present invention include incorporating a second digital audio input signal into the processor according to the audio processor 204, as Figure 2 shown in
[0047] Figure 5 is a processing block diagram of the central processor 306 according to an embodiment of the present invention. Referring Figure 4Block diagram of the target determination module 408 as shown. In some embodiments, the target determination module 408 receives a digital audio signal as an input on line 402. In some embodiments, the target determination module 408 receives data from the adaptive processing module 404 on line 412.
[0048] In some embodiments, the feature determination module 502 determines features related to sounds in the listener's environment based at least in part on the digital audio signal on line 402. In some embodiments, the feature determination module 502 determines features related to sounds in the listener's environment based at least in part on data provided from the adaptive processing module 404 on line 412. The feature determination module 502 outputs the features to the feature mapping module 504. The feature mapping module 504 determines target data based at least in part on the features provided by the feature determination module 502. The feature mapping module 504 provides the target data to the adaptive processing module 404 on line 410.
[0049] In some embodiments, the feature mapping module 504 includes components for communicating with devices or systems external to the listening device 200 of the listening device 100. For example, these components include personal devices (such as smart phones) or cloud-based systems (such as data servers). The feature mapping module 504 transmits the features to the external device or system. The external device or system uses these features to determine target data. The external device or system transmits the target data to the feature mapping module 504.
[0050] In some embodiments, the external device or system uses these features to identify the song being played in the listener's environment and determines target data based on the identified song. In some embodiments, the external device or system uses these features to identify the genre of the music being played in the listener's environment and also determines target data based on the identified genre. In some embodiments, the external device uses these features to determine the characteristics of unwanted sounds (such as background noise) in the listener's environment and determines target data based on the determined characteristics.
[0051] Figure 6 is according to the first embodiment of the present invention Figure 4 Block diagram of the adaptive processing module 404 as shown. As Figure 6 shown, the adaptive processor module 404 receives an input signal X(ω) on line 402, and the input signal X(ω) corresponds to the digital form of the signal received by a microphone externally mounted to the listening device (such as Figure 1 the microphone 102 shown). The input signal X(ω) is provided as an input to the adaptive filter 602. The transfer function of the adaptive filter 602 is in Figure 6In the frequency domain, it is represented by H(ω). The input signal is further provided as an input to the adaptation control module 606 and the leakage model 604. The leakage model 604 at least partially includes a transfer function L(ω). The output L(ω)X(ω) of the leakage model 604 corresponds to an estimate of the sound from the listener's environment leaking into the listener's ear canal through the occlusion of the listening device (e.g., the personal listening device 100 shown in Figure 1 ). Those of ordinary skill in the art will understand that ω is a frequency variable and the signals and systems herein are considered in the frequency domain. For example, X(ω) is the frequency domain representation of the input signal. Those of ordinary skill in the art will understand that, for example, in the processor 306, digital signal processing can be performed in the frequency domain by combining an appropriate forward transform before frequency domain processing and an appropriate inverse transform after frequency domain processing. For example, a sliding window short-time Fourier transform is used as the forward transform, and an inverse short-time Fourier transform with overlap-add is used as the inverse transform.
[0052] In some exemplary embodiments, this leakage is considered in the adaptation control module 606. Target data T(ω) is provided to the adaptation control module 606 on line 410. The adaptation control module 606 provides control data to the adaptive filter 602 on line 608. For example, the adaptive filter 602 can be updated based on one or more inputs to the adaptation control module 606.
[0053] In some embodiments, the adaptation control module 606 is configured to determine a filter such that the sound received by the listener has a target spectral balance T(ω). For this mathematical development, the sound received by the listener is represented as the sum of the output of the adaptive filter 602 and the leakage of the external sound into the listener's ear canal through the occlusion of the personal listening device 100.
[0054] Mathematically, this is represented as:
[0055] Y(ω) = H(ω)X(ω) + L(ω)X(ω) (1) where Y(ω) represents the sound received by the listener, H(ω)X(ω) is the output of the adaptive filter, and L(ω)X(ω) is the estimate of the leakage.
[0056] The filter H(ω) is determined by minimizing a measure of the difference between the output spectrum Y(ω) and the target T(ω). In some embodiments, the squared error measure is used as the measure of the difference. In particular,
[0057] ∈(ω) = |Y(ω) - T(ω)| 2 . (2) Using equation (1), the optimal filter for the squared error measure is determined as
[0058]
[0059] Alternatively, in terms of statistical measures,
[0060]
[0061] the spectral and statistical measures in equations (3) and (4) may vary with time, and thus the optimal filter may vary with time. The time dependence can be included in equation (4) as follows:
[0062]
[0063] Estimate the statistical measures R XT (ω,t) and R XX (ω,t) (at frequency ω and time t) corresponding to the cross-correlation between the input and the target and the autocorrelation of the input, respectively, as would be understood by a person of ordinary skill in the art. A person of ordinary skill in the art will further understand that the statistical measures may change with time as the input signal, the target, or both change. In some embodiments, at processing time t, the adaptive filter 602 is adapted based on a combination of its filter settings at the previous processing time t-1 and the optimal filter provided by the adaptation control block 606 for time t, for example as
[0064]
[0065] where a is a suitably chosen forgetting factor.
[0066] Figure 7 is a block diagram of the adaptive processing module 404 according to a second embodiment of the present invention. Referring to Figure 4 the adaptive processor module 404 receives the input signal X(ω) on line 402, which corresponds to the digital form of the signal received by an externally mounted microphone of the listening device, such as Figure 7 the microphone 202 of the personal listening device 200 shown in Figure 2 . The input signal X(ω) is provided as an input to the adaptive filter 702. The transfer function of the adaptive filter 702 is represented by H(ω).
[0067] The signal Y(ω) is provided as an input on line 710 to the adaptation control module 706. The signal Y(ω) corresponds to the signal received by an internal microphone that monitors the sound in the listener's ear canal, such as Figure 2The signals received and transduced by the internally mounted microphone 208 of the personal listening device 200 shown in FIG. correspond. The target data T(ω) is provided to the adaptation control module 706 on line 410. The adaptation control module 706 provides control data to the adaptive filter 702 on line 708, such as an updated filter determined based on one or more inputs to the adaptation control module 706.
[0068] In some embodiments, the adaptation control module 706 determines an optimal filter based on a squared error metric, such as
[0069]
[0070] where the statistical measure R XZ (ω) corresponds to the cross-correlation between the input and leakage signal Z(ω) (calculated as Y(ω) - H(ω)X(ω)). Those of ordinary skill in the art will recognize that temporal correlation can be incorporated into Equation (7), as in Equation (5), resulting in:
[0071]
[0072] Figure 8 is a flowchart according to a first set of embodiments of the present invention. The method begins with receiving an acoustic signal at an external microphone (block 802). For example, this external microphone can be the externally mounted microphone 102 in the personal listening device 100, Figure 1 as shown in FIG. Next, the acoustic signal is transduced into an electrical form by the external microphone and then converted into a digital form, for example, by an analog-to-digital converter 304 (block 804). The digital input signal generated by the analog-to-digital converter is then received by a processor (such as processor 306) (block 806).
[0073] The features of the digital input signal are determined, for example, in a feature determination module 502 (block 808). The feature determination module 502 is a component of the target determination module 408, and the target determination module 408 is a component of the processor 306. Next, for example, in Figure 5 as shown in FIG., the target data is determined at least in part by the determined signal features (block 810). The feature mapping module 504 is a component of the target determination module 408, and the target determination module 408 is a component of the processor 306.
[0074] The process continues, such as in Figure 6The determination of the estimate of the external acoustic signal leaking into the listener's ear canal in the leakage module 604 shown in continues (block 812). The leakage module 604 is a component of the adaptive processing module 404, and the adaptive processing module 404 is a component of the processor 306. The digital input signal received by the processor in block 806, the target data determined in block 810, and the leakage estimate determined in block 814 are provided to an adaptation control module (such as the adaptation control module 606) (block 814). The adaptation control module 606 is a component of the adaptive processing module 404, and the adaptive processing module 404 is a component of the processor 306.
[0075] The process continues by updating the adaptive filter (block 816). In some embodiments, the adaptation control module 606 calculates the updated adaptive filter, for example, according to equation (5), and provides the updated adaptive filter to the adaptive filter 602 on line 608. Next, the adaptive filter is applied to the digital input signal to generate a digital output signal (block 818). For example, in the adaptive processing module 404, the adaptive filter 602 is applied to the input signal on line 402 to generate the output signal on line 406. Finally, the digital output signal is converted into analog form (block 820). For example, this can be done by the digital-to-analog converter 308. Then the analog signal is transduced into acoustic form, for example, by the internally mounted speaker 106 in the personal listening device 100.
[0076] Figure 9 is a flowchart of a second set of embodiments of the present invention. The process begins with receiving an acoustic signal at an external microphone (such as the externally mounted microphone 202 in the personal listening device 200), as shown in Figure 2 Next, the acoustic signal is transduced into electrical form by the external microphone and then converted into digital form, for example, by the analog-to-digital converter 304 (block 904). The digital input signal generated by the analog-to-digital converter is then received by the processor (such as the processor 306 shown in Figure 3 ).
[0077] The process continues by determining the characteristics of the digital input signal, for example, in the feature determination module 502 shown in Figure 5 The target data is determined at least in part by the determined signal characteristics (block 910), for example, in Figure 5in the feature mapping module 504 shown in. The feature mapping module 504 is a component of the target determination module 408, and the target determination module 408 is a component of the processor 306. Next, an acoustic signal is received at an internal microphone (e.g., the internally mounted microphone 208 in the personal listening device 200), which is transduced by the internal microphone into an electrical form and then converted into a digital form by an analog-to-digital converter (block 912). This signal can be referred to as a monitoring signal because it is used to monitor the sound in the listener's ear canal.
[0078] The digital monitoring signal determined in block 912 and the target data determined in block 910 are provided to an adaptation control module (e.g., Figure 7 the adaptation control module 706 shown in) (block 914). The adaptation control module 706 is a component of the adaptive processing module 404, and the adaptive processing module 404 is a component of the processor 306. The target data is then provided to the adaptation control module 706 on line 410 and the digital monitoring signal is provided to the adaptation control module 706 on line 710. Then the adaptive filter is updated (block 916). In some embodiments, the adaptation control module 706 calculates the updated adaptive filter, e.g., according to equation (8), and provides the updated adaptive filter to the adaptive filter 702 on line 708. Then the adaptive filter is applied to the digital input signal to generate a digital output signal (block 918). For example, in the adaptive processing module 404, the adaptive filter 702 is applied to the input signal on line 402 to generate the output signal on line 406. Finally, the digital output signal is converted into an analog form (block 920). For example, this can be performed by the digital-to-analog converter 308. Then, e.g., through the internally mounted speaker 206 in the personal listening device 200, the analog signal is transduced into an acoustic form.
[0079] Alternative Embodiments and Exemplary Operating Environments
[0080] Many other variations different from those described herein will be apparent from this document. For example, depending on the embodiment, certain actions, events, or functions of any of the methods and algorithms described herein can be performed in a different order, can be added, combined, or completely removed (such as, not all described actions or events are necessary for the practice of the methods and algorithms). Also, in certain embodiments, actions or events can be performed simultaneously, such as through multi-threading, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. Additionally, different tasks or processes can be performed by different machines and computing systems that can act together.
[0081] The various illustrative logical blocks, modules, methods, and algorithmic processes and sequences described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and process operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. For each particular application, the described functionality may be implemented in a different manner, but such implementation decisions should not be construed as causing a departure from the scope of this document.
[0082] The various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or executed by a machine, such as a general-purpose processor, a processing device, a computing device having one or more processing devices, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor and a processing device may be a microprocessor, but in the alternative, the processor may be a controller, a microcontroller, or a state machine, combinations thereof, and the like. The processor may also be implemented as a computing device, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0083] Embodiments of the systems and methods described herein may operate in a variety of types of general-purpose or special-purpose computing system environments or configurations. In general, a computing environment may include any type of computer system, including but not limited to a computer system based on one or more microprocessors, mainframe computers, digital signal processors, portable computing devices, personal organizers, device controllers, computing engines in appliances, mobile telephones, desktop computers, mobile computers, tablet computers, smart phones, and appliances having embedded computers, and the like.
[0084] Such computing devices can generally be found in devices having at least some minimum computing power, including but not limited to personal computers, server computers, handheld computing devices, laptop or mobile computers, communication devices such as cell phones and PDAs, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, audio or video media players, and so on. In some embodiments, the computing device will include one or more processors. Each processor can be a dedicated microprocessor such as a digital signal processor (DSP), very long instruction word (VLIW), or other microcontroller, or can be a conventional central processing unit (CPU) having one or more processing cores, including GPU-based cores in a multi-core CPU.
[0085] The processing actions or operations of the methods, processes, or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, in software modules executed by a processor, or in any combination of the two. The software modules can be included in a computer-readable medium accessible by a computing device. The computer-readable medium includes both volatile and non-volatile media, removable or non-removable, or some combination thereof. The computer-readable medium is used to store information such as computer-readable or computer-executable instructions, data structures, program modules, or other data. By way of example and not limitation, the computer-readable medium can include computer storage media and communication media.
[0086] Computer storage media includes but is not limited to: computer or machine-readable media or storage devices such as Blu-ray Discs (BDs), Digital Versatile Discs (DVDs), Compact Discs (CDs), floppy disks, tape drives, hard drives, optical drives, solid-state memory devices, RAM memory, ROM memory, EPROM memory, EEPROM memory, flash memory, or other memory technologies, magnetic tape cartridges, tapes, magnetic disk storage, or any other device that can be used to store the desired information and can be accessed by one or more computing devices.
[0087] The software modules can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of non-transitory computer-readable storage medium, media, or physical computer storage known in the art. Exemplary storage media can be coupled to the processor such that the processor can read / write information from / to the storage media. In an alternative, the storage media can be an integral part of the processor. The processor and the storage media can reside in an application specific integrated circuit (ASIC). The ASIC can reside in a user terminal. Alternatively, the processor and the storage media can reside as discrete elements in the user terminal.
[0088] As used in this document, the phrase "non-transitory" means "persistent or long-lived". The phrase "non-transitory computer-readable medium" includes any and all computer-readable media, with the sole exception of transitory propagating signals. By way of example and not limitation, this includes non-transitory computer-readable media such as register memory, processor cache, and random access memory (RAM).
[0089] The phrase "audio signal" is a signal that represents physical sound.
[0090] The retention of information such as computer-readable or computer-executable instructions, data structures, program modules, etc. can also be achieved by using a variety of communication media to encode one or more modulated data signals, electromagnetic waves (such as carrier waves), or other transmission mechanisms or communication protocols, and includes any wired or wireless information conveyance mechanism. Generally speaking, these communication media refer to signals whose one or more characteristics are set or changed in a way that encodes information or instructions in the signal. For example, communication media includes wired media (such as a wired network or a direct wire connection carrying one or more modulated data signals), and wireless media (such as acoustic, radio frequency (RF), infrared, laser, and other wireless media for transmitting and / or receiving one or more modulated data signals or electromagnetic waves). Any combination of the above should also be included within the scope of communication media.
[0091] In addition, some or all of the various embodiments of the systems and methods described herein, any combination or portion thereof of software, programs, computer program products can be stored, received, sent, or read in the form of computer-executable instructions or other data structures from any desired combination of a computer or machine-readable medium or storage device and a communication medium.
[0092] Embodiments of the systems and methods described herein can be further described in the general context of computer-executable instructions (such as program modules) executed by a computing device. Generally speaking, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments described herein can also be practiced in a distributed computing environment where tasks are performed by one or more remote processing devices or within a cloud of one or more devices linked by one or more communication networks. In a distributed computing environment, program modules can be located in local and remote computer storage media including media storage devices. Furthermore, the above instructions can be partially or fully implemented as hardware logic circuits that may or may not include a processor.
[0093] Unless otherwise stated or otherwise understood as used in the context, conditional language used herein (such as "can", "may", "could", "for example", etc.) generally intends to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or states. Thus, such conditional language generally does not imply that the features, elements, and / or states are required in any way by one or more embodiments or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or states are included in or to be performed in any particular embodiment without the input or prompting of an author. The terms "comprising", "having", etc. are synonymous and are used inclusively in an open-ended manner and do not exclude additional elements, features, acts, operations, etc. Moreover, the term "or" is used in its inclusive sense (rather than in its exclusive sense), such that when used, for example, to connect a list of elements, the term "or" refers to one, some, or all of the elements in the list.
[0094] While the foregoing detailed description has shown, described, and pointed out novel features as applied to various embodiments, it will be understood that various omissions, substitutions, and changes in the form and detail of the devices or algorithms shown can be made without departing from the scope of the present disclosure. As will be recognized, certain embodiments of the invention described herein may be embodied in forms that do not provide the features and advantages set forth herein, as some features may be used or practiced separately from other features.
Claims
1. A method for processing an incoming audio signal, comprising: Receiving an incoming acoustic signal from an external listening environment at a microphone of an audible device, the audible device being designed to be inserted into a listener's ear canal; Transducing the incoming acoustic signal and converting it into an incoming audio signal; Actively monitoring audio characteristics of the incoming audio signal; Providing target data including target performance criteria; Determining an estimate of leakage of the incoming acoustic signal through the audible device; Adapting the audio characteristics of the incoming audio signal to the target performance criteria to obtain adapted audio characteristics of the incoming audio signal by using an adaptive filter updated based on the incoming audio signal, the target data, and the estimate of the leakage; Rendering the adapted audio characteristics in an output signal; And Converting the output signal and transducing it into an acoustic signal using a speaker of the audible device, wherein the method further includes: Identifying a song in the incoming audio signal to obtain the identified song; And Determining the target data based on the identified song.
2. The method according to claim 1, further comprising rendering an adapted audio characteristic of the audio signal on a personal listening device.
3. The method according to claim 1, wherein the target data includes a listener's audiogram or a measured hearing loss curve.
4. The method according to claim 3, wherein adapting the audio characteristic of the incoming audio signal further comprises compensating for the listener's audiogram or the measured hearing loss curve.
5. The method according to claim 1, wherein adapting the audio characteristic of the incoming audio signal further comprises emphasizing the conversational band to obtain an adapted audio characteristic.
6. The method according to claim 1, further comprising: Identifying a genre of a song in the incoming audio signal to obtain the identified genre; And Determining the target data based on the identified genre.
7. The method according to claim 1, further comprising: Identifying audio characteristics of an unwanted sound in the incoming audio signal; and Determining the target data based on the audio characteristics of the unwanted sound.
8. An audible device for processing an incoming audio signal, the audible device having a shape designed to be inserted into a listener's ear canal and comprising: A microphone, open to the external listening environment and configured to receive an incoming acoustic signal from the external listening environment and convert the received incoming acoustic signal into an electrical form to provide an incoming audio signal; A processor; And A memory storing instructions which, when executed by the processor, configure the audible device to: Actively monitor audio characteristics of the incoming audio signal; Provide target data including target performance criteria; Determine an estimate of leakage of the incoming acoustic signal through the audible device; Adapt the audio characteristics of the incoming audio signal to the target performance criteria to obtain adapted audio characteristics of the incoming audio signal by using an adaptive filter updated based on the incoming audio signal, the target data, and the estimate of the leakage; and Render the adapted audio characteristics in an output signal; and A speaker, open to the listener's ear, the speaker transducing the output signal into an acoustic output signal, wherein the instructions, when executed by the processor, further configure the audible device to: Identify a song in the incoming audio signal to obtain the identified song; And Determine the target data based on the identified song.
Citation Information
Patent Citations
Hearing aid with Anti-occlusion effect techniques and ultra-low frequency response
US20090310805A1
Methods and devices for creating and modifying sound profiles for audio reproduction devices
US20150195663A1
Equalizer controller and controlling method
US20170230024A1
Methods and Systems for Automatically Equalizing Audio Output based on Room Characteristics
US20190103848A1