Systems, devices, and methods for acoustic transparency

By combining the adaptive design of acoustic transparency filters and feedback ANC filters, and integrating adaptive and compensation filters, the acoustic transparency problem of audible devices during active noise cancellation is solved, achieving personalized acoustic transparency adaptation and improving the user's perception of external sound.

CN115804105BActive Publication Date: 2026-03-27QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

When using active noise cancellation technology, existing hearing devices struggle to maintain acoustic transparency while reducing ambient noise, resulting in users being unable to clearly hear external sounds, especially due to the blocking effect of their own voice and perceptual inconsistencies caused by individual hearing differences.

Method used

It employs a combination of acoustic transparency filter and feedback ANC filter design, combined with adaptive filter and compensation filter, to achieve personalized compensation of acoustic transparency through adaptive adjustment of external and internal microphone signals, adapting to changes in earphone fit and hearing profile.

Benefits of technology

It achieves acoustic transparency during active noise cancellation, adapts to different users and environmental changes, provides a personalized listening experience, and improves the clarity of users' perception of external sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115804105B_ABST
    Figure CN115804105B_ABST
Patent Text Reader

Abstract

Methods, systems, computer-readable media, and apparatuses for audio signal processing are presented. An apparatus for audio signal processing includes a memory configured to store instructions and a processor configured to execute the instructions. The instructions, when executed, cause the processor to receive an external microphone signal from a first microphone and generate a transmissive component based on the external microphone signal and hearing compensation data. The hearing compensation data is based on a hearing map of a particular user. The instructions, when executed, also cause the processor to cause a speaker to generate an audio output signal based on the transmissive component.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims priority to commonly owned U.S. Provisional Patent Application No. 63 / 044,201 filed June 25, 2020, and U.S. Non-Provisional Patent Application No. 17 / 357,019 filed June 24, 2021, the contents of each of which are expressly incorporated herein by reference in their entirety. TECHNICAL FIELD

[0003] Aspects of the present disclosure relate to audio signal processing. BACKGROUND

[0004] Audible devices or “hearables” (such as “smart headsets,” “smart earphones,” or “smart earpieces”) are becoming increasingly popular. Such devices, which are designed to be worn over or in the ear, have been used for a variety of purposes, including wireless transmission and fitness tracking. As shown in FIG. 1, a hardware architecture of a hearable device typically includes a speaker for reproducing sound to the user’s ear; a microphone for sensing the user’s voice and / or ambient environmental sounds, and signal processing circuitry for communicating with another device (e.g., a smartphone). The hearable device can also include one or more sensors, for example, for tracking heart rate, for tracking physical activity (e.g., body movement), or for detecting proximity. In some examples, hearable devices can be worn in pairs, such as the hearable devices D10R and D10L of FIG. 2, which can communicate using wired signals or wireless signals WS10, WS20. Figure 1A Figure 1B Figure 1B

[0005] Figure 2 A schematic diagram of an implementation of a hearable device D10R configured to be worn at a user’s right ear is shown. The hearable device D10R can include, for example, hooks 214 or wings to secure the hearable device D10R in the concha and / or pinna of the ear; an ear tip 212 that surrounds a speaker 210 to provide passive acoustic isolation; one or more inputs 204 (such as switches and / or touch sensors for user control); one or more additional microphones 202, and one or more proximity sensors 208 (e.g., to detect that the device is being worn). BRIEF DESCRIPTION OF DRAWINGS

[0006] Aspects of the present disclosure are illustrated by way of example, and in which:

[0007] Figure 1A A block diagram of a hardware architecture of a hearable device is shown;

[0008] Figure 1B ​​​communication between audible devices worn at each ear of a user is shown;

[0009] Figure 2 a schematic diagram showing an implementation of an audible device is shown;

[0010] Figure 3A a block diagram of a system including a hear-through filter V(z) is shown;

[0011] Figure 3B a block diagram of a system including a feedback ANC filter -C(z) is shown;

[0012] Figure 3C a block diagram of a system including a hear-through filter V(z) and a feedback ANC filter -C(z) is shown;

[0013] Figure 4 another block diagram of a system of Figure 3C is shown;

[0014] Figure 5A a block diagram showing an implementation of a system as shown in Figure 4 is shown;

[0015] Figure 5B a block diagram showing an implementation of a system as shown in Figure 5A that receives a reproduced audio signal RX10 is shown;

[0016] Figure 6 a block diagram showing an implementation of a system of Figure 4 is shown;

[0017] Figure 7A a block diagram showing an implementation of a system as shown in Figure 6 including an apparatus A100 according to a particular configuration;

[0018] Figure 7B a block diagram showing an implementation PF20 of a pre-filter PF10;

[0019] Figure 8A a flowchart showing a method M100 according to a particular configuration;

[0020] Figure 8B a flowchart showing a method M200 according to a particular configuration;

[0021] Figure 9 a block diagram showing an implementation of a system of Figure 4 is shown;

[0022] Figure 10 an example of an audiogram for a user's left ear is shown;

[0023] Figure 11AAn implementation of the system as shown in Figure 9 is shown in a block diagram of an implementation of the system including an apparatus A 200 according to a particular configuration;

[0024] Figure 11B A block diagram of an apparatus A 250 corresponding to another implementation of the apparatus A 200 is shown;

[0025] Figure 12 An implementation of the system as shown in Figure 9 is shown in a block diagram of an implementation of the system;

[0026] Figure 13 A block diagram of an apparatus A 300 corresponding to the apparatuses A 100 and A 200 is shown;

[0027] Figure 14 A flowchart of operations for selecting a sound-transmission compensation filter state (e.g., hearing compensation data) based on a user's biometric authentication is shown;

[0028] Figure 15 An example of a speech authentication operation using a Gaussian mixture model is shown;

[0029] Figure 16A A flowchart of operations for selecting a sound-transmission compensation filter state based on a user's face recognition is shown;

[0030] Figure 16B An example of a face recognition operation using a trained neural network is shown;

[0031] Figure 17 An example of an ANC system including a feed-forward ANC filter is shown;

[0032] Figure 18 An example of an ANC system including an ANC filter with a fixed transfer function C(z) is shown;

[0033] Figure 19 An example of an ANC system in Figure 17 with a fixed filter - H(z) on the feedback path is shown;

[0034] Figure 20 A flowchart of audio signal processing based on hearing compensation data of a particular user is shown;

[0035] Figure 21 A schematic diagram of a device configured to perform audio signal processing based on hearing compensation data of a particular user is shown;

[0036] Figure 22 A schematic diagram of a headset configured to perform audio signal processing based on hearing compensation data of a particular user is shown; and

[0037] Figure 23 A schematic diagram of an extended reality (e.g., virtual reality, mixed reality, or augmented reality) headset configured to perform audio signal processing based on hearing compensation data for a particular user is shown. DETAILED DESCRIPTION

[0038] The principles described herein can be applied to, for example, a hearable device, a headset, or other communication or sound reproduction device configured to be worn at a user’s ear (e.g., over, on, or in the ear) (“personal audio device”). For example, such a device can be configured as an active noise canceling (ANC, also referred to as active noise reduction) device (“ANC device”). Active noise canceling is a technique for actively reducing acoustic noise (e.g., ambient environmental noise) by generating a waveform of an inverse form of a noise wave (e.g., having the same level and opposite phase) (also referred to as an “anti-phase” or “anti-noise” waveform). An ANC system typically uses one or more microphones to pick up an external noise reference signal, generates an anti-noise waveform from the noise reference signal, and reproduces the anti-noise waveform through one or more loudspeakers. This anti-noise waveform destructively interferes with the original noise wave (the primary disturbance (“d”) at the user’s ear) to reduce the noise level reaching the user’s ear.

[0039] Active noise canceling techniques can be applied to personal communication devices, such as cellular telephones, and sound reproduction devices, such as headsets and hearable devices, to reduce acoustic noise from the surrounding environment. In such applications, the use of ANC techniques can reduce the background noise level reaching the ear by as much as 20 decibels or more, while transmitting useful sound signals, such as music and far-end speech. For example, in a headset for a communication application, the device typically has a microphone and a loudspeaker, where the microphone is used to capture the user’s voice for transmission and the loudspeaker is used to reproduce received signals. In this case, the microphone can be mounted on a boom or in an ear cup or earbud (also referred to as an “earpiece”), and / or the loudspeaker can be mounted in an ear cup or earbud. In another example, the microphone is mounted on a pair of glasses (of a smart pair of glasses or other head-mounted device or display) near the user’s ear.

[0040] ANC devices typically have a microphone (e.g., an external reference microphone) arranged to generate a reference signal (“x”) based on ambient environmental noise and / or a microphone (e.g., an internal error microphone) arranged to generate an error signal (“e”) based on noise-cancelled sound output. In either case, the ANC device uses the microphone input to estimate the noise at that location and produces an anti-noise signal (“y”), which is a modified version of the estimated noise. The modification typically includes filtering with a phase inversion, and can also include gain amplification.

[0041] ANC devices typically include an ANC filter that models the acoustic primary path ("P(z)") between an external reference microphone and an internal error microphone and generates an anti-noise signal that matches the acoustic noise in amplitude and opposes the acoustic noise wave in phase. In a typical feed-forward design, for example, a reference signal x is modified by passing it through an estimate of the secondary path ("S(z)") from the ANC filter output through the electro-acoustic path of, for example, a loudspeaker and error microphone to produce an estimated reference x' that is used to adapt the state of the ANC filter (e.g., filter gain and / or tap coefficient values). In a typical feedback design, the error signal e is modified to produce the estimated reference x'. The ANC filter is typically adapted according to an implementation of the least mean square (LMS) algorithm, such as the filtered reference ("filter-X") LMS algorithm, the filtered error ("filter-E") LMS algorithm, the filtered-U LMS algorithm, and variants thereof (e.g., sub-band LMS algorithms, step-size normalized LMS algorithms, etc.). Signal processing operations such as time delay, gain amplification, and equalization or low-pass filtering can be performed to improve noise cancellation.

[0042] ANC systems can be effective at canceling ambient environmental noise. Unfortunately, ANC devices can hinder a user's ability to hear desired external sounds even when the ANC system is inactive. Passive attenuation of the device can make environmental sounds difficult to perceive when the user is wearing the personal audio device. Even when the ANC system is turned off, a user wearing ear muffs or earplugs often needs to remove the device to hear announcements or converse with others because the device mutes or blocks external sounds from the user's ear canal.

[0043] For example, it can be desirable to make a personal audio device acoustically transparent so that a user hears the same sounds as when not wearing the device. The device can be configured to, for example, pass external sounds into the user's ear canal. While the device can provide a "pass-through" mode that passes ambient sounds into the ear, the perception of acoustic transparency can be inadequate and the user can be forced to remove the device because the desired perception of acoustic transparency is not met.

[0044] Several illustrative configurations will now be described with respect to the drawings, which form a part of this disclosure. Although specific configurations can be described in detail below, other configurations can be used, and various modifications can be made without departing from the scope of the disclosure or the appended claims. The solutions described herein can be implemented on a chip set.

[0045] One aspect of providing acoustic transparency is to pass ambient sounds so that a user can hear them as if not wearing the device.Figure 3A A block diagram of the system is shown, where an external reference signal x(n) (the desired air-conducted ambient sound) is filtered by a primary path P(z) (e.g., the passive attenuation of the device) to produce a primary disturbance d(n) at the user's ear. Due to the passive attenuation, the disturbance reaching the user's ear does not sound like the external reference signal x(n).

[0046] Figure 3A The system of includes a trans-aural filter V(z) designed such that its output, when passed through a secondary path S(z), adds up with d(n) to provide an acoustically transparent response. As Figure 3A shown, the trans-aural filter V(z) can be designed (e.g., based on online models of the loudspeaker response and the passive attenuation) to have a transfer function (1 - P(z)) / S(z) such that the error signal e(n) approximates x(n). The coefficients of V(z) can be computed by an iterative gradient descent algorithm, and a filter modeling the primary path P(z) can be computed using an implementation of the LMS algorithm with the internal and external microphone signals as inputs. When the acoustic models S(z) and P(z) used to compute the trans-aural filter V(z) are good enough estimates of the true time-varying responses S(t,z) and P(t,z), this structure can be expected to generate a proper transparent response.

[0047] A second aspect of providing acoustic transparency is that, in addition to blocking ambient sound, the passive attenuation can also affect the user's perception of their own sound ("self-sound"). This muffling of the air-conducted component of self-sound due to occlusion of the ear canal is known as the "occlusion effect." This occlusion effect is characterized by under-emphasis of high-frequency sounds and over-emphasis of low-frequency sounds (e.g., due to conduction through the bone and soft tissue), which can give the user the perception of speaking underwater.

[0048] In the absence of air-conducted sound (e.g., due to the passive attenuation of the device), the error signal e(n) is primarily the user's own sound as conducted within the user's head. Figure 3B A block diagram of the system is shown, where a feedback ANC filter - C(z) is used to generate an anti-noise signal y(n) to cancel the error signal e(n). As Figure 3B shown, the transfer function of this system from d(n) to e(n) (including the secondary path S(z)) can be characterized as H(z) = 1 / [1 + C(z)S(z)].

[0049] Figure 3C A block diagram of the system is shown, where the two aspects described above are combined. In Figure 3CIn this system, the output of V(z) is modified based on an estimate of the secondary path S(z) and subtracted from the internal microphone signal EM10 to generate the input to the feedback ANC filter FB10. The output of the feedback ANC filter FB10 is combined with the output of the trans-aural filter HF10 to generate the audio output signal AO10 that is used to drive the loudspeaker. The filter V(z) is designed to have a transfer function [1 - P(z)H(z)] / S(z). In this system, the output of V(z) is modified based on an estimate of the secondary path S(z) and subtracted from the internal microphone signal EM10 to generate the input to the feedback ANC filter FB10. The output of the feedback ANC filter FB10 is combined with the output of the trans-aural filter HF10 to generate the audio output signal AO10 that is used to drive the loudspeaker. Figure 4 Another block diagram of this system is shown, and Figure 5A A block diagram of an implementation of this system is shown, in which the blocks V(z), and - C(z) are implemented by the trans-aural filter HF10, the path estimate PE10, and the feedback ANC filter FB10, respectively. In Figure 5A the external microphone signal XM10 is filtered by the trans-aural filter HF10. The output of the trans-aural filter RF10 is modified based on the path estimate PE10 and subtracted from the internal microphone signal EM10 to generate the input to the feedback ANC filter FB10. The output of the feedback ANC filter FB10 is combined with the output of the trans-aural filter HF10 to generate the audio output signal AO10 that is used to drive the loudspeaker.

[0050] A user of a personal audio device can wish to listen to a reproduced audio signal (e.g., a far-end voice communication signal (e.g., a telephone call) or a multimedia signal (e.g., a music signal, which can be received via a broadcast or decoded from a stored file or other bitstream)) during ANC operation or even in the acoustically transparent mode. Figure 5B An implementation of a system including such a signal RX10 is shown in the block diagram shown in Figure 5A

[0051] When the estimates of the primary path P(z) and the secondary path S(z) on which V(z) is based are accurate, the system shown in Figure 3C , 4 , 5A, and / or 5B can be effective. However, these paths vary over time, which they are better represented as P(t,z) and S(t,z). For example, even a slight change in the fit of the earbuds can cause the secondary path S(t,z) to change significantly. Thus, a solution that is designed to work optimally in one scenario and acceptably in many scenarios can not provide the desired results in individual cases.

[0052] ​Earplugs are not equally suitable for everyone, and the variation in fit is especially true in the case of earplugs that do not use a silicone earplug sleeve to seal the ear canal (non-blocking earplugs). The result can be that the level of acoustic transparency is inconsistent or insufficient for different users. Even for the same user, the fit can vary over time: for example, when talking or when exercising. In these cases, while the fit can be good at the start, the movement can cause the fit to change over time, resulting in inconsistent performance.

[0053] It can be desirable to adapt the coefficients of the trans-aural filter based on the external and internal microphone signals. For example, the adaptation can be designed such that the internal microphone signal is equal to the external microphone signal, even when the acoustic transfer function changes (e.g. taking into account variations in fit).

[0054] Figure 6 A block diagram of an implementation of the system of Figure 4 is shown, where the trans-aural filter has a fixed part V(z) as described above and an adaptive part. The adaptive part comprises an adaptive filter W(z) whose state is updated based on the reference signal x(n) and the error signal e(n).

[0055] The adaptive part comprises an adaptive block and a pre-filter presenting the signal r(n) to the adaptive block The pre-filter ensures that the input of the adaptive filter is time-aligned, and that the signal r(n) represents the trans-aural component in the absence of W(z) (and assuming

[0056] The adaptive block filters r(n) to produce a result y(n), and updates the state of W(z) based on the difference between the result y(n) and the error signal e(n). In this example, the state of W(z) is updated according to the rule w(n+1) = w(n) - μr(n)[e(n) - y(n)], where μ is a step factor. The updated state of W(z) is then used to update the state of the filter in the processing path of x(n) (i.e. upstream of the fixed filter V(z) or at the output of V(z)).

[0057] Convergence of the adaptive filter W(z) to 1 means, for example, that there is no variation in fit, and that the static trans-aural filter V(z) achieves perfect acoustic transparency. When the secondary path S(t,z) changes such that is not equal to S(t,z), as shown in Figure 6 the solution can become particularly effective, and such a system can provide a more consistent level of acoustic transparency in the case of changes in the acoustic transfer function due to variations in fit.

[0058] Figure 7AAs shown Figure 6 The diagram shows the implementation of the system, which includes, as follows: Figure 5A The features shown and the device A100 according to a specific configuration. Device A100 includes a reference... Figure 5A The described path estimation PE10 and feedback ANC filter FB10 are described. The device A100 also includes a transmission filter HF20, which is... Figure 5A The implementation of the HF10 acoustic transmission filter. Figure 7A In this design, the acoustic transmission filter HF20 has a fixed portion HF24 and an adaptive portion HF22. The fixed portion HF22 includes a fixed filter XF10 (e.g., an implementation of the acoustic transmission filter HF10 as described above). The adaptive portion HF22 includes an update filter UF10, whose state is updated based on external microphone signal XM10 and internal microphone signal EM10 according to the adaptation performed by the adaptive filter AF10.

[0059] The adaptive section HF22 also includes a pre-filter PF10, which presents a signal representing the transmitted component to the adaptive filter AF10 in the absence of the adaptive section (and assumes that the transfer function of the path estimate PE10 is the same as the transfer function of the secondary path S(z)). Figure 7B It shows the relationship with Figure 7A A block diagram of an example of a prefilter PF20 corresponding to a specific implementation of the prefilter PF10. Figure 7B In this process, the pre-filter PF20 uses a cascade of a fixed filter XF10A (which is an instance of a fixed filter XF10) and a path estimate PE10A (which is an instance of a path estimate PE10).

[0060] Return to Figure 7A The adaptive filter AF10 filters the output of the pre-filter PF10 to produce a filtered result, and updates the state of the adaptive filter AF10 based on the difference between the filtered result and the internal microphone signal EM10 (e.g., according to the rule described for the reference filter W(z) above). The updated state of the adaptive filter AF10 is then used to update the state of the update filter UF10. In another implementation, the update filter UF10 is placed at the output of the fixed filter XF10 before branching to path estimation PE10.

[0061] For cases where the acoustic transfer function is time-varying (e.g., when the fit of the earphones changes), the response of the acoustic transmission filter HF20 can also be expected to be time-varying. By including an auxiliary filter (e.g., an update filter UF10) in series with the acoustic transmission response, the cascaded output of filters XF10 and UF10 can track changes in the acoustic transfer function.

[0062] There are no particular requirements on the structure of the update filter UF10. For example, the update filter UF10 can have a finite impulse response (FIR) or an infinite impulse response (IIR). The adaptive filter AF10 can be configured to adapt the coefficients of the update filter UF10 at a lower rate than the adaptive filter AF0 coefficients are updated and / or in the background processing. The adaptive filter AF10 can be configured to update the coefficient values of the update filter UF10 by copying the current state of the adaptive filter AF10 into the update filter UF10.

[0063] The state of the update filter UF10 (e.g., the values of its tap coefficients) can be updated periodically: e.g., according to a time interval (e.g., one second, one half second, one quarter second, or one tenth second) and / or based on an event. For example, the adaptive filter AF10 can be configured to copy the updated coefficient values into the update filter UF10 (for application to the signal path) only after a convergence criterion is reached and / or (in the case of an IIR implementation) a stability criterion is reached.

[0064] Figure 8A A flowchart of a method M100 of audio signal processing is shown, the method M100 comprising tasks T110, T120, and T130. Task T110 produces a pass-through component based on an external microphone signal (e.g., as described above with reference to the pass-through filter HF20). Task T120 produces a feedback component based on an internal microphone signal (e.g., as described above with reference to the feedback ANC filter FB10). Task T130 produces an audio output signal comprising the pass-through component and the feedback component (e.g., by mixing the signals produced by tasks T110 and T120). In this method, a relationship between the external microphone signal and the pass-through component changes in response to a change in a relationship between the audio output signal and the internal microphone signal (e.g., based on a change in an acoustic coupling between a loudspeaker that produces an acoustic signal based on the audio output signal and an internal microphone that is arranged to produce the internal microphone signal in response to the acoustic signal, where the acoustic coupling can change as a result of, e.g., a change in fit).

[0065] A device (e.g., an audible device) can be implemented to include a memory configured to store audio data and a processor configured to receive audio data from the memory and perform the method M100. An apparatus can be implemented to include means for performing each of tasks T110, T120, and T130 (e.g., as software executing on hardware). A computer-readable storage medium can be implemented to include code that, when executed by at least one processor, causes at least one server to perform the method M100.

[0066] Another reason users may experience suboptimal acoustic transparency is that not everyone hears the same sounds. Each individual's hearing profile has unique deficiencies, which can vary from ear to ear. A design that works best in one scenario and acceptable in many may not suit the user's own instinctive hearing profile.

[0067] It is expected that personalized, transparent mode design will be supported. For example, it is expected that acoustic transfer functions and / or system models will be provided tailored to each individual's hearing profile.

[0068] Figure 9 It shows Figure 4 The block diagram illustrates the implementation of a system that includes a compensation filter (also known as a "shaping filter") in the acoustic transmission filter path. The compensation filter has a transfer function A. -1 (z), which is selected to compensate for an individual's unique hearing impairment. The compensation filter can be implemented as follows: Figure 9 The pre-filter shown can be applied to the output of the transmission filter V(z) (in the secondary path estimation). (Before the branch). This system can be used to provide users with imperfect hearing profiles with a perception of acoustic transparency.

[0069] The response of the compensation filter can be based on the user's audiogram, which records a curve describing an individual's hearing impairment profile A(ω). The user's audiogram can include individual results for each ear. Furthermore, the audiogram can indicate how the user perceives sound (at various frequencies) via air conduction and / or bone conduction. Therefore, a complete user audiogram can indicate the user's perception of various frequencies of sound conducted in the air and via bone conduction in the right ear, and the user's perception of various frequencies of sound conducted in the air and via bone conduction in the left ear. Bone conduction testing can be performed using a device placed behind the ear to transmit sound through vibrations of the mastoid bone.

[0070] Figure 10 An example audiogram of a user's left ear is shown. This example shows a bone conduction hearing loss of 30 to 45 dB, with a significant deficit at 2 kHz, and a total hearing loss (including bone and air conduction) of 50 to 80 dB, with a significant deficit at 4 kHz.

[0071] In a specific implementation, the transfer function A of the compensation filter can be obtained by inverting the total hearing loss audiogram curve. -1(z) to compensate the response by providing higher levels in the frequency bands where the user's hearing is degraded. In other implementations, the air conduction audiogram curve can be inverted to obtain the transfer function A(z) of the compensation filter. For example, the air conduction audiogram curve can be determined via testing, or the air conduction audiogram curve can be determined by subtracting the bone conduction audiogram curve from the overall hearing loss audiogram curve. Assuming a suitable audiogram is available, such a system can support an acoustically transparent response in perception even for individuals with imperfect hearing. -1 (z). For example, the air conduction audiogram curve can be determined via testing, or the air conduction audiogram curve can be determined by subtracting the bone conduction audiogram curve from the overall hearing loss audiogram curve. Assuming a suitable audiogram is available, such a system can support an acoustically transparent response in perception even for individuals with imperfect hearing.

[0072] In one example, an application (e.g., executing on a smartphone or tablet computer linked to a personal audio device) is used to obtain the user's audiogram, e.g., via manual data entry or by querying another device. In another example, the application is used to measure the user's audiogram. After obtaining or generating (e.g., measuring) the user's audiogram, data describing the user's audiogram (or an inverted audiogram) can be stored in memory (e.g., memory of the personal audio device or another device) and used to configure the compensation filter. For example, the user's audiogram can be obtained at a first device (e.g., a computer, tablet computer, or smartphone), and data describing the user's audiogram can be uploaded (e.g., via a wired or wireless data link, such as a Bluetooth® data link) to the personal audio device to configure the compensation filter (Bluetooth is a registered trademark of Bluetooth SIG, INC. of Kirkland, Washington, USA). For example, the application can perform a series of tests in which the application causes a sound to be played at a particular intensity and frequency at the left or right ear, while directing the user to tap a designated portion of the touchscreen to indicate which ear (if any) perceived the sound. For example, the application can perform a series of tests in which the application causes a sound to be played at a particular intensity and frequency at the left or right ear, while directing the user to tap a designated portion of the touchscreen to indicate which ear (if any) perceived the sound.

[0073] Figure 11A A block diagram illustrating an implementation of a system as shown in Figure 9 includes features as shown in Figure 5A and an apparatus A200 configured according to a particular configuration. In addition to the features shown in Figure 5A , the apparatus A200 includes a compensation filter CF10 having a transfer function selected to compensate for the individual's unique hearing deficiency (e.g., an inversion of a user's audiogram as described herein). The compensation filter CF10 can be implemented as a pre-filter as shown in Figure 11A , or can be applied to the output of the transmissive filter HF10 (prior to the branch to the path estimate PE10).

[0074] The apparatus A200 can also be configured to receive a reproduced audio signal RX10 (e.g., as shown in Figure 5B . Figure 11BA block diagram is shown of an implementation of the system as shown in

[0075] Figure 12 A block diagram is shown of an implementation of the system as shown in Figure 9 Figure 6 Figure 13 A block diagram is shown of an implementation of the system as shown in

[0076] Figure 8B A block diagram is shown of an implementation of the system as shown in

[0077] A device (e.g., an audible device) can be implemented to include a memory configured to store audio data and a processor configured to receive audio data from the memory and perform the method M200. An apparatus can be implemented to include means for performing each of tasks T210, T220, and T230 (e.g., as software executing on hardware). A computer-readable storage medium can be implemented to include code that, when executed by at least one processor, causes the at least one processor to perform the method M200.

[0078] ​​It can be desirable for a personal audio device to support such individualized hearing compensation for more than one user. For example, the device can be configured to record and store hearing compensation data, such as vent compensation filter states (e.g., filter coefficient values), for each of a set of registered users. In use or during use, the device can select the hearing compensation data (e.g., vent compensation filter states) corresponding to the current user based on, for example, authentication of the user. To illustrate, the user can be authenticated using biometric authentication techniques such as voice authentication, fingerprint recognition, iris recognition, and / or facial recognition. The selection of the hearing compensation data based on user authentication can be incorporated into any of the systems shown in, for example, 11B, 12, or 13. Figure 9 , 11A , 11B, 12, or 13.

[0079] Figure 14 A flowchart showing operations for selecting hearing compensation data (e.g., vent compensation filter states) based on biometric data that identifies or authenticates a user is shown. An identification operation 1402 receives a signal or request that includes biometric data 1404 (such as a sample of the user’s voice) (e.g., based on an external microphone signal XM10, an internal microphone signal EM10, or a combination of both) and identifies the user as user i among a set of n registered users. At operation 1406, an indication of the identification i is used to select corresponding hearing compensation data 1408 from a set of n stored hearing compensation data. In Figure 14 , the stored hearing compensation data includes filter states for each registered user, and the selected hearing compensation data 1408 is copied into a compensation filter (e.g., compensation filter CF10). In some implementations, if the stored hearing compensation data does not include hearing compensation information associated with a particular user, the processor can execute instructions to add hearing compensation data for that particular user to the set of hearing compensation data. For example, the processor can prompt the particular user to provide an audiogram (by selecting a previously generated file or by testing the user’s hearing), and can generate hearing compensation data for the particular user based on the user’s response to the prompt.

[0080] As one example, the biometric authentication can include a voice authentication operation, which can be implemented as a classification of a voice signal of a registered user. In one example, the voice signal is a specified keyword that the user can say to initiate the compensation filter selection operation. Such an operation can be configured to classify the voice signal using, for example, a deep neural network (DNN). In another example, the voice authentication operation is configured to classify the user’s own voice, without regard to what words are spoken.

[0081] One example of a voice authentication operation uses a Gaussian Mixture Model (GMM). A GMM is a statistical method that evaluates the log-likelihood ratio of a particular utterance spoken by a hypothetical speaker. As shown in Figure 15 FIG. 1, the operation can include a front-end processing block that receives the user's voice and produces a feature vector. For each of n registered users, a corresponding GMM indicates the likelihood that the feature vector represents the voice of the corresponding user, and the voice is classified according to the GMM that indicates the highest likelihood.

[0082] The voice authentication operation can be configured to use a Deep Neural Network (DNN) to enable individualized hearing deficit compensation filters. The DNN (e.g., a fully connected neural network) can be trained to model each of N registered speakers, and the output layer of the DNN can be a 1 x N one-hot vector that indicates which of the N speakers is predicted. In one example, the DNN is trained on an array of feature vectors, where each array is computed from the speech of one of the registered speakers by forming the speech into a series of frames and computing a K-length vector of Mel-Frequency Cepstral Coefficients (MFCCs) for each frame. The voice authentication operation is then performed by computing K-length MFCC vectors in real-time from the speech signal to be classified and using these vectors as input to the trained DNN.

[0083] In another example, a Long Short-Term Memory (LSTM) network is used to perform a text-independent voice authentication operation. LSTM networks are relatively insensitive to lags of unknown duration that can occur between significant events in a time series. LSTM networks are well suited for classifying time series data, and can be particularly effective for short utterances. For example, such an operation can be configured to use MFCCs to directly capture temporal speaker information that is classified according to a set of registered users using an LSTM network.

[0084] Additionally or alternatively, the device can select hearing compensation data (e.g., see-through compensation filter states) corresponding to the current user based on recognition of the user's face. For example, the recognition operation can be performed by a device with a camera (e.g., a smartphone, tablet computer, laptop or other personal computer, smart glasses, etc.) and wirelessly linked to send an indication of the recognized user i (e.g., by a data link) to the personal audio device. In another example, the recognition operation is performed by a head-mounted device ("HMD", such as smart glasses) that includes a camera arranged to capture images of the user's face, and that also includes or is linked to the personal audio device.

[0085] Figure 16A An example is shown in FIG. 2 for selecting hearing compensation data (in this case, see-through compensation filter states) based on recognition of the user's face. In this example, the personal audio device 202 is configured to receive a signal from a camera 204 that is arranged to capture images of the user's face. The camera 204 can be a component of the personal audio device 202, or it can be a separate device that is wirelessly linked to the personal audio device 202. The camera 204 can be a component of a head-mounted device ("HMD") such as smart glasses, or it can be a component of a separate device such as a smartphone, tablet computer, laptop or other personal computer, etc. Figure 16Aa flowchart of the operation of a face recognition operation (e.g., the operation of the acoustic compensation filter state in the example of FIG. 1). The face recognition operation receives an image signal (e.g., from a camera as described above) that includes a user's face, and identifies the face as that of user i from a set of n registered users. An indication of the identity i is used to select a corresponding filter state from a stored set of n filter states, and the selected filter state is copied into the compensation filter (e.g., compensation filter CF10).

[0086] The face recognition operation can be performed using any of a variety of methods. In one example, the face recognition operation uses principal component analysis to map a face image from a high-dimensional space to a lower-dimensional space to facilitate comparison to a set of known images. Such a method can use, for example, the EigenFace algorithm.

[0087] The face recognition operation can be a DNN-based method that uses convolution and pooling layers to reduce the dimensionality of the problem. Such an operation can be configured to perform feature extraction via deep learning, and then classify the extracted features. Examples of algorithms that can be used include FaceNet and DeepFace.

[0088] The face recognition operation can be implemented as a classification of a user's face among a set of registered users. Figure 16B An example of such an operation is shown, in which a trained DNN is used to perform the classification. An image signal is pre-processed to extract a face. The extracted face can be used as a feature vector to be classified, or an operation can be performed to generate a feature vector. The feature vector is input to the trained DNN, which classifies the vector to indicate the corresponding one of a set of n registered users.

[0089] In one example of a DNN-based face recognition operation, a face detector is used to locate the face, which is then aligned to normalized canonical coordinates in the image space. The normalized image is input to a face recognition module, which uses a trained DNN to extract a feature vector from the image. The extracted feature vector is then classified (e.g., using a support vector machine) to identify one of a set of registered users.

[0090] In a particular use case, it can be desirable that the personal audio device automatically transitions to the acoustically transparent mode when the user is driving. The vehicle (e.g., a car) can comprise a camera arranged to capture images of the driver and a processor configured to perform a face recognition operation on the captured images and send an indication of the identity of the user i to the personal audio device (e.g., without any input of the user) to select the corresponding individualized hearing compensation data. The personal audio device can also be configured to automatically enter the acoustically transparent mode upon receiving the indication of the identity of the user i and / or another signal from the processor of the vehicle. In another example, the processor of the vehicle stores the acoustically transparent compensation filter state corresponding to the current user and uploads it to the personal audio device upon completion of the face recognition operation.

[0091] In another example, the personal audio device is installed in or linked to a head mounted device (HMD; e.g., smart glasses) that comprises a camera arranged to capture images of the user’s eyes (e.g., for gaze detection). In this case, the HMD is configured to perform an iris recognition operation to produce an indication of the identity of the user i that is received by the personal audio device and used to select the corresponding individualized acoustically transparent compensation filter state.

[0092] The personal audio device as described herein can also comprise an ANC system configured to perform an ANC operation (e.g., for times when noise cancellation is needed instead of acoustical transparency). Figure 17 An example of an ANC system is shown that comprises a feed-forward ANC filter whose transfer function C(z) is adapted according to a normalized filter-X LMS (nFxLMS) algorithm. Figure 18 An example of an ANC system is shown that comprises an ANC filter whose transfer function C(z) is fixed (e.g., implemented as a long-tap finite impulse response (FIR) or infinite impulse response (IIR) filter) and which comprises a gain k that is adapted according to a normalized filter-X LMS (nFxLMS) algorithm. In one example, the gain k is adapted according to the expression where μ denotes a step factor and γ denotes a leakage factor. As shown in Figure 18 and 19 It can be desirable to include a band-pass filter on the external microphone signal and / or the internal microphone signal (e.g., to focus the adaptation on low-frequency noise reduction).

[0093] It can be desirable to implement the ANC system to include a filter on the feedback path, which can be fixed or adaptive. Such a feedback filter can be provided in addition to or instead of the filter on the feed-forward path. Figure 19An example of an ANC system is shown in Figure 18 which also includes a fixed filter - H(z) on the feedback path.

[0094] As shown in Figure 18 and 19 it can be desirable to bandpass filter the signal input to the adaptive algorithm (e.g., to emphasize cancellation at low audio frequencies). It can also be desirable to implement the system as shown in Figure 18 or Figure 19 switching between different fixed C(z) and / or different H(z) at different times (e.g., to optimize cancellation of a particular audio frequency range at that time as desired).

[0095] It can be desirable to configure the ANC filter to high pass filter the signal (e.g., to attenuate high amplitude, low frequency acoustic signals). Additionally or alternatively, it can be desirable to configure the ANC filter to low pass filter the signal (e.g., such that the ANC filter reduces acoustic signals having frequencies at high frequencies). Because the anti-noise signal should be available as the acoustic noise travels from the microphone to the actuator (i.e., the loudspeaker), the processing delay caused by the ANC filter should not exceed a very short time (typically on the order of 30 to 60 microseconds). In the example shown in Figure 17 the ANC filter is executed in a first clock domain (e.g., in hardware at a clock rate of, e.g., 8 MHz) and the adaptation is executed in a second clock domain at a lower frequency (e.g., in software on a digital signal processor (DSP) clocked at a rate of, e.g., 16 kHz). Figure 18 and Figure 19 The examples shown in Figure 19 may be similarly implemented, and in the example shown in the feedback filter can also be executed in a higher rate clock domain.

[0096] As shown in Figure 1B the audible devices D10L, D10R worn at each ear of the user can be configured to communicate audio and / or control signals with each other wirelessly (e.g., via a data link or through near field magnetic induction (NFMI). In some cases, the audible devices can also be equipped with internal microphones located within the ear canal. Such microphones can be used, for example, to obtain an error signal (e.g., a feedback signal) for active noise cancellation (ANC). The audible devices can be configured to communicate wirelessly with a wearable device or "wearable" that can, for example, send volume levels or other control commands. Examples of wearable devices include (in addition to audible devices) watches, head-mounted displays, earphones, fitness trackers, and pendants.

[0097] Audible devices worn on each ear of a user can be configured to communicate audio and / or control signals wirelessly with one another. For example, a true wireless stereo (TWS) protocol allows for a stereo Bluetooth stream to be provided to a master device (e.g., one of a pair of audible devices), which reproduces one channel and sends the other channel to a slave device (e.g., the other of a pair of audible devices). Even when a pair of audible devices are linked in this way, many audio processing operations can occur independently on each device in the TWS group, such as ANC operations.

[0098] The case where each device modifies its ANC operations independently of the device at the other ear of the user can result in an unbalanced hearing experience. For wireless audible devices, a mechanism by which the two audible devices negotiate their states and share ANC-related information can help provide a more balanced ANC experience to the user. A device, method, and / or apparatus as described herein (e.g., one of a pair of audible devices) can be further configured to exchange parameter values or other indications with the other device (the other of a pair of audible devices) to provide a uniform user experience. In one example, it can be desirable for a device to attenuate or disable an ANC path in response to an indication of a howl from the other device. In another example, it can be desirable for the pair of audible devices to perform a synchronized entry into a transparent mode (e.g., from an active (ambient) noise cancellation mode).

[0099] The human ear is generally not sensitive to phase. However, a phase difference between sounds perceived at the left and right ears of a user can be important for spatial localizability. Thus, it can be desirable for the phase response of the vented paths at the left and right ears of a user to be similar (e.g., to maintain such a phase difference). In another example, parameter values generated during adaptation of a vented filter HF20 (e.g., updated coefficient values) are shared between personal audio devices (e.g., earbuds) worn at the left and right ears of a user. Such shared parameters can be used to ensure that the adaptation operations at the left and right ears produce vented filter paths with similar phase responses.

[0100] Figure 20A flowchart illustrating a method M300 for audio signal processing based on hearing compensation data for a particular user is shown. The method M300 includes tasks T310, T320, T330, and T340. Task T310 receives an external microphone signal (e.g., the external microphone signal XM10 described above) from a first microphone and an internal microphone signal (e.g., the internal microphone signal EM10 described above) from a second microphone. Task T320 generates a transmissive component based on the external microphone signal and the hearing compensation data, where the hearing compensation data is based on a hearing map for the particular user (e.g., as described above with reference to the compensation filter CF10 and the transmissive filter HF20). Task T330 generates a feedback component based on the internal microphone signal (e.g., as described above with reference to the feedback ANC filter FB10). Task T340 causes a loudspeaker to generate an audio output signal based on the transmissive component and the feedback component (e.g., by mixing the signals generated by tasks T320 and T330 and driving the loudspeaker based on the result of the mixing). In this method, a relationship between the external microphone signal and the transmissive component changes in response to a change in a relationship between the audio output signal and the internal microphone signal (e.g., a change in an acoustic coupling between the loudspeaker that generates an acoustic signal based on the audio output signal and the internal microphone that is arranged to generate the internal microphone signal in response to the acoustic signal, where the acoustic coupling can change as a result of, for example, a fit change). Moreover, in this method, the hearing compensation data based on the user-specific hearing map is used to improve the perceived sound quality of audio provided to the user based on the user's own hearing deficiencies.

[0101] A device (e.g., an audible device) can be implemented to include a memory configured to store audio data and a processor configured to receive the audio data from the memory and perform the method M300. An apparatus can be implemented to include means for performing each of the tasks T310, T320, T330, and T340 (e.g., as software executing on hardware). A computer-readable storage medium can be implemented to include code that, when executed by at least one processor, causes the at least one processor to perform the method M300.

[0102] Reference is made to Figure 21 A block diagram of a particular illustrative implementation of a device is depicted and generally designated 2100. In the illustrative implementation, the device 2100 includes a signal processing circuit 2140, which can correspond to or include any of the filters, signal paths, or other audio signal processing components described above with reference to any of Figures 1A-20 In the illustrative implementation, the device 2100 can perform one or more of the operations described above with reference to Figures 1A-20

[0103] In Figure 21 ​In the illustrated example, the device 2100 is configured to communicate with a second device 2190. For example, the second device 2120 can store a plurality of hearing compensation data sets 2192. In this example, the device 2100 can retrieve particular hearing compensation data from the second device 2190 for use by the signal processing circuit 2140. To illustrate, the device 2100 can authenticate a user based on biometric data and send information identifying the authenticated user to the second device 2190. In this illustrative example, the second device 2190 selects particular hearing compensation data corresponding to the user from among the hearing compensation data sets 2192 and sends the particular hearing compensation data to the device 2100 for use.

[0104] Alternatively, the second device 2190 can authenticate the user. To illustrate, the second device 2190 can include one or more sensors (e.g., a fingerprint scanner, a camera, a microphone, etc.) to collect biometric data for authenticating the user. As another illustrative example, the device 2100 can collect biometric data and send the biometric data to the second device 2190. In this illustrative example, the second device 2190 authenticates the user based on the biometric data received from the device 2100.

[0105] In a particular implementation, the device 2100 includes a processor 2106 (e.g., a central processing unit (CPU)). The device 2100 can include one or more additional processors 2110 (e.g., one or more DSPs). The processor 2110 can include a voice and music coder-decoder (CODEC) 2108 including a voice coding ("vocoder") encoder 2136, a vocoder decoder 2138, a signal processing circuit 2140, or a combination thereof.

[0106] The device 2100 can include a memory 2186 and a CODEC 2134. The memory 2186 can include instructions 2156 executable by the one or more additional processors 2110 (or the processor 2106) to implement the references Figures 1A-20One or more of the functions described above can be performed by device 2100. Device 2100 can include modem 2154 coupled to antenna 2152 via transceiver 2150. Modem 2154, transceiver 2150, and antenna 2152 can facilitate data exchange with another device, such as second device 2190. For example, second device 2150 can store a plurality of sets of hearing compensation data corresponding to a plurality of users. In this example, device 2100 can send (via modem 2154, transceiver 2150, and antenna 2152) a request including user identification information, such as a user identification for a particular user or biometric data associated with a particular user. In this example, second device 2190 can select particular hearing compensation data associated with the particular user from sets of hearing compensation data 2192, such as a transmissive compensation filter state determined based on an audiogram for the particular user, as described above with reference to, for example Figures 9-16B In some implementations, if sets of hearing compensation data 2192 do not include any hearing compensation data associated with the particular user, processor 2106 or processor 2110 can execute instructions 2156 to add hearing compensation data for the particular user to sets of hearing compensation data 2192. For example, processor 2106 or processor 2110 can prompt the particular user to provide an audiogram (by selecting a previously generated file or by testing the user’s hearing), and can generate hearing compensation data for the particular user based on the user’s response to the prompt. In this example, device 2100 can send the hearing compensation data to second device 2190 for addition to sets of hearing compensation data 2192.

[0107] Device 2100 can include display 2128 coupled to display controller 2126. One or more speakers 2146 and one or more microphones 2142 can be coupled to CODEC 2134. CODEC 2134 can include digital-to-analog converter (DAC) 2102 and analog-to-digital converter (ADC) 2104. In a particular implementation, CODEC 2134 can receive an analog signal from microphone 2142, convert the analog signal to a digital signal using analog-to-digital converter 2104, and send the digital signal to speech and music codec 2108. In a particular implementation, speech and music codec 2108 can provide a digital signal to CODEC 2134. CODEC 2134 can convert the digital signal to an analog signal using digital-to-analog converter 2102, and can provide the analog signal to speaker 2146.

[0108] In particular implementations, device 2100 can be included in a system-in-a-package or system-on-a-chip device 2122. In particular implementations, memory 2186, processor 2106, processor 2110, display controller 2126, CODEC 2134, modem 2154, and transceiver 2150 are included in system-in-a-package or system-on-a-chip device 2122. In particular implementations, input device 2130 and power supply 2144 are coupled to system-in-a-package or system-on-a-chip device 2122. Also, in particular implementations, as illustrated in FIG. 21 A, display 2128, input device 2130, speaker 2146, microphone 2142, antenna 2152, and power supply 2144 are external to system-in-a-package or system-on-a-chip device 2122. In particular implementations, each of display 2128, input device 2130, speaker 2146, microphone 2142, antenna 2152, and power supply 2144 can be coupled to a component of system-in-a-package or system-on-a-chip device 2122, such as an interface or a controller. Figure 21

[0109] Device 2100 can include an audible device, a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular telephone, a laptop computer, a computer, a tablet computer, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a virtual reality headset, a drone, a home automation system, a voice-activated device, a wireless speaker with voice-activated device, a portable electronic device, an automobile, a vehicle, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0110] In various implementations, device 2100 can have more or fewer components than Figure 21 those shown. For example, when device 2100 corresponds to an audible device, in some implementations device 2100 can omit display 2128 and display controller 2126. In some implementations, device 2100 corresponds to a smartphone or another portable electronic device that provides audio data to an audible device (not shown in FIG. 21 A). In such implementations, signal processing circuitry 2140 can be included in the audible device, instead of (or in addition to) being included in device 2100. Figure 21 Figure 22 and 23 ​​A diagram illustrates an example of a hearable device that includes an instance of signal processing circuit 2140. In such an implementation, second device 2190 can include a server or other computing device that stores a set of hearing compensation data 2192 and provides particular hearing compensation data to device 2100 based on a request from device 2100.

[0111] Figure 22 A diagram illustrates a schematic of a headset device 2200 configured to perform audio signal processing based on hearing compensation data for a particular user. In Figure 22 In particular examples, headset device 2200 includes one or more hearable devices, such as hearable devices D10L and D10R, each of which can include or be coupled to an instance of signal processing circuit 2140. To illustrate, hearable device D10L can include or be coupled to signal processing circuit 2140A, and hearable device D10R can include or be coupled to signal processing circuit 2140B.

[0112] Figure 23 A diagram illustrates a schematic of an extended reality (e.g., virtual reality, mixed reality, or augmented reality) headset 2300 configured to perform audio signal processing based on hearing compensation data for a particular user. In Figure 23 In particular examples, headset 2300 includes a visual interface device 2302 mounted to be in front of a user’s eyes to enable display of augmented reality or virtual reality images or scenes to the user while wearing headset 2300. Headset 2300 also includes one or more microphones 2304, 2306 to capture ambient environmental sounds (e.g., external microphone signal XM10 described above), to capture error signals (e.g., internal microphone signal EM10 described above), and the like. Headset 2300 also includes Figure 20 one or more instances of signal processing circuit 2140, such as signal processing circuits 2140A and 2140B. In particular examples, a user of headset 2300 can participate in a conversation with a remote participant, such as via a video conference using microphones 2304, 2306, audio speakers, and visual interface device 2302.

[0113] Any of the systems described herein can be implemented as (or as part of) an apparatus, device, assembly, integrated circuit (e.g., a chip), chip set, or printed circuit board. In one example, such a system is implemented within a cellular telephone (e.g., a smartphone). In another example, such a system is implemented within a hearable device or other wearable device.

[0114] Unless the context explicitly limits it, the term "signal" is used herein to indicate any general meaning, including the state of a storage location (or set of storage locations) as expressed on a wire, bus, or other transmission medium. Unless the context explicitly limits it, the term "generate" is used herein to indicate any general meaning, such as calculating or otherwise producing. Unless the context explicitly limits it, the term "compute" is used herein to indicate any general meaning, such as calculating, evaluating, estimating, and / or selecting from a plurality of values. Unless the context explicitly limits it, the term "obtain" is used herein to indicate any general meaning, such as calculating, deriving, receiving (e.g., from an external device), and / or retrieving (e.g., from an array of memory elements). Unless the context explicitly limits it, the term "select" is used herein to indicate any general meaning, such as identifying, indicating, applying, and / or using at least one of two or more sets, but not all of them. Unless the context explicitly limits it, the term "determine" is used herein to indicate any general meaning, such as deciding, establishing, summarizing, calculating, selecting, and / or evaluating. Where the term "comprising" is used in this specification and claims, it does not exclude other elements or operations. The term “based on” (as in “A is based on B”) is used to indicate any of its general meanings, including (i) “derived from” (e.g., “B is a cause of A”); (ii) “based on at least” (e.g., “A is based on at least B”); and (iii) “equal to” (e.g., “A equals B”) if applicable in a particular context. Similarly, the term “in response to” is used to indicate any of its general meanings, including “in response to at least”. Unless otherwise stated, the terms “at least one of A, B, and C”, “one or more of A, B, and C”, “at least one of A, B, and C”, and “one or more of A, B, and C” indicate “A and / or B and / or C”. Unless otherwise stated, the terms “each of A, B, and C” and “each of A, B, and C” indicate “A and B and C”.

[0115] Unless otherwise indicated, any disclosure of operations of a device having specific features is also expressly intended to disclose a method having analogous features (and vice versa), and any disclosure of operations of a device according to a particular configuration is also expressly intended to disclose a method according to the analogous configuration (and vice versa). The term "configuration" can be used to refer to methods, devices, and / or systems as indicated by their particular context. The terms "method," "process," "procedure," and "technique" can be used interchangeably and generically unless otherwise indicated by the particular context. A "task" having multiple sub-tasks is also a method. The terms "device" and "apparatus" can also be used generically and interchangeably unless otherwise indicated by the particular context. The terms "element" and "module" are generally used to indicate a part of a larger configuration. The term "system" is used herein to mean any of its ordinary meanings, including "a group of elements interacting to serve a common purpose," unless otherwise explicitly limited by context.

[0116] Unless originally introduced by a definite article, ordinal terms used to modify a claim element (e.g., "first," "second," "third," etc.) do not themselves connote priority or order of one claim element over another but merely distinguish one claim element from another claim element having a same name (but used to introduce the ordinal term). Each of the terms "a plurality" and "a set" are used herein to mean an integer quantity greater than one, unless the context clearly limits.

[0117] The terms "encoder," "codec," and "coding system" are used interchangeably to mean a system that includes at least one encoder and a corresponding decoder, the encoder configured to receive and encode frames of an audio signal (possibly after one or more pre-processing operations such as perceptual weighting and / or other filtering operations), the corresponding decoder configured to produce a decoded representation of the frames. Such encoders and decoders are typically deployed at opposite ends of a communication link. The term "signal component" is used to mean a constituent part of a signal that can include other signal components. The term "audio content from a signal" is used to mean an expression of audio information carried by the signal.

[0118] Various elements of implementations of the apparatuses or systems as disclosed herein can be implemented as any combination of hardware and software and / or firmware, which combinations are deemed suitable for the intended application. For example, such elements can be fabricated as one or more electronic and / or optical devices residing, for example, on the same chip or among two or more chips of a chip set. One example of such a device is a fixed or programmable array of logic elements, such as transistors or logic gates, and any of these elements can be implemented as one or more such arrays. Any two or more, or even all, of these elements can be implemented together in the same array or arrays. Such an array or arrays can be implemented within one or more chips (for example, within a chip set including two or more chips).

[0119] Processors or other components for processing as disclosed herein can be fabricated as one or more electronic and / or optical devices residing, for example, on the same chip or among two or more chips of a chip set. One example of such a device is a fixed or programmable array of logic elements, such as transistors or logic gates, and any of these elements can be implemented as one or more such arrays. Such an array or arrays can be implemented within one or more chips (for example, within a chip set including two or more chips). Examples of such arrays include fixed or programmable arrays of logic elements, such as microprocessors, embedded processors, IP cores, DSPs (digital signal processors), FPGAs (field-programmable gate arrays), ASSPs (application-specific standard products), and ASICs (application-specific integrated circuits). Processors or other components for processing as disclosed herein can also be implemented as one or more computers (for example, machines including one or more arrays programmed to execute one or more sets or sequences of instructions) or other processors. The processors described herein can be used to perform tasks or execute other instruction sets not directly related to the implementation of the method M100, M200, or M300 (or another method disclosed with reference to the operation of the apparatuses or systems described herein), such as tasks related to another operation of a device or system in which the processor is embedded (for example, a voice communication device such as a smartphone or smart speaker). A portion of the method disclosed herein can also be performed under the control of one or more other processors.

[0120] Particular aspects of the present disclosure are described in the following first set of related clauses:

[0121] According to Clause 1, an apparatus for audio signal processing comprises a memory configured to store instructions and a processor configured to execute the instructions to: receive an external microphone signal from a first microphone; generate a transmissive component based on the external microphone signal and hearing compensation data, wherein the hearing compensation data is based on a hearing map of a particular user; and cause a loudspeaker to generate an audio output signal based on the transmissive component.

[0122] Clause 2 includes the apparatus of Clause 1, wherein the hearing map represents a hearing deficiency profile of the particular user.

[0123] Clause 3 includes the apparatus of Clause 1 or Clause 2, wherein the processor is configured to execute the instructions to generate the hearing compensation data based on an inversion of the hearing map.

[0124] Clause 4 includes the apparatus of any of Clauses 1-3, wherein the processor is configured to execute the instructions to receive the hearing compensation data from a second apparatus.

[0125] Clause 5 includes the apparatus of Clause 4, wherein the hearing compensation data is accessed based on authentication of the particular user.

[0126] Clause 6 includes the apparatus of Clause 5, wherein the particular user is authenticated based on voice recognition.

[0127] Clause 7 includes the apparatus of Clause 5 or Clause 6, wherein the particular user is authenticated based on facial recognition.

[0128] Clause 8 includes the apparatus of any of Clauses 5-7, wherein the particular user is authenticated based on iris recognition.

[0129] Clause 9 includes the apparatus of any of Clauses 5-8, wherein the memory is configured to store a set of hearing compensation data corresponding to a plurality of users, and wherein the request to retrieve the hearing compensation data is sent to the second apparatus based on a determination that the set of hearing compensation data does not include any hearing compensation data associated with the particular user.

[0130] Clause 10 includes the apparatus of any of Clauses 5-9, wherein the second apparatus performs a user authentication operation and provides the hearing compensation data to the apparatus in response to authentication of the particular user.

[0131] Clause 11 includes the apparatus of Clause 10, wherein the processor is further configured to execute the instructions to add the hearing compensation data to the set of hearing compensation data.

[0132] Clause 12 includes the device of any of clauses 1-11, wherein the processor is further configured to execute instructions to update the hearing compensation data based on a hearing test of the particular user.

[0133] Clause 13 includes the device of any of clauses 1-12, wherein the relationship between the external microphone signal and the sound-transmission component varies in response to a change in placement of the earphone within the ear canal.

[0134] Clause 14 includes the device of any of clauses 1-13, wherein the memory, the processor, the first microphone, and the speaker are integrated in at least one of a headset, a personal audio device, or an earphone.

[0135] Clause 15 includes the device of any of clauses 1-14, wherein the relationship between the external microphone signal and the sound-transmission component varies in response to a change in a relationship between the audio output signal and the internal microphone signal.

[0136] Clause 16 includes the device of any of clauses 1-15, wherein the processor is further configured to execute instructions to receive a reproduced audio signal, wherein the audio output signal is based on the reproduced audio signal.

[0137] Clause 17 includes the device of any of clauses 1-16, wherein the processor is further configured to execute instructions to dynamically adjust the sound-transmission component to reduce a blocking effect.

[0138] Clause 18 includes the device of any of clauses 1-17, wherein the processor is further configured to: receive the internal microphone signal from a second microphone; and generate a feedback component based on the internal microphone signal, wherein the audio output signal is further based on the feedback component, wherein the feedback component is to reduce components of the internal microphone signal other than the sound-transmission component.

[0139] According to clause 19, a method of audio signal processing includes: receiving an external microphone signal from a first microphone; generating a sound-transmission component based on the external microphone signal and hearing compensation data, wherein the hearing compensation data is based on a hearing profile of a particular user; and causing a speaker to generate an audio output signal based on the sound-transmission component.

[0140] Clause 20 includes the method of clause 19, further comprising receiving a reproduced audio signal, wherein the audio output signal comprises the reproduced audio signal, and wherein a relationship between the external microphone signal and the sound-transmission component varies when the reproduced audio signal is non-active.

[0141] Clause 21 includes the method of clause 19 or clause 20, wherein the relationship between the external microphone signal and the sound-transmission component varies in response to a change in placement of the device within the ear canal.

[0142] Clause 22 includes the method of any of clauses 19-21, wherein the hearing compensation data is selected from a set of hearing compensation data corresponding to a plurality of users based on a signal identifying a particular user.

[0143] Clause 23 includes the method of clause 22, wherein the signal identifying the particular user is generated based on a voice authentication operation.

[0144] Clause 24 includes the method of clause 22 or clause 23, wherein the signal identifying the particular user is generated based on a facial recognition operation.

[0145] Clause 25 includes the method of any of clauses 22-24, wherein the signal identifying the particular user is generated based on a biometric identification operation.

[0146] Clause 26 includes the method of any of clauses 20-25, further comprising receiving an internal microphone signal from a second microphone; and generating a feedback component that is out of phase with the internal microphone signal, wherein the audio output signal is further based on the feedback component.

[0147] According to clause 27, an apparatus for audio signal processing comprises means for receiving an external microphone signal from a first microphone; means for generating a transmissive component based on the external microphone signal and hearing compensation data, wherein the hearing compensation data is based on a hearing map of a particular user; and means for causing a speaker to generate an audio output signal based on the transmissive component.

[0148] Clause 28 includes the apparatus of clause 27, further comprising means for selecting the hearing compensation data from a set of hearing compensation data based on a signal, wherein the set of hearing compensation data corresponds to a plurality of users, and wherein the signal identifies a particular user.

[0149] Clause 29 includes the apparatus of clause 28, wherein the signal identifying the particular user is generated by a biometric authentication operation.

[0150] Clause 30 includes the apparatus of any of clauses 27-29, wherein a relationship between the external microphone signal and the transmissive component varies in response to a change in placement of the device within an ear canal of the particular user.

[0151] Clause 31 includes the apparatus of any of clauses 27-30, further comprising means for receiving an internal microphone signal from a second microphone; and means for generating a feedback component that is out of phase with the internal microphone signal, wherein the audio output signal is further based on the feedback component.

[0152] According to Clause 32, a non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause the at least one processor to: receive an external microphone signal from a first microphone; generate a transmissive component based on the external microphone signal and hearing compensation data, wherein the hearing compensation data is based on a hearing profile of a particular user; and cause a speaker to generate an audio output signal based on the transmissive component.

[0153] Clause 33 includes the non-transitory computer-readable storage medium of Clause 32, wherein the hearing compensation data is selected from a set of hearing compensation data based on a signal, wherein the set of hearing compensation data corresponds to a plurality of users, and wherein the signal identifies the particular user based on biometric authentication.

[0154] Clause 34 includes the non-transitory computer-readable storage medium of Clause 28 or Clause 33, wherein a relationship between the external microphone signal and the transmissive component varies in response to a change in placement of the device within an ear canal.

[0155] Clause 35 includes the non-transitory computer-readable storage medium of Clause 28 or Clause 34, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to: receive an internal microphone signal from a second microphone, and generate a feedback component that is out of phase with the internal microphone signal, wherein the audio output signal is further based on the feedback component.

[0156] Each of the tasks of the methods disclosed herein can be directly implemented in hardware, a software module executed by a processor, or a combination of the two. In a typical application of implementation of the methods as disclosed herein, an array of logic elements (e.g., logic gates) is configured to perform one, more or even all of the various tasks of the method. One or more (possibly all) of these tasks can also be implemented as code (e.g., one or more sets of instructions), embodied in a computer program product (e.g., one or more data storage media such as floppy disks, CD-ROMs, or other nonvolatile storage cartridges, semiconductor memory chips, etc.), which is readable and / or executable by a machine (e.g., a computer) including an array of logic elements (e.g., a processor, microprocessor, microcontroller, or other finite state machine). The tasks of implementations of the methods disclosed herein can also be performed by more than one such array or machine. In these or other implementations, the tasks can be performed by a device for wireless communication such as a cell phone or other device having such communication capability. Such a device can be configured to communicate with circuit- switched and / or packet-switched networks (e.g., using one or more protocols such as VoIP). For example, such a device can include RF circuitry configured to receive and / or transmit encoded frames.

[0157] In one or more exemplary embodiments, the operations described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, such operations can be stored or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer-readable storage media and communication media, except that computer-readable storage media does not include transitory media. By way of example, and not limitation, computer-readable storage media can include an array of storage elements, such as semiconductor memory (which can include without limitation dynamic or static RAM, ROM, EEPROM, and / or flash TM RAM), or ferroelectric, resistive, ovonic, polymeric, or phase-change memory; CD-ROM or other optical disk storage; and / or magnetic disk storage or other magnetic storage devices. Such storage media can store data which is accessible by a computer, the instruction or the data representing an aspect of the present disclosure. Communication media can include any medium that facilitates the transfer of a computer program from one place to another, such as over a network. Further, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and / or microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and / or microwave are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray Disc (Blu-ray Disc Association, Universal City, Calif), where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. TM

[0158] The previous description is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these implementations will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other implementations without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the implementations shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein and made apparent to others skilled in the art by the teachings herein.

Claims

1. A device for audio signal processing, the device comprising: a memory configured to store instructions; and a processor configured to execute the instructions to: receive an external microphone signal from a first microphone; receive an internal microphone signal from a second microphone; update filter coefficients of an adaptive filter based on a difference between the internal microphone signal and an output of the adaptive filter based on the external microphone signal; update filter coefficients of an update filter at a lower rate than the filter coefficients of the adaptive filter are updated based on the updated filter coefficients of the adaptive filter; output an output signal from a trans-ear filter coupled to the first microphone and configured to facilitate generating acoustic transparency in an audio output signal produced by a loudspeaker, wherein an input to the trans-ear filter is based on the external microphone signal, and wherein the trans-ear filter comprises a fixed portion and an adaptive portion, wherein the adaptive portion comprises the update filter and the adaptive filter; and cause the loudspeaker to produce the audio output signal based on the output signal of the trans-ear filter.

2. The apparatus of claim 1, wherein, The input to the trans-ear filter is further based on hearing compensation data, wherein the hearing compensation data is based on a hearing map of a particular user.

3. The apparatus of claim 1, wherein, Hearing compensation data is applied to the output signal of the trans-ear filter, wherein the hearing compensation data is based on a hearing map of a particular user.

4. The apparatus of claim 2 or 3, wherein, The hearing map represents a hearing deficiency profile of the particular user.

5. The apparatus of claim 2 or 3, wherein, The processor is configured to execute the instructions to generate the hearing compensation data based on an inverse of the hearing map.

6. The apparatus of claim 2 or 3, wherein, The processor is configured to execute the instructions to receive the hearing compensation data from a second device.

7. The apparatus of claim 6, wherein, The hearing compensation data is accessed based on authentication of the particular user.

8. The apparatus of claim 7, wherein, The particular user is authenticated based on voice recognition.

9. The apparatus of claim 7, wherein, The particular user is authenticated based on facial recognition.

10. The apparatus of claim 7, wherein, The particular user is authenticated based on iris recognition.

11. The apparatus of claim 7, wherein, The memory is configured to store a set of hearing compensation data corresponding to a plurality of users, and wherein, based on a determination that the set of hearing compensation data does not include any hearing compensation data associated with the particular user, a request to retrieve the hearing compensation data is sent to a second device.

12. The apparatus of claim 11, wherein, The processor is further configured to execute the instructions to add the hearing compensation data to the set of hearing compensation data.

13. The apparatus of claim 2 or 3, wherein, The processor is further configured to execute the instructions to update the hearing compensation data based on a hearing test of the particular user.

14. The apparatus of claim 1, wherein, A relationship between the external microphone signal and the trans-ear filter changes in response to a change in placement of an earphone within an ear canal.

15. The apparatus of claim 1, wherein, The memory, the processor, the first microphone, and the loudspeaker are integrated in at least one of a headset, a personal audio device, or an earphone.

16. The apparatus of claim 1, wherein, The processor is further configured to: produce a feedback component based on the internal microphone signal, wherein the audio output signal is further based on a feedback component, wherein a relationship between the external microphone signal and the acoustic transparency filter varies in response to a change in a relationship between the audio output signal and the internal microphone signal, and wherein the feedback component is to reduce components in the internal microphone signal other than the output signal of the acoustic transparency filter.

17. The apparatus of claim 1, wherein, The processor is further configured to execute the instructions to receive a reproduced audio signal, wherein the audio output signal is based on the reproduced audio signal.

18. The apparatus of claim 1, wherein, The processor is further configured to execute the instructions to dynamically adjust the acoustic transparency filter to reduce blocking effects.

19. A method of audio signal processing, the method comprising: receiving an external microphone signal from a first microphone; receiving an internal microphone signal from a second microphone; updating filter coefficients of an adaptive filter based on a difference between the internal microphone signal and an output of the adaptive filter based on the external microphone signal; updating filter coefficients of an update filter at a lower rate than at which the filter coefficients of the adaptive filter are updated based on the updated filter coefficients of the adaptive filter; outputting an output signal from an acoustic transparency filter, the acoustic transparency filter being coupled to the first microphone and configured to facilitate generation of acoustic transparency in an audio output signal produced by a loudspeaker, wherein an input of the acoustic transparency filter is based on the external microphone signal, and wherein the acoustic transparency filter comprises a fixed portion and an adaptive portion, wherein the adaptive portion comprises the update filter and the adaptive filter; and causing the loudspeaker to produce the audio output signal based on the output signal of the acoustic transparency filter.

20. The method of claim 19, wherein, The input of the acoustic transparency filter is further based on hearing compensation data, wherein the hearing compensation data is based on a hearing map of a particular user.

21. The method of claim 19, wherein, Hearing compensation data is applied to the output signal of the acoustic transparency filter, wherein the hearing compensation data is based on a hearing map of a particular user.

22. The method of claim 19, further comprising receiving a reproduced audio signal, wherein, The audio output signal comprises the reproduced audio signal, and wherein a relationship between the external microphone signal and the acoustic transparency filter varies when the reproduced audio signal is non-active.

23. The method of claim 19, wherein, The relationship between the external microphone signal and the acoustic transparency filter varies in response to a change in placement of a device within an ear canal.

24. The method of claim 20 or 21, wherein, The hearing compensation data is selected from among a set of hearing compensation data corresponding to a plurality of users based on a signal that identifies the particular user.

25. The method of claim 24, wherein, The signal that identifies the particular user is produced based on a voice authentication operation.

26. The method of claim 24, wherein, The signal that identifies the particular user is produced based on a facial recognition operation.

27. The method of claim 19, further comprising: producing a feedback component that is out of phase with the internal microphone signal, wherein the audio output signal is further based on the feedback component.

28. An apparatus for audio signal processing, the apparatus comprising: means for receiving an external microphone signal from a first microphone; means for receiving an internal microphone signal from a second microphone; means for updating filter coefficients of the adaptive filter based on a difference between the internal microphone signal and an output of the adaptive filter based on the external microphone signal; means for updating filter coefficients of an update filter based on the updated filter coefficients of the adaptive filter at a lower rate than at which the filter coefficients of the adaptive filter are updated; means for outputting an output signal from a transaudio filter coupled to the first microphone and configured to facilitate generation of acoustic transparency in an audio output signal produced by a speaker, wherein an input to the transaudio filter is based on the external microphone signal, and wherein the transaudio filter comprises a fixed portion and an adaptive portion, wherein the adaptive portion comprises the update filter and the adaptive filter; and means for causing the speaker to produce the audio output signal based on the output signal of the transaudio filter.

29. The apparatus of claim 28, wherein, an input to the transaudio filter is further based on hearing compensation data, wherein the hearing compensation data is based on an audiogram of a particular user.

30. The apparatus of claim 28, wherein, hearing compensation data is applied to the output signal of the transaudio filter, wherein the hearing compensation data is based on an audiogram of a particular user.

31. The apparatus according to claim 29 or 30, further comprising means for selecting the hearing compensation data from among a set of hearing compensation data based on a signal, wherein, the set of hearing compensation data corresponds to a plurality of users, and wherein the signal identifies the particular user.

32. The apparatus of claim 31, wherein, the signal identifying the particular user is produced by a biometric authentication operation.

33. The apparatus of claim 28, wherein, a relationship between the external microphone signal and the transaudio filter varies in response to a change in placement of a device within an ear canal of a particular user.

34. A non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause the at least one processor to: receive an external microphone signal from a first microphone; receive an internal microphone signal from a second microphone; update filter coefficients of an adaptive filter based on a difference between the internal microphone signal and an output of the adaptive filter based on the external microphone signal; update filter coefficients of an update filter based on the updated filter coefficients of the adaptive filter at a lower rate than at which the filter coefficients of the adaptive filter are updated; output an output signal from a transaudio filter coupled to the first microphone and configured to facilitate generation of acoustic transparency in an audio output signal produced by a speaker, wherein an input to the transaudio filter is based on the external microphone signal, and wherein the transaudio filter comprises a fixed portion and an adaptive portion, wherein the adaptive portion comprises the update filter and the adaptive filter; and cause the speaker to produce the audio output signal based on the output signal of the transaudio filter.

35. The non-transitory computer-readable storage medium of claim 34, wherein, an input to the transaudio filter is further based on hearing compensation data, wherein the hearing compensation data is based on an audiogram of a particular user.

36. The non-transitory computer-readable storage medium of claim 34, wherein, hearing compensation data is applied to the output signal of the transaudio filter, wherein the hearing compensation data is based on an audiogram of a particular user.

37. The non-transitory computer-readable storage medium of claim 35 or 36, wherein, The hearing compensation data is selected from a set of hearing compensation data based on a signal, wherein the set of hearing compensation data corresponds to a plurality of users, and wherein the signal identifies the particular user based on biometric authentication.

38. The non-transitory computer-readable storage medium of claim 34, wherein, The relationship between the external microphone signal and the sound-transmission filter varies in response to a change in placement of the device within the ear canal.

39. A computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 19 to 27.

Citation Information

Patent Citations

  • Personal communication device with hearing support and method for providing the same

    US20130243227A1

  • Granting access rights to a sub-set of the data set in a user account

    US20170318400A1